---
title: Job Object orphan-kill (Windows)
---

# Job Object orphan-kill (Windows)

## What it is

### What problem this solves

When Omniscio spawns a Claude CLI session, that CLI process spawns its own children — MCP servers (playwright, gcloud, mempalace), an OS shell, sometimes a network tool. If Omniscio dies cleanly, Omniscio's graceful-shutdown code uses `taskkill /F /T` to walk the tree and end each process. That works for a polite quit.

It does **not** work for the failure modes that matter:

- Omniscio's main process **crashes** (uncaught exception, segfault in a native module).
- The user **force-kills** Omniscio from Task Manager or `Stop-Process -Force`.
- The machine **BSODs** or loses power.
- A debugger detaches Omniscio's process tree.

In any of those, Omniscio has no chance to run its `taskkill /F /T` cleanup. Every Claude CLI process and every MCP grandchild is left behind as an orphan. The user opens Omniscio again the next morning and finds a dozen `claude-code` processes burning CPU, holding file locks, blocking ports.

The orphan reaper that runs on next Omniscio startup did clean these up eventually — but only when Omniscio was launched again. Until then, the machine kept paying.

## Where to find it

There is nothing to switch on and nothing to see — it is a safety net under the sessions you already run, on **Windows only**. Its effect is what you *don't* find the next morning: a machine not left carrying a dozen stray processes from an app that died overnight.

## How it behaves

### What Omniscio does now (Windows only)

On startup, Omniscio asks the Windows kernel to create a single **Job Object** with the flag `JOB_OBJECT_LIMIT_KILL_ON_JOB_CLOSE` set. A Job Object is a kernel construct — a named bag of process IDs the kernel manages on Omniscio's behalf. The flag means: **when the last handle to this job is closed, the kernel terminates every PID in the job.**

The only handle to that job is held by Omniscio's main process. So:

| Omniscio exit path | Handle state                                           | Kernel action                                    |
| ------------------ | ------------------------------------------------------ | ------------------------------------------------ |
| Clean quit         | `taskkill` already killed children, then handle closes | KILL_ON_JOB_CLOSE has nothing to do              |
| Crash              | Process death closes all handles                       | Kernel terminates every assigned PID             |
| Force-kill         | Process death closes all handles                       | Kernel terminates every assigned PID             |
| BSOD / power loss  | All handles invalidated                                | On reboot, no orphans existed in the first place |
| Debugger detach    | Handle still owned by Omniscio                         | No effect — same as running                      |

After the job is created, every long-running spawn Omniscio issues calls `assignToSessionJob(child.pid)` immediately after `spawn()` returns. The new PID joins the job. **Agents go one step further:** a post-spawn assignment can land after `cmd.exe` has already started the real agent, so every agent launch goes through the job-gate launcher (`amc-hidden-launch.exe`), which cannot start the agent until Omniscio has put it inside the job and signalled a release event. If that can't be done, the launch is refused (backend-spawn-contract rule 4). Job membership is **inherited transitively**: if the assigned process spawns its own children (like an MCP server), those grandchildren are automatically in the job too — unless someone explicitly sets `JOB_OBJECT_LIMIT_BREAKAWAY_OK`, which Omniscio never does.

### Which spawns are covered

The main long-running spawn surfaces (local agents are held until they are inside the job; the SSH wrapper for a remote agent is still assigned after it starts):

| Spawn site                                             | What it spawns                                                | Held until in the job? |
| ------------------------------------------------------ | ------------------------------------------------------------- | ---------------------- |
| `src/main/process/process-manager.ts` (main path)      | The persistent Claude CLI for each session                    | Yes                    |
| `src/main/process/process-manager.ts` (SSH path)       | The local SSH client wrapping a remote Claude CLI             | No (post-spawn)        |
| `src/main/services/engines/codex-app-server-client.ts` | The OpenAI Codex app-server child                             | Yes                    |
| `src/main/services/engines/pi-rpc-client.ts`           | The `pi --mode rpc` child (one per session)                   | Yes                    |
| `src/main/services/engines/opencode-server-supervisor.ts` | The one shared `opencode serve` server                     | Yes                    |
| `src/main/process/gemini-acp-client.ts`                | The long-lived `gemini --acp` ACP child (one per session)     | Yes                    |
| `src/main/process/slim-acp-session-client.ts`          | The Kimi / Hermes ACP child (one per session)                 | Yes                    |
| Cursor / Grok / Antigravity turn runners               | One child per turn                                            | Yes                    |
| `src/main/process/aside-runner.ts`                     | One-shot `claude -p --resume <id> --fork-session` per aside   | Yes                    |
| `src/main/services/context-query.ts`                   | One-shot CLI run that asks the agent for a summary            | No (post-spawn)        |
| `src/main/services/contextdock/spawn.ts`               | Vendored ContextDock CLI invocations (`runCli` / `runCliRaw`) | No (post-spawn)        |

**Out of scope (intentionally not assigned)**:

- `git clone`, `npm install`, `taskkill`, and other short-lived utility spawns. They finish in seconds and don't survive Omniscio by definition.
- User-launched bookmarks (`Start-Process` from the bookmarks popover). The user wants those to survive Omniscio closing; that's the whole point of a bookmark.
- Anything spawned by the renderer (the renderer can't spawn).
- **Cron script children** (`cron/cron-script-executor.ts`) — deliberately in a SEPARATE, non-killing job instead (2026-07-08 hang mitigation, explicit user decision): a **cron background job** with the CPU rate cap + BELOW_NORMAL priority class but **no** `KILL_ON_JOB_CLOSE`, so a mid-run cron script (e.g. an MP3 conversion) survives an Omniscio restart. The cron engine's own timeout + zombie-run machinery bounds them instead. The spawn audit reports these as `inJob: false, backgroundJob: true`. See [backend-spawn-contract.md](../../.claude/memory/contracts/backend-spawn-contract.md).

### What happens if it fails

The wrapper is conservative. Any failure logs a warning and returns false; Omniscio continues normally. Possible failures:

- `CreateJobObjectW` failed at startup — `initSessionKillJob()` returns `false`. Reaper-on-next-startup is the fallback. Logs `[JobObject] Init returned false — orphan-on-crash safety net DISABLED`.
- `AssignProcessToJobObject` returned `ERROR_ACCESS_DENIED` — Windows refused the job's nesting (on 2026-09-24 this turned out to depend on which child joined the job first, so the job is now pinned with an anchor child before anything else can join it). **Every local agent launch refuses to start** rather than run an agent the app cannot end — a Claude session reports "Windows process-tree ownership could not be established before agent launch", and Codex, Pi, OpenCode, Gemini, Kimi/Hermes, Cursor, Grok, Antigravity and asides show "Failed to start <engine> session: Windows process-tree ownership …". Other children log the error code and continue unprotected; the reaper covers them.
- `koffi` failed to load `kernel32.dll` — basically impossible on Windows but caught by try/catch. Same outcome: warn, continue, reaper covers.

The reaper is still wired and runs on every startup. It's no longer the primary defense, but it's the second line.

### Non-Windows behavior

`initSessionKillJob()` returns `null` on macOS and Linux. Those platforms reparent orphaned children to init / launchd, which immediately reaps them. There's no orphan-burning-CPU class of bug to solve, so no analogue is needed.

### Why not just rely on the reaper?

The reaper only runs when Omniscio is launched. If Omniscio crashes and the user doesn't reopen it for hours — or never reopens it on that machine — the orphans burn CPU, hold file locks, and rack up API token usage in the background until the user notices. Job Objects fix the leak at the kernel level so it can't happen at all.

### If an agent survives anyway

Two backstops catch an agent that escaped the job (on 2026-09-24 one did: a since-reverted change let agents launch outside it, and a restart's shutdown ran out of time):

- **The orphan reaper** reads the app's whole launch record — every rotated copy of `spawn-audit.log`, about the last 10–12 hours on a busy machine, not just the newest two — so it can prove an agent is Omniscio's even when it was launched hours before the app died. It ends only what it can prove, and never a process that started under a reused process number after that app run had ended.
- **A resumed session ends its own leftover copy first.** When a session starts again after a restart, it looks up its last launch from the previous app run and ends that copy (with its tools) before the new one takes over, so one session never runs two agents on one branch. If the old copy was still running real work, such as a build, the session is told what was ended.

### Why a job per Omniscio instance, not per session?

Per-session jobs would require one job handle per running session, plus careful cleanup when sessions end normally. The shared single-job design has one handle for Omniscio's whole lifetime, which is exactly what we want: when Omniscio ends, every Claude process ends. We don't need finer-grained control because session-end already calls `taskkill /F /T` for that specific PID tree.

### What a session stop ALSO ends: the agent's background work

Job membership is **inherited and permanent**. Every process a session's Claude CLI starts — the OS
shell, and anything that shell runs — joins the session's job automatically, and nothing can leave a
job it is already inside. Probed directly on the owner's box: a fresh `node.exe` started from a
session shell reads `IN-JOB`.

That has a consequence worth stating plainly, because it cost a week of misdiagnosis in September
2026:

> When a session's child is stopped, `terminateScopedJob()` calls the kernel's `TerminateJobObject`.
> That ends **every** process in the job — including a long-running command the agent deliberately
> started in the background (`run_in_background`) and is waiting on, such as a build or a test run.

A kernel job terminate is completely silent. There is no exception, no exit code, no signal handler,
no crash event, and no surviving process. To the agent, a 15-minute build simply **stops writing to
its log** — which reads exactly like a hung gate, and got diagnosed as one three times.

**What Omniscio does about it.** It cannot save that work once a stop happens: a process cannot be
excluded from a job terminate it is already inside, and the alternative (skipping the job kill and
the tree kill) would disable the orphan protections this whole page exists for. So it avoids the
OPTIONAL stops, and makes the rest **honest**:

- A per-session **live-task ledger** (`live-background-tasks.ts`) records each background command
  the agent launched, from the stream the app already reads, until the CLI reports it finished —
  a `system`/`task_notification` event on its stdout, sent whether the agent is mid-turn or idle.
- **Optional stops wait or skip.** Idle release and the account-switch idle kill skip a session with
  a live command; a sign-in re-run on an automated turn waits for the command to finish — it goes
  ahead the moment the command is reported finished, and after 30 minutes at most, so a watcher that
  never ends cannot keep the session signed out (a message from a person, or from another agent that
  arrived while the command was already running, still goes through at once).
- **Every other stop announces what it ended, before the job kill** — the commands the ledger holds,
  each with its own runtime and the stop's real reason (`stop-reasons.ts`). A stop that ended
  nothing says nothing.
- `reapEscapedSessionTree()` reports the WORK processes it had to reap to the **session**, not only
  to `main.log` — never the leftover agent processes.
- Both use the shared `endedWorkNotice()` line.

**If you need a long build to survive a session restart**, run it detached — a Windows Scheduled Task
is outside the session's job and outside its process tree, so neither kill can reach it.

## For agents

### How this is verified

- **Unit**: [tests/unit/process/job-object-noop.test.ts](../../tests/unit/process/job-object-noop.test.ts) — three tests that mock `os.platform` to `darwin` and assert `initSessionKillJob()` returns `null` and `assignToSessionJob()` returns `false` on non-Windows.
- **Integration (Windows-only)**: [tests/integration/job-object-orphan-kill.test.ts](../../tests/integration/job-object-orphan-kill.test.ts) spawns a standalone host helper script that creates a real Job Object, assigns a long-running child, then hard-exits. The test reads the printed PID and asserts the kernel killed the child within 500 ms. A control case uses a bogus PID to confirm the "child should be alive" branch can fire (protects against false-positives if `pidExists` ever broke).
- **Manual smoke**: dev launch Omniscio → spawn a session running a 30s prompt → `Get-Process electron* | Stop-Process -Force` → verify zero `claude-code` rows in the process table.

### Implementation pointers

- [src/main/process/job-object.ts](../../src/main/process/job-object.ts) — cross-platform public API: `initSessionKillJob()`, `assignToSessionJob(pid)`. No-op on non-Windows.
- [src/main/process/job-object-win32.ts](../../src/main/process/job-object-win32.ts) — koffi wrapper around `CreateJobObjectW`, `SetInformationJobObject`, `AssignProcessToJobObject`, `OpenProcess`, `CloseHandle`. Loads `kernel32.dll` lazily.
- [src/main/index.ts](../../src/main/index.ts) — `initSessionKillJob()` is called once at startup, BEFORE the orphan reaper and the resume-sessions step. The job exists before any spawn could happen.

### Known gotcha: koffi struct marshalling

The koffi FFI library (used to call Win32 functions without a native addon) cannot auto-marshal a plain JavaScript object onto a `void*` parameter — it has no struct layout to use. The `SetInformationJobObject` 3rd parameter MUST be declared as `koffi.pointer(JOBOBJECT_EXTENDED_LIMIT_INFORMATION)`, not `LPVOID`. If declared as `LPVOID`, koffi throws `TypeError: Unexpected Object value, expected void *`, `initJob()` catches it, and the safety net silently disables.

The same fix applies to the test helper at [tests/helpers/job-object-orphan-kill-host.cjs](../../tests/helpers/job-object-orphan-kill-host.cjs).

## Related

What Omniscio does on a normal shutdown, and how a session is stopped by hand, is described on the [Pause or stop a session](pause-or-stop-a-session.md) page. Crash recovery after an unexpected exit is on [Crash recovery](crash-recovery.md).
