---
title: Job Monitor (test/build/dev process dashboard)
---

# Job Monitor (test/build/dev process dashboard)

## What it is

A read-only sidebar dashboard inside Omniscio that surfaces every long-lived test, build, dev-server, and sandbox process held by the filesystem-lock concurrency gates — plus any vitest / Playwright process the gates missed because the runner forgot to acquire a slot. The dashboard runs alongside the **Resources** diagnostic panel but is dedicated to the parallel-worktree-agent workflow: when 10–20 Claude Code agents are working concurrently in their own worktrees and one of them spawns a runaway test loop, the Job Monitor is where you spot it and kill it. It also shows the **wait queue** — the agents blocked waiting for a free slot — as a section in the dashboard and a small ⏳ badge on each waiting session in the sidebar (see "Wait queue" below).

**Why this exists.** You run many Claude agents in parallel across worktrees. Each agent that runs `npm run test:agent` reserves a slot under `%TEMP%/amc-test-slots/` (3 slots total, default) and each Electron-spawning E2E spec reserves a slot under `%TEMP%/amc-e2e-slots/` (2 slots). If an agent crashes mid-test, or a vitest worker hangs after the parent exited, or someone runs bare `vitest` outside the gate, a slot can stay held by a process that no longer exists — or a process can be running without holding any slot at all. Both states cause "why is my CPU pinned at 100% with no visible work" surprise. The Job Monitor reads the slot directory plus a walker that scans every `node.exe` / `electron.exe` for vitest- and Playwright-shaped command lines, joins the two, and shows the user a flat list with Kill and Release-dead-slot actions.

## Where to find it

As a **sidebar tile**, as a compact row strip inside the **Resources** diagnostic panel, and as a **floating always-on-top window** you can keep beside your work. On the web and mobile it is read-only and reduced.

## How it behaves

### Sidebar tile

**Job Monitor** appears in the left sidebar with a small activity-pulse icon. It is a built-in tile, not a real folder: clicking it opens the panel and nothing else, and you cannot start a session inside it.

### Three view states

Like the Resources panel, the Job Monitor mirrors the state of the underlying subscription:

1. **Loading** — first mount, before the SUBSCRIBE + GET_INITIAL round trip resolves. Shows a centered "Loading job monitor…" message.
2. **Unavailable** — the IPC handler is missing or rejected. Shows "Job monitor unavailable on this build." (partial-deploy gate against an old main process without the handler; fails closed instead of spinning forever). **Web / mobile clients no longer hit this** — the panel works over the web bridge now (see "On web and mobile"); earlier the web subscribe threw because the handler expected a desktop window, which surfaced as this message.
3. **Ready** — initial tick arrived. Renders the header counts and the job list. Subsequent ticks update in place via the `job-monitor:tick` push channel.

### What you see in the Ready state

- A header with the active count, plus a `· N stale` and `· N dead` summary when those counts are non-zero. **"Active" means genuinely running** — it EXCLUDES the stale and dead rows listed beside it, so 3 running + 1 stale + 2 dead reads `3 active · 1 stale · 2 dead`, not `6 active`. (Those counts arrive from the service as subsets of one `jobs` array; `activeCount` is derived once in `useJobMonitorSubscription` so the full panel and the Resources-pane strip can't drift apart.) A **Pop out** button on the right opens the always-on-top floating window (see "Floating window" below).
- An empty-state message ("No active jobs. Runaway vitest, build, and sandbox processes will appear here.") when nothing is running.
- A column header strip (Type · Label · Workspace · PID · Uptime · CPU · Mem · Status · Actions) followed by one row per job, sorted attention-worthy first (`dead > stale > running`, oldest within bucket).
- Each row shows the job type chip (`vitest` / `e2e` / `sandbox` / `build` / `typecheck` / `lint` / `dev` / `other`), the runner label, the workspace path (last two segments, full path on hover), the PID (clickable to copy), uptime, CPU percent of one logical core, resident memory, and a status chip (`running` / `stale` / `dead`).
- **Kill** button on running and stale rows — hard-kills the entire Windows process tree rooted at the row's PID via `taskkill /F /T`. Confirm dialog summarises the consequence ("In-flight test output, and any uncommitted changes inside the runner, will be lost"). The copy lives in ONE translated builder (`job-confirm-copy.ts`) shared with the Resources-pane strip, so the two dialogs for this one action can't drift into making different safety promises.
- **Release slot** button on dead-status rows whose slot is still held — frees the sidecar lock and the heartbeat sibling. Hidden on alive jobs (must Kill first); hidden on walker-discovered ungated strays (no slot to release). Confirm dialog explains the action is safe — the process is already gone. Same shared builder.

### Wait queue — what's waiting for a slot

The dashboard surfaces both halves of the slot system. The job list above is everything _holding_ a slot; the **wait queue** is the mirror image — the agents _blocked_ waiting for one to free.

When 10–20 agents share 3 test slots (and 2 E2E slots), the ones that don't get a slot immediately sit in a poll loop until one frees. Before this, that waiting was invisible: you'd see your machine busy but couldn't tell which sessions were stuck in line or for how long. The wait queue makes it visible in two places:

- **A "Waiting for a slot" section** at the top of the Job Monitor panel (and its pop-out window). One row per blocked agent, longest wait first: the session name · which pool it's queued on (`test pool` / `e2e pool`) · the command it's trying to run · a live "waiting 2m 14s" timer. It is read-only — a waiting agent is queued work, not a runaway, so there is no Kill button. When nothing is waiting it reads "No agents waiting — slots are free."
- **An amber ⏳ badge** on each waiting session in the sidebar, showing the wait in minutes (just the hourglass under a minute). This is the glanceable signal — you can see a session is stuck in line without opening the Job Monitor at all. It clears the instant that session gets a slot.

**How it knows.** While an agent is blocked, its gated runner drops a tiny "I'm waiting" file (which session, which pool, what it's running, since when) and deletes it the moment it acquires a slot or exits. The backend scans those files on a cheap ~3-second directory poll — no process walking, so the sidebar badges stay live the whole time Omniscio is open at negligible cost (this is a _separate_ feed from the job-list discovery, which is heavier and only runs while the panel is open). A waiter is shown only while its process is alive and is not already holding a slot, so the brief moment between "got a slot" and "deleted my waiting file" never shows a phantom. Runs started outside Omniscio (no session id) still appear in the panel by their workspace, just without a sidebar badge.

Like the job list, the wait queue also rides into the always-on-top pop-out window — and (as of the 2026-06 web-parity change) into the browser / phone client too, so the badge and the "Waiting for a slot" section are visible from any device (see "On web and mobile" below).

### Compact row strip (Resources panel embed)

A second presentation — `JobMonitorRows` — lives below `ProcessTree` in **Settings → Tools & Maintenance → Diagnostics** (Resources tab). It uses the same `useJobMonitorSubscription` hook (the backend ref-counts subscribers, so two consumers cause no extra work), but renders one compact line per job and drops the Workspace column (the label already names the worktree). No Pop-out button, no Release-dead-slot button — the full panel handles the rarer slot-release case. Kill confirms + IPCs through the same path. The user can find a runaway runner without leaving the Resources surface they're already on.

### Floating window (always-on-top)

The **Pop out** button in the panel header invokes `JOB_MONITOR_TOGGLE_WINDOW`. The main process lazy-creates a frameless always-on-top `BrowserWindow` on first click via `src/main/services/job/job-monitor-window.ts`, registered with role `'job-monitor'` in the window registry, and loads a dedicated renderer entry `src/renderer/src/job-monitor.tsx` that mounts the same `JobMonitorPanel` component — no `App.tsx`, no shared providers, no sidebar chrome. Subsequent clicks alternate show / hide: visible-but-unfocused → focus; visible-and-focused → hide. Because the floating window renders the same panel, clicking **Pop out** inside the focused floating window flips to hide — naturally giving the button a "close" affordance there.

The push to that detached window is deliberately narrow: only the job list and the wait queue cross to it, so both halves render in the pop-out and nothing else follows you out.

### On web and mobile

The Job Monitor works in the browser / phone client, not just the desktop app — so you can spot and kill a runaway test process on your PC from anywhere. Both feeds reach web clients, so the panel, the "Waiting for a slot" section and the sidebar ⏳ badge all render live on web/mobile.

**Kill / Release dead slot / Clean up orphans** all act on the real processes on your PC, from any device. The only desktop-only piece is the **Pop out** button — a floating always-on-top window has no meaning on a phone — so it is hidden on web (gated behind `isElectron` in `JobMonitorPanel`).

When you navigate away from the panel the renderer unsubscribes; and if the tab closes hard (swipe-away / network drop), the WebSocket disconnect cleanup replays the unsubscribe so the backend process sampler stops once the last web viewer is gone. One trade-off: while anyone has the monitor open, the ~3s/15s job tick is broadcast to every connected web client (a client not viewing the panel just ignores it) — emission is subscriber-gated at the source, so nothing runs when no one is watching.

## For agents

**Sidebar tile identity.** The tile is registered as an **amc-builtin** virtual project in
`src/shared/integration-registry.ts`, with the Lucide `Activity` icon (amber-500 — a heartbeat cue beside Inbox
Pilot's amber-400). Its virtual-project sentinel is `__job_monitor__`, and it is non-spawnable by construction: no
case in `resolveProjectWorkDir()`, no entry in `SESSION_HOSTING_VIRTUAL_PROJECT_PATHS`, and
`isSpawnableProjectPath()` refuses it.

**Push allowlists.** The detached window's fan-out is `JOB_MONITOR_ALLOWLIST` in
`src/main/services/detached-push-filter.ts`, which contains exactly two channels — `IPC.JOB_MONITOR_TICK` (the
job list) and `IPC.JOB_WAITERS_TICK` (the wait queue). See `job-monitor-push-listener-coverage.test.ts` for the
lint that pins both directions. On the web bridge the same two channels ride `INBOX_ALLOWED_CHANNELS`, and the
renderer hooks send a stable per-tab `subscriberId` with their subscribe/unsubscribe calls so the backend can
ref-count the subscriber over the WebSocket bridge — there is no Electron `webContents` id there, and that
mismatch is what used to surface as "Job monitor unavailable on this build".

### How discovery works

The service ticks at two cadences, both subscriber-gated (zero subscribers → zero ticks → zero idle CPU cost):

- **HEAVY tick (~15s)** — rediscovers everything AND samples CPU/memory. Walks both slot directories (`%TEMP%/amc-test-slots/` + `%TEMP%/amc-e2e-slots/`) reading every `.lock`/`.meta.json`/`.beat` triple — both filesystem slot pools were retired on 2026-09-03 and nothing writes them any more, so this walk finds only the detached sandbox's slot (if one is running) until the discovery residue is removed; it is inventoried as `retired-residue` in the [agent-load governance registry](../../.claude/memory/agent-load-governance-registry.md) — runs one Windows `Get-CimInstance Win32_Process` query for vitest- and Playwright-shaped command lines across all `node.exe` / `electron.exe` PIDs, then `reconcileCache()` unions the two sources keyed by `jobId`. Immediately after reconciling, it runs ONE more scoped `Get-CimInstance Win32_Process` query for just the tracked PIDs to read each one's resident memory (`WorkingSetSize`) and CPU time — `cpu%` is computed from the change in CPU time between consecutive ticks (so a fresh row reads 0% until its second sample). Walker-discovered strays use `jobId = discovered:${pid}:${spawnTimeMs}`; slot-pool jobs use the slot path as the `jobId`. The walker skips Omniscio's own PID and every slot descendant so it never double-counts.
- **LIGHT tick (~3s)** — emit-only. Re-publishes the cached job list so uptime and status stay fresh between heavy ticks; it does no process sampling of its own. `reconcileCache()` on each HEAVY tick deliberately preserves the prior CPU/memory sample so the numbers don't blink to zero when a stats query transiently misses a PID.

**Why CPU/memory comes from a CIM query, not `pidusage`.** Earlier builds sampled per-process stats with the `pidusage` npm package on the light tick. On Windows + modern Node, `pidusage` always falls back to spawning a full `Get-WmiObject | format-table` PowerShell process **per PID** and treats any stderr as a fatal error — which silently failed inside Omniscio's windowless main process (the failure was logged only at debug level), so every CPU and memory cell rendered `0` even for long-running jobs. The fix routes stats through the same robust `Get-CimInstance … | ConvertTo-Json` mechanism (`-NoProfile`, exit-code-only, with a fallback) that the discovery walker already uses successfully in the same process — collapsing the old per-PID PowerShell storm (24+ spawns every 3 s) into one scoped query every 15 s. A failed stats query now logs a **visible warning** (not a hidden debug line) and keeps the last-known values. `pidusage` is still installed because the **Resources** panel uses it separately.

**Status semantics.** A row is `dead` when the PID is gone; `stale` when `lastBeat > 0` and `now - lastBeat > 60s` while the PID is still alive (the authoritative heartbeat path) or when no heartbeat ever arrived but `now - lastSampleMs > 60s` (the stats-sample fallback for walker strays and heartbeat-exempt sandbox slots); `running` otherwise.

### Heartbeat contract

Every slot-holding runner under `scripts/` MUST either call `writeHeartbeat()` from `scripts/lib/job-registry.mjs` on a 5-second `setInterval` cadence (and `clearJobMeta(slot.path)` in the release path so the heartbeat file is unlinked alongside the sidecar), OR appear in the ALLOW_LIST at `tests/unit/lint/runners-write-heartbeat.test.ts` with a short `reason:` explaining why it can't beat. The only current exemption is `scripts/sandbox-launch.cjs` — it hands its slot to a detached Electron child via `slot.rekey()` and exits, so there's no surviving event loop to tick a beat. Sandbox-window liveness is PID-based via the rekeyed lock; the Job Monitor's `stale` fallback path treats sandbox slots correctly.

The heartbeat file is a bare epoch-ms integer string (no JSON), written atomically via tmp+rename, sibling to the slot lock at `${slotPath}.beat`. Round-trip semantics pinned by `tests/unit/scripts/run-tests-gated-heartbeat.test.ts`.

## Related

- [resources-diagnostic-panel.md](resources-diagnostic-panel.md) — Settings → Tools & Maintenance → Diagnostics (Resources tab), the sibling read-only process monitor for Omniscio's own child processes (Claude sessions, MCP servers, asides). The Job Monitor focuses on test/build/dev/sandbox processes; Resources focuses on Claude session lifecycle.
- Omniscio’s internal testing guide (developer-only) — covers the slot-gated `npm run test:agent` and `launchElectronGated()` mechanisms that the Job Monitor surfaces.
- **Feature contract**: [`.claude/memory/contracts/job-monitor-contract.md`](../../.claude/memory/contracts/job-monitor-contract.md) — invariants, dead/museum code, postmortem ledger. READ before any change to discovery, classification, status semantics, the action surface, or the floating window's push allowlist.
