---
title: Session CPU cap (Windows soft-cap + below-normal priority)
---

# Session CPU cap (Windows soft-cap + below-normal priority)

## What it is

Omniscio can have 20+ Claude CLI sessions running at the same time, and each one of those can spawn its own subagents (Task / Agent), MCP servers, `npm test`, vitest, builds — every child process that the CLI starts. On a contended host this aggregates into a CPU pile-up: foreground apps stutter, fans roar, the user notices the machine "feels owned by Claude". Even at lower session counts, a single runaway subagent loop can grab every core.

The default Windows scheduler treats every Omniscio child equally to the user's browser, IDE, and Zoom call. There's no built-in throttle.

## Where to find it

### Where to find the toggle

**Settings → Performance section**:

- **Cap CPU when many sessions are running (Windows only)** — master toggle (off by default).
- **Max combined CPU %** — number input revealed when the toggle is on. Range 25–90, step 5, suffix `%`. Lower = friendlier to foreground apps, slower for batch work. Higher = closer to native speed.
- **Strict ceiling — never exceed the cap, even on an idle machine** — toggle revealed when the cap is on (off by default). Opts into the hard cap (see [Strict ceiling](#strict-ceiling-hard-cap) below).

All reapply live without restart — toggling on, changing the percent, or flipping strict takes effect on the next scheduler tick.

## How it behaves

### What Omniscio does (Windows only)

When the user enables **Settings → Performance → "Cap CPU when many sessions are running"**, Omniscio asks the Windows kernel to enforce two things on the existing Job Object that already owns every long-running Omniscio child (see [Job Object orphan-kill](job-object-orphan-kill.md) — same job, two new attributes):

1. **Soft CPU rate cap.** A target percentage (default 50%, configurable 25–90% in 5% steps) of total host CPU that the entire Omniscio process tree can use **combined** when the host is contended. "Soft" means: when nothing else on the machine wants CPU, Omniscio can still use the full machine — the cap only kicks in once other processes are competing. An isolated single session is not punished.
2. **Below-normal process priority class.** Every child inherits `BELOW_NORMAL_PRIORITY_CLASS`, so the Windows scheduler hands the user's foreground apps (browser, IDE, Zoom, games) a slice before Omniscio's children even when the cap isn't engaged. This is the "yield to the user" half — quiet by default, present even at low session counts.

   Note: below-normal priority does **not** actually depend on this cap toggle. A separate, default-ON setting — `sessionTreeBelowNormalPriority` ([system-settings.ts](../../src/shared/types/settings/system-settings.ts), default `true`) — already stamps every Claude CLI child below-normal **per-PID at spawn time** via `applySessionTreeBaselinePriority` ([priority-boost.ts](../../src/main/process/priority-boost.ts)), called from [spawn-cluster-manager.ts](../../src/main/process/spawn-cluster-manager.ts) regardless of the cap. So out of the box your sessions already yield priority; enabling the cap adds the _combined-rate_ throttle (and the Job-Object-wide priority class) on top.

Both apply to the **whole tree**: sessions, the Task/Agent subagents they spawn, `npm test`, vitest, builds, MCP servers — anything that was assigned to the Job Object. Inheritance is transitive (a subagent's subagent is in the job too).

### When to enable it

Turn it on if any of these match:

- You routinely run **10+ concurrent sessions** and want the host to stay usable for browsing/coding/calls while they grind.
- A single session is firing parallel subagents (Task / Agent dispatch) and your machine becomes unresponsive.
- You leave Omniscio running overnight with scheduled cron jobs and recipe runs, and want the morning's CPU/heat profile to be friendlier.

Leave it **off** if:

- You only run 1–3 sessions at a time and want maximum throughput per session.
- You're on macOS or Linux (the toggle has no effect — see below).
- You're running large `npm run build` jobs and want them to finish as fast as possible (the cap will slow them under contention).

### Related performance toggles

Several adjacent toggles live in the same Performance section and address different symptoms — see [amc-priority-boost.md](amc-priority-boost.md) for the full breakdown (a sibling **[session process cap](#session-process-cap-fork-bomb-backstop)** — a concurrent-process fork-bomb backstop on the same Job Object — is documented at the end of this page):

- **Boost Omniscio interface priority** — raises Omniscio's own Electron processes to Above-Normal. Pairs well with this cap: cap holds CLI children at Below-Normal _combined-rate_; the boost lifts Omniscio's UI/main _above_ normal, widening the priority gap from one notch (normal vs below-normal) to two (above-normal vs below-normal). Best for users who feel Omniscio's window itself stutter under load.
- **Push sessions to efficient cores when CPU is high** — adaptive E-core delegation on hybrid CPUs (Intel 12th-gen+ / Core Ultra). When host CPU stays above 85% for 5s, sessions are confined to E-cores via Job Object affinity; released back to all cores when host drops below 60%. Complements the cap on hybrid hardware — the cap throttles per-second, the affinity physically removes them from P-cores while pressure is high.
- **Let sessions yield the disk to the app** — the DISK/MEMORY sibling of the below-normal CPU baseline (default ON, `sessionTreeLowIoPriority`). CPU priority class inherits to children, but IO priority and memory (page) priority do NOT — so [session-io-priority.ts](../../src/main/process/session-io-priority.ts) stamps every process in the session Job Object to IO=Low + memory=Below-Normal (spawn-time root stamp + a 30s job-PID sweep, recycle-proof via `IsProcessInJob`). Fleet disk reads queue behind the UI's page-ins, and Windows evicts the fleet's cached file pages before Omniscio's. This is the lever for the STARVED input-delay class (UI blocked on disk with CPU protections already on — the 2026-07-03 parallel-slowdown fix). Kill switch `AMC_DISABLE_SESSION_IO_PRIORITY=1`; invariants in [session-io-priority-contract.md](../../.claude/memory/contracts/session-io-priority-contract.md).

These toggles are independent — any combination is supported. The cap + priority-boost pair is the recommended starting point for most users; add adaptive E-core delegation if you're on an Alder Lake / Raptor Lake / Core Ultra machine. (Two always-on levers — the **below-normal session-tree priority** baseline and its **low disk/memory priority** sibling, both default-on — already run underneath all of these; see §2 above and the bullet just before this paragraph.)

### Soft-cap vs hard-cap (why isolated sessions aren't punished by default)

By default the cap is implemented via `JobObjectCpuRateControlInformation` with the `JOB_OBJECT_CPU_RATE_CONTROL_ENABLE` flag — but **without** `JOB_OBJECT_CPU_RATE_CONTROL_HARD_CAP`. This is deliberate. The kernel enforces the percentage only when there's contention: if the only thing wanting CPU is Omniscio's tree, the kernel lets it use 100%. The moment another process (browser, IDE, antivirus) wants a slice, the kernel starts throttling Omniscio down to the configured target.

The deliberate downside: a pure soft cap **can't bound total CPU** when Omniscio's own sessions are the only load — "no contention" means they keep the whole machine. Users who want a guaranteed ceiling can opt into the hard cap via the **Strict ceiling** sub-toggle below.

### Strict ceiling (hard cap)

**Settings → Performance → "Strict ceiling — never exceed the cap, even on an idle machine"** is a sub-toggle revealed under the cap (off by default). When on, Omniscio adds `JOB_OBJECT_CPU_RATE_CONTROL_HARD_CAP` to the rate-control flags, so the kernel enforces the percentage as a **true ceiling** — Omniscio's tree can never exceed it, even when the machine is otherwise idle.

- **Use it** when you want a hard guarantee — a laptop you need to keep cool and quiet, or a shared box where Omniscio must never dominate regardless of what else is (or isn't) running.
- **Tradeoff:** a single session you're actively waiting on can no longer burst to full speed on an idle machine — it's held at the cap like everything else. The default soft cap exists precisely to avoid that, so turn strict on only when the guarantee matters more than lone-session latency.
- **Windows only**, and — like every session limiter — **suspended while a global CPU Burst window is open** (a deliberate full-machine run lifts it, then it's restored on expiry). See the [CPU-burst governor contract](../../.claude/memory/contracts/cpu-burst-governor-contract.md), invariant **`hard-cap-follows-cap-and-strict`**.
- Reapplies live (no restart): the governor re-reconciles whenever the strict toggle changes, so `hard = cap && strict` is always current.

### Subagent inheritance (it's automatic)

When a Claude CLI session spawns a Task/Agent subagent via the harness, that subagent process is parented under the CLI process. The Job Object owns the CLI. New processes spawned by a job member are automatically members of the same job — that's how Windows Job Objects work. So the subagent inherits both the soft cap and the below-normal priority without Omniscio having to wire anything per-subagent.

Same logic applies to `npm test`, vitest workers, MCP server processes, and any other child the CLI fires.

### Why Windows only

The cap is a Win32-specific feature. macOS and Linux have:

- `nice` (priority adjustment) — analogous to the priority half, but per-process, not tree-wide; no kernel construct equivalent to a job.
- `cpulimit` / cgroups — closer to the rate-cap half, but requires elevated privileges and per-PID configuration.

The user-base impact doesn't justify shipping equivalents on the other platforms. The toggle is hidden on the **Performance** section but the helper text is explicit: "No effect on macOS or Linux." The renderer doesn't gate by platform — the backend just no-ops, so the UI stays consistent across machines a user might switch between.

### What happens if it fails

The wrapper is conservative. Any failure logs a warning and returns false; Omniscio continues normally. Possible failures:

- `initSessionKillJob()` returned false (Job Object init failed at startup — rare, e.g. `ERROR_ACCESS_DENIED` if Omniscio was launched inside another job). `applyJobResourceLimits()` becomes a no-op. The user-visible toggle still flips, but the cap doesn't engage. The log line from `initSessionKillJob` already documented the underlying failure.
- `SetInformationJobObject` failed for `JobObjectCpuRateControlInformation` or for `JobObjectBasicLimitInformation` (priority). Logs `[JobObject] setCpuRateControl(<pct>) failed: <code>` or `[JobObject] setPriorityClassBelowNormal(<bool>) failed: <code>`. The other half may still have applied; `applyJobResourceLimits` returns true only if both calls succeeded.

The cap is best-effort. The session keep-alive, orphan reaper, and graceful-shutdown paths are unaffected.

### How this is verified

- **Unit (non-Win shim)**: [tests/unit/process/job-object-resource-limits.test.ts](../../tests/unit/process/job-object-resource-limits.test.ts) — verifies the shim threads through to the Win32 impl when initialized, and short-circuits when not.
- **Unit (non-Win no-op)**: [tests/unit/process/job-object-noop.test.ts](../../tests/unit/process/job-object-noop.test.ts) — verifies `applyJobResourceLimits()` returns false on macOS/Linux.
- **Settings re-apply**: [tests/unit/services/settings-apply.test.ts](../../tests/unit/services/settings-apply.test.ts) — verifies the apply-settings pipeline re-reconciles the governor (calling `applyJobResourceLimits()` with the post-update cap / percent / hard) whenever the cap toggle, percent, or strict toggle changes, and not when no governor field is in the delta.
- **Manual smoke**: enable the toggle at 50%, spawn 10 sessions running a heavy prompt (e.g. "scan the codebase and write a report"). Watch Task Manager → Performance. Omniscio's process group should plateau near 50% of total host CPU. Then open Chrome and start watching YouTube. Chrome's CPU should stay healthy; Omniscio's tree should yield down further.

### Why these specific defaults

- **Off by default.** Most users run 1–3 sessions and would be confused if their first session felt artificially slow.
- **50% default percent.** Empirically leaves enough headroom for a browser + IDE + Zoom on a typical 8-core host. Users who want more aggressive yielding can drop to 40%; users who want max throughput at the edge of yield can raise to 80%.
- **25–90% range.** Below 25% the host starts starving the agent runtime itself; above 90% the cap is functionally absent so we don't pretend to enforce one.
- **Step 5.** Granularity finer than 5% isn't perceptible in scheduler behavior; coarser would feel constraining.

### Session process cap (fork-bomb backstop)

The CPU cap above bounds how much CPU a session's tree uses. A **second, separate** Job-Object limiter bounds a different failure mode: a single session spawning an unbounded **number** of processes at once — a runaway loop, a fork bomb, `xargs -P 500`, a `find … | parallel` that fans out to thousands. On a box running many agents, one such runaway can flood the scheduler and freeze the whole machine. This cap stops that at the kernel level, and — because the limit is enforced by Windows on the whole session tree — no shell trick, absolute path, or alternate tool can route around it.

**Settings → Performance → "Cap runaway processes per session (Windows only)"** — a master toggle, **default ON**. When on, each session's own scoped Job Object carries a `JOB_OBJECT_LIMIT_ACTIVE_PROCESS` ceiling (default **512**). At the ceiling the kernel simply FAILS the next `CreateProcess` in that session's tree — the shell reports a spawn error and the agent moves on; **nothing already running is killed**, and other sessions are untouched. 512 is far above any normal session (typically 5–20 concurrent processes, ~120 even for an aggressive parallel job), so it only ever bites a genuine runaway.

- **Tune / disable via env**: `AMC_SESSION_MAX_PROCESSES=<n>` sets the ceiling; `AMC_DISABLE_SESSION_MAX_PROCESSES=1` turns it off. A garbage / non-positive value falls back to 512 — never a 0-process lockout. Either the toggle OR the kill switch being "off" disables it.
- **Applied at spawn**: the toggle / env take effect on new and restarted sessions, not already-running ones. (Unlike the CPU cap — which re-applies live because it's one process-wide call — this is per-session and sits at a ceiling that essentially never binds, so a live re-apply wasn't worth the cost.)
- **Honest scope**: it caps CONCURRENT processes, not the RATE of process creation. A strictly-serial fast spawner (e.g. `find … -exec cmd {} \;` one file at a time) stays low-concurrency and slips this limit — that pattern is handled by other guards.
- **Scoped job only, fail-open**: the ceiling lives on the per-session scoped Job Object, never the process-wide singleton (so it can never disturb the orphan-kill safety net that owns `KILL_ON_JOB_CLOSE`); any failure is a silent no-op that never blocks a spawn. Windows only — a no-op on macOS/Linux.

**Implementation**: [session-max-processes.ts](../../src/main/process/session-max-processes.ts) (the pure policy resolver — toggle + env → ceiling), [job-object.ts](../../src/main/process/job-object.ts) / [job-object-win32.ts](../../src/main/process/job-object-win32.ts) (`applyActiveProcessLimitTo`), wired at the shared spawn chokepoint `assignJobAndRecordSpawn` in [spawn-cluster-manager.ts](../../src/main/process/spawn-cluster-manager.ts). Invariants + the scoped-vs-singleton rationale live in [backend-spawn-contract.md §4](../../.claude/memory/contracts/backend-spawn-contract.md). Verified by [job-object-active-process-cap.test.ts](../../tests/unit/process/job-object-active-process-cap.test.ts) (resolver + public wrapper) + the `applyActiveProcessLimitTo` describe in [job-object-win32-workingset-affinity.test.ts](../../tests/unit/process/job-object-win32-workingset-affinity.test.ts).

## For agents

### Implementation pointers

- [src/main/process/job-object.ts](../../src/main/process/job-object.ts) — cross-platform public API: `applyJobResourceLimits(enabled, percent, hard)`. No-op on non-Windows or when Job Object init failed.
- [src/main/process/job-object-win32.ts](../../src/main/process/job-object-win32.ts) — koffi wrappers around `SetInformationJobObject` for both `JobObjectCpuRateControlInformation` (the soft cap) and `JobObjectBasicLimitInformation` (the priority class). Both load `kernel32.dll` lazily.
- [src/main/index.ts](../../src/main/index.ts) — applies the configured cap on startup, immediately after `initSessionKillJob()`, when `sessionCpuCapEnabled` is true in the loaded settings.
- [src/main/services/settings-apply.ts](../../src/main/services/settings-apply.ts) — re-reconciles the governor live (via `reconcileJobGovernor()`) when `sessionCpuCapEnabled`, `sessionCpuCapPercent`, or `sessionCpuCapStrict` changes via Settings, without restart.
- [src/renderer/src/features/settings/PerformanceSettings-definitions.tsx](../../src/renderer/src/features/settings/sections/performance/PerformanceSettings-definitions.tsx) — toggle + number + strict-toggle definitions in `PERFORMANCE_DEFINITIONS` (imported and rendered by [PerformanceSettings.tsx](../../src/renderer/src/features/settings/sections/performance/PerformanceSettings.tsx)). The percent and the strict toggle are both gated `visibleWhen: (s) => s.sessionCpuCapEnabled === true`.

## Related

[amc-priority-boost.md](amc-priority-boost.md) lays out the whole family of performance toggles this cap belongs to — interface priority boost, adaptive E-core delegation and the disk and memory sibling — and how each differs. [job-object-orphan-kill.md](job-object-orphan-kill.md) explains the Windows job object these limits ride on, the same one that keeps a session's children from outliving it.
