---
title: Performance Monitor (how the machine and the app are really doing)
---

# Performance Monitor — the app's one canonical performance readout

## What it is

### What it is

One answer to "how is this machine and this app actually performing right now?", available three
ways from the same measurement:

- a **panel** — Performance Monitor, open to everyone since 2026-09-11,
- an **HTTP route** — `GET /perf/status` on the local control server,
- a **command** — `npm run perf:status`.

All three read the same snapshot from the same place in the app, so **any number they both show**
agrees.

> **They do NOT all show the same numbers — but the whole-box gap is CLOSED.** The 2026-09-07/08
> false green (a spot `Kernel share 41%` with no duty beside it, reading busy-but-fine on a hanging
> box) was fixed: the panel now renders **`Kernel-bound share`** (`kernelBoundPct`, alarm at ≥50%),
> **`Storm verdict on`** (`stormOnPct`) and **`Sensor stood down`** (`standDownPct`) — see
> `PerfStatusPanel.tsx`, whose own comment records that their absence _was_ the green.
>
> **One duty is still CLI-only: `kernel-share SHADOW would-pace %`** (`kernelShareShadowPct`), which
> reads how much the kernel-share lever _would_ have paced while switched off. No renderer code
> references it. So the panel can answer "is the whole box hung?"; what it cannot yet answer is
> "what would the switched-off lever have done?" — for that, run `npm run perf:status`.

Alongside the readout, the app **watches** these numbers on its own and raises an inbox card when
one crosses its bound — so a machine in trouble reaches you without anybody looking.

## Where to find it

The **Performance Monitor** panel in the sidebar's **Developer Tools** group, open to everyone (how long your agents' actions take is its **Where agents spend their time** section); the same numbers over the local control server's performance route; and **`npm run perf:status`** from a terminal.

## How it behaves

### Why it exists

Performance _rules_ were already centralized: every throttle, gate and pacer in the app is a row in
one registry with one off-switch accessor. The **watching** was not. It lived inside a single
long-running AI session with a 30-minute alarm, a script under one person's home directory, and a
notes file that was that session's memory. Findings were passed to other sessions by message — and
when those sessions were archived, the findings went with them.

Worse, part of that instrument went blind exactly when it was needed. Three of its measurements
worked by starting a small external program, and when the machine ran short of capacity to start
programs — the precise condition they existed to detect — they timed out and returned nothing. A
measurement that dies of the thing it measures reports that thing as silence.

The rest of the same instrument kept working right through those episodes, because it read files
rather than starting programs. That contrast is the rule this is built on: **reading a log cannot
be starved by a program-starting pile-up; starting a program can.** Everything here reads logs and
values the app already holds.

### What it reports

| Block               | What it answers                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                             |
| ------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Box**             | CPU busy · kernel share · pressure faults per second · event-loop lag · RAM free **and** RAM available, labelled separately. RAM available is the reading taken at the **last stall**, not a live one, so it says when (`at the last stall, 14:05`, with the date once it is not from today, or `at the last stall, time unknown`) or why there is no number: `no stall in this window`, `no reading on the last stall`, or `stall log not read` — blind, which is **not** a calm window                                                                                                                                                                    |
| **Birth rate**      | How fast the whole box is manufacturing processes, **per live session** — demand-zero faults per second divided by the number of live agent sessions. Carries the threshold it was compared against, how many consecutive readings have been over it, and how many child launches the app's own spawn seam has made. Read it **beside** Box, never instead of it: they answer different questions, and they genuinely diverge — a box can be at 100% CPU without a birth storm, and a birth storm is what a box full of 200 ms processes looks like. `n/a` means the reading could not be taken, which is **not** a calm box.                               |
| **App main thread** | What share of the window the app's main thread was stalled, how many stalls, the longest one, memory in use, blocking file-system holds, garbage-collection pauses. Plus **what the app was actually doing** while frozen, from two separate instruments that are never merged — see “Two instruments name a freeze” below |
| **Memory pool**     | The external-memory pool size, how fast it is refilling, and confirmed refill episodes                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                      |
| **Fleet**           | Sessions by status, and how fast processes are being started, by program name                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                               |
| **Disks**           | Per volume: free space, **runway hours** at the measured rate of change, and live worktree directories. Box-wide loose directories are reported separately, never mixed in. The rate is a **fit across hours of samples, not the gap between two of them**, and it prints the span it actually used (`9.4 GB/h over 6.0 h`) — see “Runway is a trend, not a stopwatch” below                                                                                                                                                                                                                                                                                                                                                 |
| **Cloud verdicts**  | Attempted · green · red · **no verdict** — the last being a run that answered nothing, which is a different failure from a red                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                              |
| **Auto-lander**     | Lands in the window and per hour, branches waiting, git wedge events and refused reads                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                      |
| **Load gates**      | The harm verdict share, how often the older busy gate disagreed with it, and the pack-storm share                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                           |
| **Governor**        | Every governor policy on one screen, each as THREE separate facts — what it OBSERVED, what it REQUESTED on the strength of that, and what was actually APPLIED — plus the mode it is really in, the gap between the last two, and the card that raises if it fails. Covers fleet pacing, Job containment, maintenance yield and cooperative drain. A policy whose records cannot be read says so with the reason; it never shows a green it did not measure. Each policy also carries a **spread** line — how many sessions were actually being held back, and how hard the worst-off one was hit — which today reports that it cannot be measured, and why |
| **Hook host**       | The resident host that runs the agent hooks Omniscio injects into every session (so a tool call starts one tiny native program instead of a Node runtime): up/down/held and why, its process id and uptime, warm and in-flight workers, hooks served, held and fallen back (by reason), p50/p95, restarts, and whether THIS computer's sessions are contained by the native client at all — other systems read "not contained" and keep the per-call behaviour |
| **Agent action time** | Where your agents' time actually went, grouped by what they were doing and ranked by total wait — with every row's quiet-box and busy-box typical wait side by side, which is the pair that says whether an action is expensive or merely queued. Plus the model half, so a slow model is separable from a slow tool. See "Where the agents' own time went" below |
| **Work rate**       | What the fleet actually finished at each live-session count — completed turns, tool calls and branches landed per session-hour, beside lands per wall-clock hour — with the two peaks named as readings, the land record's own coverage stated, and a thin bucket printing counts rather than a rate. See "What the fleet finishes, by session count" below |

### Where the agents' own time went

**Agent action time** is the one block about your agents rather than your computer, and the one that answers "why do they feel slow".

It takes every tool call the agents made in the window and groups them by **what the agent was doing** — `Bash:git`, `Bash:test`, `Read`, sub-agents — then ranks those groups by the total time they consumed. Each row carries its call count, its typical wait, its slow-tail wait, its worst single call, its share of all agent waiting, and its error share.

**Every row also carries its typical wait on a quiet box beside its typical wait on a busy one, and that pair is the point rather than a detail.** The number here is what the agent *waited* — the hook time and the queue hold included — not what the command cost. Measured on this machine: an `echo`, which does nothing, has a typical wait of 883 ms with 0–25 sessions running and **10.6 seconds with 100–125**. Every class behaves the same way. So a ranking on its own would tell you shell commands are your bottleneck, which is the opposite of the truth: the box was holding everything.

Read it the same way you read the rest of this panel — the row that grows between its two columns is the box queueing work, and the row whose two columns stay close together while its total climbs is genuinely doing something expensive.

The block also reports the **model half**: the vendor's own per-turn thinking time, recorded beside the tool waits, so you can tell "the model is slow" apart from "the tools are slow". Until the app is restarted onto a build that records it, that half reads as *not measured* rather than as zero. A turn whose thinking time cannot be told apart from time carried over from earlier work, such as the first turn after a resumed agent relaunches, also reads as *not measured*, and it is left out of the model-versus-tools comparison instead of counting as one very long think.

It states on its own face what it cannot see: that this is agent wait and not command runtime, that it covers app-spawned sessions only, and that a sub-agent's calls are counted against the session that carries them.

### What the fleet finishes, by session count

**Work rate** is the second block about your agents, and the one that answers "at what load does the fleet stop getting more done". Where agent action time says how long each step *waited*, this block says how much work was actually *finished* while a given number of sessions was live.

It takes three records the app already keeps — completed agent turns and tool calls from the timing tape, branches the auto-lander recorded as landed, and the governor tape's once-a-second count of live sessions — and folds them into the same six live-session buckets the load curve uses. For each bucket it shows the minutes of tape covered, the session-hours accrued (live sessions × covered time) and the work finished per session-hour, beside branches landed per wall-clock hour. Both rates are shown on purpose: per session-hour says whether each session gets less done as load rises; per hour says whether the fleet as a whole finishes more or less.

A bucket with under ten covered minutes shows its counts and `thin`, never a rate, and a rate the fold withheld shows as a dash — never the word null and never a zero. An event whose session count could not be read is counted on its own `unknown` row rather than dropped into the quietest bucket. The land record states its own coverage — `complete`, `partial` with how far back it reaches, or `not measured` with the reason — so a window older than the record can never read as a slow lander.

The two peaks — the bucket with the most turns per session-hour and the bucket with the most branches landed per hour — are readings of the data. The block never recommends a session count; a falling rate at a higher bucket names where the app's own per-session cost has to shrink. It also says on its face that every bucket is time-correlated: the minutes the box spent at one load may be one incident, so compare a heavy bucket across days before reading it as a load effect. The same fold renders the `WORK FINISHED` section of `npm run perf:load-curve`, so the two can never disagree.

### Two instruments name a freeze, and they answer different questions

When the app freezes, the tape records it twice, and the readout prints both without ever folding one into the other.

- **`stall by`** is the *operation tracker*: a JS-level operation the app had started and can prove it was inside — a database write, a sweep. When it says `(off-cpu-no-blocker)` it is not failing; it is correctly refusing, because the thread was not running any app operation at all.
- **`stack top`** is the *CPU profiler*: the actual frames sampled during the freeze, including native ones the operation tracker structurally cannot see — `run (native)` is the database driver's synchronous write, `spawn (native)` is the operating system starting a child process on the main thread.
- **`stack site`** names the app function that CALLED that native frame. It is the only row that points at somewhere you can go and edit.
- **`stack conf`** says what the two rows above are worth.

The two routinely disagree and both are right. A 27-second freeze on 2026-09-17 reported `(off-cpu-no-blocker)` from the first instrument and `spawn (native) 45%` from the second.

### Read `stack conf` before you trust a percentage

**The percentages are a share of the profiling window, not of the freeze.** The profiler stays armed across a whole load window, which is typically far longer than the gap it is describing — a measured median of 28.6x, and at worst 569x. So `45%` means 45% of everything the profiler watched, and on a badly over-long window the top frame is often `(idle)`, meaning the profiler saw nothing in particular at all. `stack conf` states that ratio, how many freezes had no capture, and how many were led by an idle frame, so a number is never read as more than it is.

One number there moves the opposite way to the rest: when freezes the operation tracker declined to name start getting named by the profiler, the unexplained pile **shrinks**. That is coverage arriving, not freezes disappearing — the row says so in those words.

### Runway is a trend, not a stopwatch

A worktree volume is a **sawtooth**: agents create checkouts and the reclaim drain removes them,
so free space can swing by gigabytes every few minutes while its real trend is a few GB an hour.
So `runway hours` is fitted across **hours** of recorded samples, never measured between two of
them.

**The line prints the span it actually fitted** — `9.4 GB/h over 6.0 h` — because that span is not
the window in the header and nothing else on the page says so. Read `over 6.0 h` as the rate's own
basis, not the `event window` above it. Before 2026-09-22 the rate printed with no span at all,
under a header reading `(window 30 min)`; a reader took the two together, concluded our basis was
5.6x _shorter_ than their own multi-hour series, reversed a conclusion they had right the first
time, and carried the wrong version to the owner. The instrument was computing correctly the whole
time — it was describing itself wrongly.

The **disk card** quotes that same fitted rate, so it states the same span. One formatter renders
both, so the readout and the card cannot drift apart on how a span is written.

That matters when you read a scary number:

- **A short, sharp dip is not a deadline.** Before 2026-09-12 the rate was the difference between
  the first and last sample in a thirty-minute window. On 2026-09-12 those two happened to land on
  a local peak and a local trough, the readout printed `466.9 GB (26.4 h left, 17.26 GB/h)`, and an
  emergency was opened against a clock that did not exist — across those same six hours the
  volume _gained_ 30 GB.
- **`runway n/a` is an answer, not a gap, and the line now says which one.** It prints
  `runway n/a (not filling)` when the volume is draining, and `runway n/a (too little history)`
  when there is not yet enough recorded history to see past the swing (under thirty minutes of
  samples, which is normal for the first half hour after a restart). It never means the disk is
  fine, and it never means the readout broke — but those two readings are **opposite** (one is
  nothing to do, the other is nothing measured), so the line states which one rather than leaving
  you to guess between them. A rate that was never fitted prints a bare `n/a GB/h` with **no**
  span: a span beside an `n/a` rate would name the basis of a number that does not exist.
- **What a real alarm looks like**: a sustained drop over hours. The one this deliberately keeps is
  a genuine 2026-09-08 event — 312 GB down to 200 GB, monotone, over six hours.

If you need the raw series rather than the verdict, it is on disk at
`<log dir>/perf-status-samples.jsonl`, one row per volume per tick.

### Four states, never two

Every block reports one of:

- **ok** — it has an answer, and it saw a real stretch of time. If it saw less than the window you
  asked for, it says how much it saw.
- **point reading** — it answered, but it saw no stretch of time at all: what it holds is an instant,
  not a measurement over a period. A live snapshot belongs here. So does any number that would
  otherwise be presented as a rate with nothing behind it — the readout refuses to call those `ok`,
  because an instant and a 30-minute window look identical to a reader scanning for good news.
- **unavailable** — it could not read its source right now, and says why.
- **not on this build** — the detector that produces it is not present in this build.

A `point reading` is not a fault: the numbers on it are real, and they are labelled with what they
are. Ask "did this answer?" rather than "is this `ok`?" when you are driving a decision off a block.

This is the design's core rule: **a missing number must never look like a healthy zero.** "No
refills were detected" and "nothing in this build can detect a refill" are different facts, and
collapsing them is how a broken measurement reads as good news for hours. The same rule is why a
**per-session figure now prints the session count it was divided by** — an average over 34 sessions
and the same average over 75 are different claims, and a row that names neither cannot be argued
with.

The readout also states **what it cannot see** — most importantly, a window in which the app was
not running at all. The tapes simply stop, so a dead app and a quiet app look the same here; a
separate watchdog owns that.

### The cards it raises

Each is sustain-gated (it must hold across consecutive checks before a card appears), withdrawn
automatically when the condition clears, and carries a one-click action. **None of them page a
phone** — they are inbox-only, by the owner's explicit decision.

**Each card raises at most once per 24 hours, and an empty inbox is therefore not a verdict.** The
cap is deliberate: a condition that stays true must not re-nag, and a card you dismissed stays
dismissed for the day. One case is exempt — a card the app itself _withdrew_ when the box recovered
does not spend the day's budget, so a genuine second episode raises its own card. When the cap does
hold a raise back, the scorecard's `checks` row says so next to the tripped check
(`[raise suppressed — the day's card for this check was raised at …]`) rather than leaving a
sustained check with no card and no explanation.

The readout also carries the checks themselves, so you can see every bound and how close each one
is without waiting for a card. That list follows the same three-state rule as everything else: until
the first check has run it says **"not evaluated yet"** rather than showing an empty list, because
"nothing is over its limit" and "nothing has been looked at" are different facts.

- The app's main thread stalled 20% or more of the last 30 minutes.
- The harm verdict was on 50% or more of the last 30 minutes.
- A worktree volume has under 2 hours of disk runway left.
- 40% or more of cloud runs in the last hour produced no verdict.
- No branches landed for an hour while 5 or more were waiting — the _dark lander_ case, distinct
  from the auto-lander's own backlog card, which needs the lander to have tried and given up.
- A confirmed external-memory refill.
- Child processes are being started faster than 8 per second (every image combined — there is no
  per-program bound, so this is the whole box's spawn rate, not any one tool's).
- **The BOX was kernel-bound 50% or more of the last 30 minutes** — the whole-computer hang that the
  app-side cards structurally cannot see. Every other card above reads Omniscio's own main thread,
  which is priority-boosted and reads fine while the rest of the machine crawls through the memory
  manager; this one reads the sampler's box-wide kernel-share verdict instead.
- **The desktop window was frozen by long frames for 30 seconds or more in the last hour** — the
  DURATION twin of the stuttering card. That one counts long frames; this one adds up how long they
  blocked the window you work in, so an hour with a few very long freezes shows up here even when
  the frame count looks ordinary. It reads only the main desktop window, not a detached or web
  window, and it stays "unmeasured" until the tracing has covered most of an hour.

Four more watch the **governors themselves** rather than the machine. They are the odd ones out on
this list, and deliberately so: each says that the thing built to PREVENT a slow box has stopped
working, or has stopped being able to tell. None of them is claiming the app is slow right now.

- **The governor stopped recording what it decided.** Every pacing threshold in the app is set by
  replaying that recording, so a gap in it cannot be filled in later.
- **The kernel disagrees with the CPU limits the app thinks it set** — seen twice in a row, so it is
  a real disagreement rather than a check that arrived mid-write. Exactly one part of the app is
  supposed to set those limits, so a disagreement means something else did.
- **The governors were deciding without being able to see you.** They all key on one signal: whether
  you are actually being slowed. While that signal is blind they are not cautious, they are
  uninformed — a window nobody could measure reads exactly like a calm one.
- **A governor asked for something that never took effect.** "Asked for a limit" and "a limit is in
  force" are different facts; this is the card for the gap between them. It reads only policies
  in full `enforce` — which fleet pacing is by default since 2026-09-15.

### Who receives which card

**Customers and developers read different versions of these cards (2026-09-24).** A customer
reported that "Agent work has been hurting the app" sent them to Settings → Performance with
nothing there matching the card — the cards had been written for the people who build Omniscio.

- **Nine cards reach everyone** — the harm verdict, the three main-thread freeze cards (share,
  count, one long freeze), the three window cards (slow clicks, stutter, frozen time), the
  whole-computer card and the waiting-launches card. On an installed copy they say what happened
  with its number, offer one thing to try (turn on Lite mode, or "Lite mode is already on"), and
  say the card clears itself and appears at most once a day. Only the harm card and the three
  freeze cards may say the app "holds back its background agents", and only while agent pacing
  reports full strength — the window cards never do, because that reading does not drive pacing.
- **A copy running from source** gets the same text plus every engineering sentence the card used
  to carry, verbatim, under "Technical details".
- **Their button is "Go to Lite mode"** — Settings → Performance, scrolled to the Lite mode switch
  and highlighted. The two low-memory cards carry it too, and no longer tell anyone to close
  sessions. The below-spec card instead turns Lite mode on in one click ("Turn on Lite mode").
- **Fourteen cards are developer-only** (inbox-alert `I2-gate-developer`): the four governor cards
  above, the governor-window card, the memory-pool refill, the process-start rate, the kernel-share
  pacing lever, the unpaced-storm and storm stand-down cards, the file-read and database holds, the
  garbage-collection pause, and the blind freeze instrument. An installed copy drops them at the
  alert checkpoint; the watch skips them there rather than retrying every tick, withdraws any left
  from before the update on its first tick, and Settings → Alert types leaves them out with a
  one-line note.

### Using it

```
npm run perf:status                   # the plain scorecard, last 30 minutes
npm run perf:status -- --json         # the full report
npm run perf:status -- --window-min 60
```

The command is a thin client over the route: it holds no parser and no thresholds of its own, so it
cannot drift from what the app believes. If the app is not running it says so plainly, rather than
reading a stale file and reporting an old machine as the current one.

**The panel shipped on 2026-09-11** and is no longer hidden behind Settings → Lab. It was gated
while it was a grid of raw counters — faults per second, kernel-bound percentage, event-loop lag —
which is a developer's readout, not an answer. It graduated once it could open with a one-line
plain-language verdict and say what the app is doing about the numbers.

The watching, the alerts and the route were **never** gated on it and still are not — they are the
product's performance ownership and run regardless. Gating them would mean the machine is only
watched while somebody has a toggle switched on, which is the session-shaped ownership this
replaced.

### What the app is doing about it — the governance section

Below the readout, the panel answers the other half of the question: not "how is the machine
doing?" but "what is being done about it?" Every agent-load governor is grouped under the
plain-English mechanism titles the registry itself defines, each group showing how many are on, how
many cannot be observed, and how many you can change. The mechanisms are named in exactly one place
— the registry — and the panel reads those titles rather than restating them, which is why this
page does not list them either. (Restating them is how the taxonomy previously drifted into four
disagreeing copies, one of which said "priority" where the registry says `demotion`, so an agent
grepping the prose found nothing.)

**Where there is a switch and where there deliberately is not.** Admission, pacing, priority,
reclaim and machine strength get a toggle, because turning one off is a sane operator choice.
Correctness fences, starvation floors, off-thread moves and sensors do not, and the panel says
"always on" instead: a fence has to hold with every governor off — that is its definition — and a
floor is the _bound_ on how long background work may be deferred, so removing it is how a
background pass waits forever.

**Three answers it must always be able to give, for the same reason every block above states its
own state:**

- **unknown**, never off — 53 of the 150 governors run in another process or bundle, so this
  process genuinely cannot read their live state;
- **not counted**, never zero — 96 keep no act counter, and "nothing counts this" is a different
  fact from "it has never acted";
- **your settings could not be read** — rather than resolving every setting-backed governor as off
  and printing a confident tally.

Why it exists: of 150 governors a human could reach 24 before this
([the postmortem](../../.claude/memory/postmortems/agent-load-governance-had-no-control-surface-postmortem.md)).

## Related

### Related

- **The judgment half** — the periodic review a threshold cannot do:
  `.claude/memory/performance-watch-checklist.md`.
- **Fleet status** answers a different question: what is out there across the box, rather than how
  it is performing.
- **The load-governance registry** holds every throttle and pacer this readout reports on.

