Omniscio documentation
Browse all documentation
  1. Getting Started13
  2. Sessions & Agents115
  3. Inbox & Notifications59
  4. Projects & Tasks95
  5. Automation & Scheduling75
  6. Knowledge & Memory26
  7. AI Features60
  8. Integrations100
  9. Plugins & Marketplace33
  10. Cloud & Teams56
  11. Settings & Customization58
  12. Account & Billing28
  13. Troubleshooting84
  14. CLI & API Reference22
  15. Legal & Policies4
  16. Uncategorised22

Governance control and triage (what is being paced, and why)

The one screen that says what the app is currently holding back and which governor decided it, and the one ordered runbook for working out why the machine is slow — both reading the app's own numbers rather than taking their own sample of the computer. They exist because three overlapping readouts used to answer the question differently.

What it is

What they are

Two commands that answer the two different questions people actually arrive with.

You want to know Run It answers
What is being paced right now, by which governor, and why? npm run perf:control the harm signal and its inputs, the governor's band and mode and who last flipped it, every registered governor with its counters and its kill-switch name, and the top load producers by class
Why is it slow right now? npm run perf:triage the existing readouts in a fixed order, ending in ONE ranked verdict: the most likely class, the measurement that ranked it, and the next command to run

Both are also available as data: GET /perf/control serves the control screen's block, and the control section of the Performance Monitor panel renders the same block. All three come from one assembler, so a panel reading and a scripted reading cannot disagree.

Contract: agent-load-governance-contract.md — clauses one-control-surface-and-one-triage-entry, control-covers-the-registry-or-names-what-it-missed, a-flip-from-the-control-surface-names-who-and-why and triage-ends-in-a-ranked-class-and-a-next-command.

Where to find it

Three doors onto the same numbers: the npm run perf:control command, the governance section of the Performance Monitor panel, and the local control server's control route. The triage side is the npm run perf:triage command.

How it behaves

Why they exist

Asked on 2026-09-16 whether governance and troubleshooting on this box were centrally organised, the honest answer was no. There were three overlapping readouts (perf:status, fleet:vitals, overseer:vitals) plus fleet-status, four contracts describing the family, no single screen that said WHAT is being paced, BY WHICH governor, and WHY, and no single entry that ran the readouts in the right order. The numbers all existed. Nothing composed them.

So these two commands compose. They are the front door; everything else in the family is a drill-down reached from one of them.

perf:control — the screen

What is on it

  1. The harm signal and its inputs — the fused harm verdict with every contributing and every blind term, main-thread lag p99 against its SLO bound, caught stalls per hour and the commonest blockedOp causes behind them, kernel share, per-volume disk runway, and userInteractiveHeldMs, which must read 0.
  2. The governor — its band, its live mode, its admission coverage, and who last flipped it and when, read back from the attributed-flip audit (actor kind, opaque actor id, the field that moved, from → to, age).
  3. Every registered governor — one row per registry row, with its live state, its four counts (offered / admitted / held / refused), its coverage where it measures one, and its kill-switch name, which is the field you need in order to act.
  4. The top load producers by class — from the process-start tape the app already keeps, never a fresh probe.
  5. One switch — --mode observe|enforce --source "<who/why>".

The words that are not numbers, and why they matter more

Rendered Means Never rendered as
not-counted nothing counts this governor, so no observation is possible 0 acts
unproven the counter moved, but the gate has never once held or refused anything ok, engaged
unknown its live state cannot be observed from this process off
never no attributed or observed flip is on record "never changed"
not-on-this-build this build does not serve the block an empty roster

unproven is the one worth internalising. Measured 2026-09-15 on this box: fleet-safety-governor read mode enforce · band severe · 3536 of 4396 units granted while its real coverage of offered tool calls was 8%, and gate-admit read IN FORCE — 392 of 392 local admission decisions gated in 1h (392 granted · 0 refused). Both were green readings of a gate that paced nothing. A counter that moves proves a gate is reached; only a hold or a refusal proves it acts. Folding those two together is how a governor stays credited for a year without doing anything, so they are separate words on the screen.

userInteractiveHeldMs is a pass/fail, not a metric

It must read 0. The focused session's grants are released at the top of advance, above the mode branch and above the token bucket, so a non-zero value cannot be dispatch jitter — it means a change stopped releasing the class, which is a user-never-slowed violation. The screen says that in those words rather than printing a number next to every other number.

When the focus seam is blind the screen names userInteractiveBlindTo (killed / settings-unreadable / no-session-selected) instead of printing a 0. The seam fails toward pacing: an unresolvable focus classifies everything as agent work, so a blind seam means the user could be paced — the opposite of what a reassuring zero would say.

The switch

npm run perf:control -- --mode enforce --source "my-laptop/box hung at 14:20, arming to measure"

--source is mandatory, and the command refuses without it before it makes the request.

Authentication is not attribution. The bearer token proves the caller holds a token; it says nothing about who flipped the governor. On 2026-09-15 at 20:39Z an unidentified bearer re-armed this governor's enforce and stalled every agent for 18 minutes, and the only record was a verbose log line that rotated away. Lane AD closed that at the route — it now refuses an actor it cannot name with a 400. This command refuses one step earlier, as a usage error naming the missing flag, so you read "name yourself" instead of decoding an HTTP status.

Two deliberate limits:

  • One door. --mode calls POST /perf/fleet-governor/mode and nothing else. It writes no setting and holds no second copy of the mode vocabulary — a control surface with its own writer is how a live mode and a saved setting come to disagree, which this governor already shipped once as a one-way latch.
  • Two modes, not four. off and enforce-maintenance are session-only overrides that vanish at the next restart; offering them here would hand you a setting that silently reverts. The real off is AMC_DISABLE_FLEET_SAFETY_GOVERNOR, it outranks the dial, and the screen names it on the row.

perf:triage — the ordered runbook

The six stages, in this order, because the cheap discriminator comes first

  1. perf:status — the harm block, the box block, and the file-ops metadata-to-data ratio. That ratio is the storm class: a box at 100% CPU with ~1M faults/sec reads identically whether it is paging or answering a million filesystem questions a second, and those have opposite cures.
  2. The fleet governor — band, mode, admission coverage, userInteractiveHeldMs. A governor in enforce at severe with 8% coverage is a different problem from one that is not armed.
  3. fleet:vitals — the rows sitting off the fleet median. Not a verdict; the row that stopped matching its peers.
  4. gate:honesty — the no-verdict rate. A gate that answered nothing is infrastructure, never a code failure, and reading it as one sends the next hour in the wrong direction.
  5. Git vitals — ref-read latency, loose and stranded worktree dirs, stale locks.
  6. The harm tape's last caught-stall causes — topBlockedOps, which names what actually parked the main thread.

A stage that cannot answer says so and the run continues; it narrows the verdict toward unmeasured rather than aborting.

The unassigned bucket is not a runner — the one false positive this already produced

Stage 3 asks "do some runners answer while others do not?", because a symptom that is selective cannot have a whole-box cause. fleet:vitals buckets every row carrying no instance under the synthetic vm name (never reached a VM) — jobs that died before any runner claimed them.

On 2026-09-16 the very first live run counted that bucket as a 28th runner and reported "1 of 28 runners answered NOTHING while 27 answered." That reads as a wedged VM, and it was wrong. Grouping cloud-runs.jsonl by instance over the same 90 minutes: all 27 named runners answered (20–71 rows each), and the 28th was 22 unclaimed rows, every one no-verdict — control-plane-unavailable, preempted and attempt-abandoned during the 01:30 / 01:50 / 02:10Z generation swaps.

So the 21% no-verdict share was claim-side, not a runner fault. Two consequences, both now in the code:

  • the bucket is excluded from the runner comparison, so the selective-symptom test cannot fire on it, and the note reads "N of M NAMED runners";
  • it is printed as its own labelled line (unassigned 22 row(s) never reached a runner, 100% of them no-verdict — claim-side, not a wedged VM), because a pile of unclaimed jobs is a real and important signal, just a different one.

Why this is worth the paragraph: unfixed, it would have fired on every cloud generation swap. A false positive that recurs on a schedule is the worst kind — it stops looking like noise and starts looking like a pattern. The fix is pinned by a red-first test built on the measured shape (27 named runners plus the one bucket), together with a discriminating case proving a genuinely silent named runner is still caught.

The verdict

A closed enum, so the answer is a routing decision rather than a paragraph:

main-thread-work · kernel-storm-producer · git-lock-contention · disk-park · renderer-throttle · gate-infra · unmeasured

Every verdict carries the measurement that ranked it and the next command to run.

Three rules about what it may never say:

  • unmeasured outranks a guess. With the discriminating inputs dark, the honest answer is which input was dark and how to restore it. Naming a class it cannot measure is the one thing this command must never do.
  • A selective symptom can never be ranked to a global cause. If some runs are slow and some are fast, "the box was loaded" is not a verdict — every whole-box class is refused and the answer falls to unmeasured with the selectivity it observed. That hand-wave is the commonest way a real performance bug here gets closed unfixed.
  • Never a capacity verdict. There is no "too many sessions" member and there never will be. The hardware on this machine is fine; count-is-never-the-cause.

What neither command will do

  • Take its own sample of the machine. No WMI or CIM, no process enumeration, no filesystem walk. Get-CimInstance Win32_Process was measured at 1,491 ms for 1,342 processes — six times slower than the interval needed to see a process that lives 200 ms — so on a storming box a CIM poll is simultaneously the slowest option and the least accurate one.
    • The one exception, and why it is not one. Stages 3 and 4 of perf:triage have no HTTP route, so they are fleet:vitals and gate:honesty re-run with --json. Those read JSONL ledgers; they take no sample of the box, so re-running one reuses a guarded reader rather than inventing a second opinion. Each gets its own 60-second bound and fails to unread — never to a zero.
  • No second reading of the tapes. The tapes live under a userData directory only the app resolves, and on a box that has relocated its data directory %APPDATA% still holds a plausible-but-stale twin of every log. A script reading it would print real-looking numbers from hours ago and nobody could tell. The app is the only reader; these commands ask it.
  • No measurement of its own. The assembler selects and joins blocks that already exist and are already guarded. One that derived its own number would become a sixth readout able to disagree with the other five, which is the defect this surface exists to remove.

A request that did not answer is itself a reading

Both commands classify their own failure instead of asserting one:

Class What it means
no-answer the app holds the port and did not answer inside the deadline. The main thread is stalled — the loudest possible triage result, not a missing reading
unreachable nothing accepted the connection: the app is stopped, and no current reading exists
answered the app replied and refused — a token or route problem, demonstrably not a stopped app
usage the arguments were refused before any request was made, so this says nothing about the box

For years the first two printed the same line, which sent operators to restart a running app — banned by owner rule — and threw away the measurement that was sitting in front of them.

Turning it off

Registry row perf-control-surface, kill switch AMC_DISABLE_PERF_CONTROL_SURFACE, read through governorSwitch() at the route. With it set, GET /perf/control answers state: disabled naming the switch, the panel section renders nothing, and both commands print the disabled line. Never an empty green.

Related

Where to go next

  • The machine's own numbers in full → perf-status.md
  • Who is doing the work → npm run perf:whodunit
  • The whole perf front door, symptom routing and the do-not-retry digest → performance.md
  • What the governance family owes → agent-load-governance-contract.md

Last verified 2026-10-06