---
title: Governance control and triage (what is being paced, and why)
---

# `perf:control` and `perf:triage` — the one governance screen and the one triage entry

## What it is

### What they are

Two commands that answer the two different questions people actually arrive with.

| You want to know | Run | It answers |
| --- | --- | --- |
| **What is being paced right now, by which governor, and why?** | `npm run perf:control` | the harm signal and its inputs, the governor's band and mode and who last flipped it, every registered governor with its counters and its kill-switch name, and the top load producers by class |
| **Why is it slow right now?** | `npm run perf:triage` | the existing readouts in a fixed order, ending in ONE ranked verdict: the most likely class, the measurement that ranked it, and the next command to run |

Both are also available as data: `GET /perf/control` serves the control screen's block, and the
control section of the Performance Monitor panel renders the same block. All three come from one
assembler, so a panel reading and a scripted reading cannot disagree.

Contract: `agent-load-governance-contract.md`
— clauses `one-control-surface-and-one-triage-entry`,
`control-covers-the-registry-or-names-what-it-missed`,
`a-flip-from-the-control-surface-names-who-and-why` and
`triage-ends-in-a-ranked-class-and-a-next-command`.

## Where to find it

Three doors onto the same numbers: the **`npm run perf:control`** command, the governance section of the **Performance Monitor** panel, and the local control server's control route. The triage side is the **`npm run perf:triage`** command.

## How it behaves

### Why they exist

Asked on 2026-09-16 whether governance and troubleshooting on this box were centrally organised, the
honest answer was no. There were three overlapping readouts (`perf:status`, `fleet:vitals`,
`overseer:vitals`) plus `fleet-status`, four contracts describing the family, **no single screen that
said WHAT is being paced, BY WHICH governor, and WHY**, and no single entry that ran the readouts in
the right order. The numbers all existed. Nothing composed them.

So these two commands compose. They are the front door; everything else in the family is a
drill-down reached from one of them.

### `perf:control` — the screen

### What is on it

1. **The harm signal and its inputs** — the fused harm verdict with every contributing and every
   blind term, main-thread lag p99 against its SLO bound, caught stalls per hour and the commonest
   `blockedOp` causes behind them, kernel share, per-volume disk runway, and
   **`userInteractiveHeldMs`, which must read 0**.
2. **The governor** — its band, its live mode, its admission coverage, and **who last flipped it and
   when**, read back from the attributed-flip audit (actor kind, opaque actor id, the field that
   moved, from → to, age).
3. **Every registered governor** — one row per registry row, with its live state, its four counts
   (`offered` / `admitted` / `held` / `refused`), its coverage where it measures one, and **its
   kill-switch name**, which is the field you need in order to act.
4. **The top load producers by class** — from the process-start tape the app already keeps, never a
   fresh probe.
5. **One switch** — `--mode observe|enforce --source "<who/why>"`.

### The words that are not numbers, and why they matter more

| Rendered | Means | Never rendered as |
| --- | --- | --- |
| `not-counted` | nothing counts this governor, so **no observation is possible** | `0` acts |
| `unproven` | the counter moved, but the gate has never once held or refused anything | `ok`, `engaged` |
| `unknown` | its live state cannot be observed from this process | `off` |
| `never` | no attributed or observed flip is on record | "never changed" |
| `not-on-this-build` | this build does not serve the block | an empty roster |

`unproven` is the one worth internalising. Measured 2026-09-15 on this box:
`fleet-safety-governor` read `mode enforce · band severe · 3536 of 4396 units granted` while its real
coverage of offered tool calls was **8%**, and `gate-admit` read `IN FORCE — 392 of 392 local
admission decisions gated in 1h (392 granted · 0 refused)`. Both were green readings of a gate that
paced nothing. A counter that moves proves a gate is *reached*; only a hold or a refusal proves it
*acts*. Folding those two together is how a governor stays credited for a year without doing
anything, so they are separate words on the screen.

### `userInteractiveHeldMs` is a pass/fail, not a metric

It must read `0`. The focused session's grants are released at the top of `advance`, above the mode
branch and above the token bucket, so a non-zero value cannot be dispatch jitter — it means a change
stopped releasing the class, which is a `user-never-slowed` violation. The screen says that in those
words rather than printing a number next to every other number.

When the focus seam is blind the screen names `userInteractiveBlindTo` (`killed` /
`settings-unreadable` / `no-session-selected`) instead of printing a `0`. The seam **fails toward
pacing**: an unresolvable focus classifies everything as agent work, so a blind seam means the user
*could* be paced — the opposite of what a reassuring zero would say.

### The switch

```
npm run perf:control -- --mode enforce --source "my-laptop/box hung at 14:20, arming to measure"
```

`--source` is **mandatory**, and the command refuses without it *before it makes the request*.

Authentication is not attribution. The bearer token proves the caller holds a token; it says nothing
about who flipped the governor. On 2026-09-15 at 20:39Z an unidentified bearer re-armed this
governor's `enforce` and stalled every agent for 18 minutes, and the only record was a verbose log
line that rotated away. Lane AD closed that at the route — it now refuses an actor it cannot name
with a `400`. This command refuses one step earlier, as a usage error naming the missing flag, so you
read "name yourself" instead of decoding an HTTP status.

Two deliberate limits:

- **One door.** `--mode` calls `POST /perf/fleet-governor/mode` and nothing else. It writes no
  setting and holds no second copy of the mode vocabulary — a control surface with its own writer is
  how a live mode and a saved setting come to disagree, which this governor already shipped once as a
  one-way latch.
- **Two modes, not four.** `off` and `enforce-maintenance` are session-only overrides that vanish at
  the next restart; offering them here would hand you a setting that silently reverts. The real off
  is `AMC_DISABLE_FLEET_SAFETY_GOVERNOR`, it outranks the dial, and the screen names it on the row.

### `perf:triage` — the ordered runbook

### The six stages, in this order, because the cheap discriminator comes first

1. **`perf:status`** — the harm block, the box block, and the `file-ops` metadata-to-data ratio.
   That ratio *is* the storm class: a box at 100% CPU with ~1M faults/sec reads identically whether
   it is paging or answering a million filesystem questions a second, and those have opposite cures.
2. **The fleet governor** — band, mode, admission coverage, `userInteractiveHeldMs`. A governor in
   `enforce` at `severe` with 8% coverage is a different problem from one that is not armed.
3. **`fleet:vitals`** — the rows sitting off the fleet median. Not a verdict; the row that stopped
   matching its peers.
4. **`gate:honesty`** — the no-verdict rate. A gate that answered nothing is **infrastructure**,
   never a code failure, and reading it as one sends the next hour in the wrong direction.
5. **Git vitals** — ref-read latency, loose and stranded worktree dirs, stale locks.
6. **The harm tape's last caught-stall causes** — `topBlockedOps`, which names what actually parked
   the main thread.

A stage that cannot answer says so and the run continues; it narrows the verdict toward `unmeasured`
rather than aborting.

### The unassigned bucket is not a runner — the one false positive this already produced

Stage 3 asks "do some runners answer while others do not?", because a symptom that is selective cannot
have a whole-box cause. `fleet:vitals` buckets every row carrying no `instance` under the synthetic vm
name `(never reached a VM)` — jobs that died **before any runner claimed them**.

On 2026-09-16 the very first live run counted that bucket as a 28th runner and reported "1 of 28
runners answered NOTHING while 27 answered." That reads as a wedged VM, and it was wrong. Grouping
`cloud-runs.jsonl` by instance over the same 90 minutes: all **27 named runners answered** (20–71 rows
each), and the 28th was **22 unclaimed rows, every one no-verdict** — control-plane-unavailable,
preempted and attempt-abandoned during the 01:30 / 01:50 / 02:10Z generation swaps.

So the 21% no-verdict share was **claim-side**, not a runner fault. Two consequences, both now in the
code:

- the bucket is **excluded** from the runner comparison, so the selective-symptom test cannot fire on
  it, and the note reads "N of M **NAMED** runners";
- it is printed as **its own labelled line** (`unassigned  22 row(s) never reached a runner, 100% of
  them no-verdict — claim-side, not a wedged VM`), because a pile of unclaimed jobs is a real and
  important signal, just a different one.

Why this is worth the paragraph: unfixed, it would have fired on **every** cloud generation swap. A
false positive that recurs on a schedule is the worst kind — it stops looking like noise and starts
looking like a pattern. The fix is pinned by a red-first test built on the measured shape (27 named
runners plus the one bucket), together with a discriminating case proving a **genuinely** silent named
runner is still caught.

### The verdict

A closed enum, so the answer is a routing decision rather than a paragraph:

`main-thread-work` · `kernel-storm-producer` · `git-lock-contention` · `disk-park` ·
`renderer-throttle` · `gate-infra` · `unmeasured`

Every verdict carries the **measurement that ranked it** and the **next command to run**.

Three rules about what it may never say:

- **`unmeasured` outranks a guess.** With the discriminating inputs dark, the honest answer is which
  input was dark and how to restore it. Naming a class it cannot measure is the one thing this
  command must never do.
- **A selective symptom can never be ranked to a global cause.** If some runs are slow and some are
  fast, "the box was loaded" is not a verdict — every whole-box class is refused and the answer falls
  to `unmeasured` with the selectivity it observed. That hand-wave is the commonest way a real
  performance bug here gets closed unfixed.
- **Never a capacity verdict.** There is no "too many sessions" member and there never will be. The
  hardware on this machine is fine; `count-is-never-the-cause`.

### What neither command will do

- **Take its own sample of the machine.** No WMI or CIM, no process enumeration, no filesystem walk.
  `Get-CimInstance Win32_Process` was measured at **1,491 ms for 1,342 processes** — six times slower
  than the interval needed to see a process that lives 200 ms — so on a storming box a CIM poll is
  simultaneously the slowest option and the least accurate one.
  - **The one exception, and why it is not one.** Stages 3 and 4 of `perf:triage` have no HTTP route,
    so they are `fleet:vitals` and `gate:honesty` re-run with `--json`. Those read JSONL **ledgers**;
    they take no sample of the box, so re-running one reuses a guarded reader rather than inventing a
    second opinion. Each gets its own 60-second bound and fails to `unread` — never to a zero.
- **No second reading of the tapes.** The tapes live under a `userData` directory only the app
  resolves, and on a box that has relocated its data directory `%APPDATA%` still holds a
  plausible-but-stale twin of every log. A script reading it would print real-looking numbers from
  hours ago and nobody could tell. The app is the only reader; these commands ask it.
- **No measurement of its own.** The assembler selects and joins blocks that already exist and are
  already guarded. One that derived its own number would become a sixth readout able to disagree with
  the other five, which is the defect this surface exists to remove.

### A request that did not answer is itself a reading

Both commands classify their own failure instead of asserting one:

| Class | What it means |
| --- | --- |
| `no-answer` | the app holds the port and did not answer inside the deadline. **The main thread is stalled** — the loudest possible triage result, not a missing reading |
| `unreachable` | nothing accepted the connection: the app is stopped, and no current reading exists |
| `answered` | the app replied and refused — a token or route problem, demonstrably not a stopped app |
| `usage` | the arguments were refused before any request was made, so this says nothing about the box |

For years the first two printed the same line, which sent operators to restart a running app — banned
by owner rule — and threw away the measurement that was sitting in front of them.

### Turning it off

Registry row `perf-control-surface`, kill switch `AMC_DISABLE_PERF_CONTROL_SURFACE`, read through
`governorSwitch()` at the route. With it set, `GET /perf/control` answers `state: disabled` **naming
the switch**, the panel section renders nothing, and both commands print the disabled line. Never an
empty green.

## Related

### Where to go next

- The machine's own numbers in full → [perf-status.md](perf-status.md)
- Who is doing the work → `npm run perf:whodunit`
- The whole perf front door, symptom routing and the do-not-retry digest →
  `performance.md`
- What the governance family owes →
  `agent-load-governance-contract.md`

