---
title: Test Regime Monitor
---

# Test Regime Monitor

## What it is

A read-only developer-tools panel that shows, at a glance, the health of your whole test setup.
It lives in the sidebar's **Developer Tools** group and has **three tabs**:

- **Overview** (the default) — one verdict for the whole local + cloud test regime, the six
  end-to-end stages (local checks → cloud fleet → cloud runs → landing → master, plus the observers
  watching them), and a worst-first list of what is breaking.
- **System** — the machinery under the tests: the cloud fleet, the dispatchers, and the throttles
  that run every check, plus the full pipeline tables.
- **Branches** — each active branch's merge checks (typecheck / lint / tests).

Ctrl+Tab / Ctrl+Shift+Tab cycles the three tabs from anywhere in the panel.

In-development: hidden until you turn it on (Settings → Lab / the `regime-status` feature toggle).

## Where to find it

It lives in the sidebar's Developer Tools group, alongside the Dev Pipeline panel. It opens on the Overview tab, and the three tabs — Overview, System and Branches — cycle with Ctrl+Tab and Ctrl+Shift+Tab from anywhere in the panel.

It is still in development, so it stays hidden until you turn it on from the Lab section of Settings. Once it is on, the panel is read-only: it reports, and never changes anything.

## How it behaves

### What it shows

### Overview tab — is anything breaking?

The tab opens on **one verdict** in plain words, with a one-line reason:

- **Healthy** — every stage is observed and inside its normal band.
- **Degraded** — at least one warning, nothing critical.
- **Breaking** — at least one critical finding.
- **Inconclusive** — nothing is critical or warning, but at least one stage has no data. The panel
  never infers health from silence: a stage it cannot see keeps the whole verdict from reading
  Healthy.

The rule is "worst stage wins", and it is written on the banner so you never have to remember it.

**The six stages** — each card shows a status chip (a shape plus a colour, so it reads without
colour: ok / warn / bad / unobserved), a headline that carries the raw number, a muted detail
line, and a "where to look next" hint:

| Stage | What it reads | Reads bad / warn when |
| --- | --- | --- |
| **Local checks** | running / failed / passed branch counts, plus the repeat-offender tally | a check has failed 5+ times in a row across branches (bad) or 3+ (warn); a check fails at least half of 20+ recent runs (warn); a box pack-storm, disk saturation, broker shedding or slow gate runs (attached findings) |
| **Cloud fleet** | working runners out of the total, and whether the fleet is serving `--cloud` runs right now | the fleet health probe proves it down (bad); no working runner (bad); fewer than half working (warn) |
| **Cloud runs** | the canonical cloud diagnosis — the same verdict `npm run cloud:diagnose` prints — with its reason code, the failing stage if any, and "N of M recent runs failed · K in flight" | the diagnosis says broken (bad) or congested / degraded (warn); stuck live jobs, a high failure rate at a stage, a fallback storm or slow dispatch (attached findings) |
| **Landing** | the share of ready branches that landed in 24h, stuck branches, real conflicts | land success is low, false merge-failures are elevated, or a branch is at the escalation threshold (attached findings) |
| **Master** | the last master build (outcome + how long ago it was verified) and how many tests are failing in master's own baseline | master does not build (bad); the build net is blocked by infra or has gone quiet past its 24h window (warn); the failing-test count has climbed 10%+ across the baseline trail, or the baseline producer's last run errored / timed out (warn) |
| **Observers** | how many observability tapes are readable, and how old the fleet health probe is | a tape is unreadable, the probe is older than two 15-minute ticks or never ran, the fleet roster is stale, or the cloud verdict names a dimension it cannot observe (warn) |

A stage with no data reads **unobserved**, never ok — off the cloud-operator machine the Fleet and
Master stages are unobserved by design, and the verdict says Inconclusive rather than pretending.

**What's breaking** — every warning and critical finding across all six stages in one list, worst
first, each prefixed with its stage and showing the raw number beside the claim, a plain-English
"Means:" line and a "Check:" pointer. The findings the System tab's Diagnosis already produces
appear here too, attributed to their stage; the Overview adds the gate streaks, the master trend,
the build-net staleness, fleet capacity and the observer gaps that no verdict showed before.

The Cloud runs stage does not re-derive anything: it consumes the canonical cloud diagnosis
verbatim (the one engine behind `cloud:diagnose`, `cloud:chain` and Cloud Health) and only folds
the panel's own threshold findings on top, so the panel and the command line can never disagree
about the cloud. That same engine is now fed the last verdict's **exit code** (1 = code failed,
75 = infrastructure) instead of a bare ok / not-ok, which is why a plain failing test now reads
"degraded · code-failed" rather than the old, wrong "inconclusive — evidence missing".

### System tab — the machinery under the tests

The dense operator view. It opens on the **Diagnosis** block (the canonical cloud verdict with its
reason code and failing stage, then the threshold findings), followed by the cards and tables below.

**Cloud health** (stat cards):

- **Fleet health** — healthy / down / unknown, plus a "serving now" note. A fleet that's actively
  serving `--cloud` runs reads **Healthy** even when the 15-minute probe hasn't landed a clean verdict
  (a busy box) — but a genuinely _down_ fleet still reads Down.
- **Last master build** — the 6-hourly cloud build net's outcome (green / broken / infra / stale) and
  how long ago it last verified. A healthy build (no active alert + a recent verification) reads
  **green**, not "Unknown".
- **In-flight & recent cloud runs**, and **24h cloud spend** against the alert threshold. Each run row
  **expands** to its per-job phase timeline (below).

**Per-job phase timeline** — click any cloud run to expand it and watch that job move through its
stages: **dispatch → submit → wait-for-a-runner → prepare → run → done**, with how long each stage took,
which VM ran it, how long it waited in the queue before a runner picked it up, and — once finished — its
pass/fail counts and cost. It's the in-app version of the `cloud:trace` command. This works on **every
developer's machine**, not just the cloud-operator box: it reads your own machine's local job trace, so
you always see the phases of the jobs you ran.

**Cloud chain verdict** — the timeline above shows you **one job's** path. This block answers the other
question: *which of those stages is failing right now, across every job.* It names the single worst
stage as the **choke point** on one line, then lists every stage with its state and the evidence
behind it — `8/24 passed (33%)`, `2 stuck`, `18 flowed through, 3 in progress` — and the top fail
reasons on anything failing.

It runs the **same classifier** as `npm run cloud:chain`, literally the same function, so the panel
and the command can never name different choke points from the same data. Two rules it inherits from
the command, both deliberate:

- **A stage with no traffic reads "no traffic", never healthy.** Silence is not health — a stage
  nothing entered has proved nothing.
- **If the job trace cannot be read, it says so** instead of showing eight stages claiming nothing
  ran. "I cannot see" and "nothing happened" are different facts.

It does not print a fixed time window, because it does not have one: the panel reads the most recent
5,000 trace rows, and the funnel footer just below states the real span those cover.

You may still prefer the command in a terminal — it adds three things this block deliberately does
not duplicate, because the panel already shows them elsewhere: whether finished results are getting
back, the fleet snapshot, and the last master build (the Cloud health cards above, and the Overview
tab's Cloud fleet and Master stages).

**Fleet** — how many runners exist and how many actually work: total with the Windows/Linux split,
healthy vs broken / disk-full, live SSH tunnels, recent reliability strikes, and how fresh the roster is.

**Dispatch & queue** — how fast submitted jobs get picked up by a runner (median + worst recent
submit→claim latency).

**Gates & throttles** — the limiting factors: whether cloud-enforce is routing work onto the fleet, how
many concurrency slots are in use right now, and how long jobs wait for a slot.

Off the cloud-operator machine these degrade gracefully — each group shows an "unavailable" note
instead of looking broken.

### System tab — the full-pipeline view (everything, deliberately dense)

Below the cards, the System tab surfaces the ENTIRE pipeline from the observability tapes the
system already records — nothing hidden, every number backend-computed:

- **Diagnosis** (top of the tab) — automatic ok / warn / bad verdicts with the raw number shown
  beside every claim, a plain-English "Means:" line, and a "Check:" pointer for what to look at
  next; unreadable data sources surface here too (the panel tells you when it's flying blind).
- **Cloud job funnel** — all 8 stages a cloud job passes through (dispatch → submit → wake →
  await-claim → prepare → env → run → complete): live jobs sorted most-stuck-first with
  time-in-stage, per-stage pass/fail/stall counts, top failure reasons, and dwell-time
  percentiles (p50 / p95 / max, always with sample counts).
- **Dispatch & queue timing** — how long work takes to reach the dispatchers: submit→claim wait
  percentiles per job kind, how many jobs sit unclaimed right now (and the longest such wait),
  plus slot/lane waits per outcome (acquired, and the give-ups: failed-open / congestion-local /
  population-bound / lane-over-share / lane-parked-out — the set is `WAIT_OUTCOMES` in
  `scripts/cloud/wait-metrics.mjs`; a value outside it is quarantined and does not vote).
- **Local gates, admissions & floors** — every registered chokepoint on the box (heavy-job
  broker, check/build/install/e2e slots, memory/CPU floors, route + dispatch deciders): ok /
  hold / timeout / fail / skip / heal counts, top reasons, wait percentiles, and last-event age.
  A silent system renders an explicit "no data yet" row — it never just disappears.
- **Land pipeline & box health** — land success % (24h), false merge-failures per hour, real
  conflicts, stuck branches (with the top offenders listed), gate p95, fallback counts (1h/24h)
  with the recent fallback events, and the current disk-pressure / pack-storm state.
- **Live event feed** (bottom) — the merged tail of both trace tapes (local gate spine + cloud
  job spine), newest first: every recorded transition with its source, event, reason, wait, and
  branch/jobId.

All of it is read-only over existing ledger files (bounded tail reads, ~10s memo) — the panel
can never slow the pipeline it watches.

### Branches tab — merge gate per branch

For every active worktree branch, the three merge checks — **typecheck**, **lint**, **tests** — each
render as running / passed / failed / **stale** (passed on an _older commit_) / not run, with a one-word
rollup, grouped by repo (running + failed first).

Where the data comes from: the gate already writes a per-worktree `landing-proof.json` receipt (the
record the auto-lander reads to trust a `ready-to-merge` tag). The panel reads that for pass/fail, and a
tiny "running" marker the check writes at start (and deletes at exit) for the live state. **Nothing new
is stored in a database.**

**Known limitation:** the gate view reflects the checks agents actually iterate with (the `:agent`
variants + the guard lane), not a hand-run stock `npm run typecheck`/`lint`. The "stale" flag covers the
edge where a final hand-run check ran against a different commit.

### How fresh is it

- The **Branches** tab updates within a few seconds while the panel is open (it polls on its own
  interval, and costs nothing when closed).
- The **System** tab's heavy fleet/build cards reflect the latest 15-minute / ~6-hour snapshot,
  memoized ~60 seconds so the panel's fast poll never re-reads the big ledgers. The **live bits** —
  in-flight/recent runs and the per-job phase timelines — refresh within a few seconds, so a job's
  stages advance as you watch. The header shows an honest "updated N ago".

## For agents

### Controlling it via the CLI

It's a read-only panel; there's no state to mutate. The one read is the `regime-status:get` IPC
channel (also reachable over the mobile/web bridge for an at-a-glance check on your phone).
Agents get the SAME full snapshot over the CLI control server: `GET /gate/regime-status`
(bearer-gated, read-budgeted) returns all four halves — gate, cloud, system, and the
full-pipeline surface (diagnosis · funnel · timing · spine · land/feed) — plus the `overview`
field the Overview tab renders (`verdict`, `verdictLine`, the six `stages`, the `breaking` list
and its `counts`), mirroring the IPC payload field-for-field. An agent asking "is anything
breaking?" reads `overview.verdict` and `overview.breaking` and is done; `overview` is `null`
only if the fold itself faulted (the rest of the snapshot still arrives).

## Related

- The gate itself: [../../.claude/memory/contracts/checks-gate-contract.md](../../.claude/memory/contracts/checks-gate-contract.md)
  (the running marker is documented there).
- The panel contract: [../../.claude/memory/contracts/regime-status-panel-contract.md](../../.claude/memory/contracts/regime-status-panel-contract.md).
- The pipeline surface's own invariants: [../../.claude/memory/contracts/pipeline-visibility-surface-contract.md](../../.claude/memory/contracts/pipeline-visibility-surface-contract.md).
- The Dev Pipeline panel (live pipeline runs + the auto-lander): a sibling Developer-Tools panel.
