---
title: Verdict panel
---

# Verdict panel

## What it is

> **What it is for.** One screen that tells you whether the verdict road (the system that runs an agent's test/typecheck/lint check and hands the answer back) is healthy right now, and if it is not, exactly where the line is backed up — without opening a ledger or running a command-line readout.

## Where to find it

Developer Tools → **Verdict** in the left sidebar. It is an in-development feature: turn it on in Settings (search _Verdict panel_) or launch the app with `AMC_SHOW_VERDICT_PANEL=1`. It works on desktop and on the phone web view; the two "reveal this file" buttons are desktop-only.

## How it behaves

### How to read it, top to bottom

**1. The sentence.** The first line answers the question: _"The line is flowing."_ or _"The line is backed up at Dispatch — 14 waiting, oldest 11 min."_ If the broker is not running on this machine it says so instead of guessing. A quiet grey line under it names any source that could not be read this tick, so "flowing" is never claimed over a dark instrument. Click the sentence to open the **machinery**: the broker's process id, generation and keeper heartbeat, plus a timeline of recent incidents (broker installs and restarts, fleet deploys, app restarts) and the broker's own cards and readout.

**2. Alarms.** Shown only when something tripped: a headline stat outside its band, a marked station, a runner far off the fleet median, answers that went back to pending after being delivered, worktrees the ledger does not know, stalled deliveries, or the wake gate refusing to start a runner while jobs wait. Each alarm names its source and as-of time. When nothing tripped, nothing is shown.

**3. Six headline stats.** Each tile shows a value, the normal band it is judged against, and a colour stripe that follows the verdict (green inside the band, amber near its edge, red outside, grey for no reading):

| Stat                   | What it measures                                                                             | Normal                                    |
| ---------------------- | -------------------------------------------------------------------------------------------- | ----------------------------------------- |
| Answered · 3 h         | Share of gate checks that got a real answer, by gate-honesty's own judged/answered semantics | at least 80% (gate-honesty's floor)       |
| Time to answer · p90   | The slow ninetieth percentile from ask to answer over the last 3 h                           | within 1.5× the last 24 h                 |
| In the line            | Jobs anywhere in the eight stations right now, with the oldest one's age                     | within 1.5× the last 24 h                 |
| Fleet healthy          | Runners current and not benched, out of runners awake; plus month spend on pace for the cap  | every awake runner healthy; spend on pace |
| This box · user-harm   | Share of the last 30 min the user-harm verdict was on                                        | under 10%                                 |
| Landing · oldest ready | How long the longest-waiting ready-to-merge branch has waited                                | under 30 min                              |

Where a rule already sets the bar (the 80% floor, the spend cap) the band is that rule. Everywhere else the band is 1.5× the last 24 hours of the same reading, and the tile says how many hours of history actually backed it. Click a stat to open its detail.

**4. The line.** The road as eight stations, left to right: **Ask → Seal → Dispatch → Run → Verdict → Deliver → Tag → Land**. Each station shows how many jobs are waiting in front of it, how long the oldest has waited, and a thin bar for how many it finished in the last hour. Two bottleneck marks can appear, and they are different things: **biggest pile** (the station with the most waiting) and **slowest vs normal** (the station whose oldest wait is furthest above its own 24-hour normal). Both are marked when they differ; neither appears while every wait is within normal or fewer than three jobs sit in the line. Click a station to open its detail.

**5. One click deeper.** A detail opens below the line and closes with the × or by clicking the same thing again:

- **Ask** — accepted, deferred and refused asks (by reason) in the last hour, the backlog and the oldest parked ask.
- **Seal** — seal and pack timings against their baseline, the seal width in force, torn seals.
- **Dispatch** — the local fast lane (slots, in flight, declines by reason), cloud queue depth, and the wake gate's verdict per runner.
- **Run** — checks per hour by kind (vitest, tsc, eslint, e2e) and by route, durations p50/p90/max, executions right now, runner outliers.
- **Verdict** — answered rate over 1 h / 3 h / 24 h against the floor, non-answers by their real names (`refused:unauthorized-repo`, `no-evidence`, `attempt-abandoned`, `no-verdict-line`, `infra-permanent`, …), and the split "road failed / caller withdrew / policy refused".
- **Deliver** — pending, handed to host, stalled, delivered, undeliverable; re-pends after delivered (must read 0); scope-mismatch adoptions; and answers waiting on sessions you paused.
- **Tag** and **Land** — ready branches with their wait; landed in the last hour, the auto-lander's state, worktrees on disk versus the ledger.
- **Fleet** — the runner table (version, busy/idle, jobs, INF% and NOVER% against the fleet median with `!` on outliers, p50, lanes), deploy state, month spend against the cap.
- **This box** — CPU, kernel share, faults, the user-harm and busy-gate shares, git wedge signals, the session census, and agent 429s in the last 30 minutes by session name.

**6. Two clicks.** Every detail ends with a **Raw rows** disclosure: the last rows of the ledger behind it, one line each.

### The three states every number can be in

- **Measured** — the value, with its source and as-of time underneath ("gate-honesty · as of 02:41:05Z").
- **Stale** — the value plus an amber _stale_ pill: the source has not written for longer than its normal cadence, so the number is real but old.
- **No reading** — grey italic "no reading" with the reason in the tooltip (a missing file, an unreadable ledger, a readout not present on this build). It is never shown as zero and never coloured healthy.

### Answers waiting on paused sessions

On this app a paused session is always a deliberate pause by you. An answer waiting to be delivered to such a session is shown quietly on the Deliver station as _"6 waiting on 3 sessions you paused"_, is excluded from every red count and from the bottleneck marks, and is cleared by unpausing the session. When the installed broker cannot yet split those answers out, the Deliver detail says so in amber rather than counting them as stalled.

### What it does not do

It is read-only: **Refresh**, **Copy readout** (the sentence, the six stats and the line as plain text), and the two reveal-a-file buttons are its only levers. Restart, drain and rollback stay the owner's terminal commands (`npm run verdict:rollback`, the keeper task). It never starts a cloud machine, never spawns a process, and reads only bounded tails of the ledgers while the panel is open.

## For agents

The same snapshot the panel renders is available over the local control server: `GET http://127.0.0.1:19519/verdict/observatory` (bearer token; `?refresh=1` forces a fresh read). It returns the sentence, the headline stats, the line, the alarms and every detail as JSON, so an agent reads exactly what the operator sees.

`?refresh=1` re-reads every source from scratch — the tickets directory, the metrics tail, the three gate ledgers, the fleet ledgers, the telemetry day files and the repo's git reads — so it carries its own **per-source budget of 30 refreshes a minute** on top of the ordinary read budget; past it the route answers `429` with a `Retry-After`. The unbudgeted read (`GET /verdict/observatory` with no query) is the cached snapshot and costs nothing, so poll that and refresh only when you actually need a fresh number.

### Under the hood (for readers with the repo)

- Contract: `.claude/memory/contracts/verdict-panel-contract.md` (twenty named obligations). Map: `.claude/memory/verdict-panel-map.md`.
- Main process: `src/main/services/verdict/observatory/` — one reader per source (the broker tick ledger, the telemetry day files, gate attempts via gate-honesty, the fleet-vitals ledgers, the owner policy and rollback watch, the performance readout and sessions table, the worktree ledger and auto-lander), a line builder, a band judge, a sentence builder and an alarm deriver, folded by one composer into the existing Verdict panel feed. The two script-side readouts run in a worker thread, never on the main thread.
- Renderer: `src/renderer/src/features/verdict/observatory/` — the sentence, alarms, headline row, the line and station boxes, the drill-down host and the eleven detail views. The renderer computes nothing.

## Related

The panel watches the whole gate road; [Job Monitor](job-monitor.md) is where you kill one wedged check, and [Regime Status](regime-status.md) covers the test regime that road is running. [Stats](stats.md) holds the spend and usage numbers the panel only summarises, and [Logs and Debugging](logs-and-debugging.md) is where the raw ledgers behind every station can be read in full.
