Work output by load (what the fleet actually finished)
The readout that answers "what did the fleet actually FINISH at this load?" — commits landed, ready-to-merge tags applied, and items brought to the owner's inbox, each per hour for the whole fleet and per session-hour. Tokens are a supporting column, because they measure how much was SAID, not what was DONE. Beneath it sit two more axes: a SESSION LIFECYCLE section, and a SESSION TIME section naming where every second of a session's clock went, with the leftover reported as UNACCOUNTED.
What it is
A per-load reading of finished work, in three countable measures, next to the tokens the same window produced. It exists because tokens measure how much was SAID: a fleet that talks more while finishing less looks identical to a productive one when the only column is output tokens.
The question it answers is narrow and specific: as the live-session count rises, does the fleet finish less per session-hour — and does it still finish more per hour overall?
Where to find it
npm run perf:work-output— the full report on the terminal, including the plain-English answer at the top.npm run perf:work-output -- --share— the same report published to the owner's Omniscio Shares.- The Performance Monitor panel renders the same fold, so a bucket reads identically on both.
How it behaves
The three measures
Each is a record the app already wrote — nothing is inferred from elapsed time or activity:
- Ready-to-merge tags applied — the mint's own record in
worktree_events(ops-ready,ops-ready-user-override). A tag minted by hand, or on another machine, is absent. This leads, because a tag is stamped the moment an agent FINISHES the branch. - Commits landed — this app's own lander record:
auto_lander_eventsin the live database, read windowed and read-only,outcome = 'landed'. NOT theGET /auto-lander/eventsactivity route: it serves the newest 500 rows — a display limit for the panel's log — while the table keeps 72 h, so a busy window came back short under a coverage line that readok. - Items brought to the owner's inbox —
inbox_alert_items, counted at the moment the card REACHED the inbox. A card still held back is not yet a delivery.
Lands are not a load's output. The lander works a queue, so a bucket's lands include branches finished at an earlier load — at 0-25 sessions one night it was draining earlier tags, 21.84 lands/h against only 4.37 tags/h. Both surfaces print one line saying so and make NO causal claim in either direction: read on tags, the same rows rose 7x with load.
Reading the two tables
The report prints the same rows twice, on purpose:
- Counts, and the whole fleet per hour — how much the fleet finished in an hour, whatever its size.
- Counts per session-hour — how much each individual session finished in an hour.
A per-hour figure that holds while the per-session-hour figure falls means the busier load bought session-hours rather than finished work. Where the per-hour figure itself falls, the busier load simply finished less.
Both are read inside matched hours where possible: the load comparison in the plain-English answer at the top only ever compares two live-session buckets observed in the SAME clock hour, so the box's own drift between hours cannot be mistaken for a load effect. The pooled table below it is labelled as confounded and is never the answer.
When a number is missing
- A bucket with under ten covered minutes of tape prints its counts and
thininstead of a rate, because a rate off a few minutes is noise. - A source that cannot be read, or whose history starts inside the window, reads
partialorunmeasuredwith its reason — never a low number that looks like a slow fleet. - An event whose live-session count could not be read is counted on its own
unknownrow, never folded into the quietest bucket, so the parts still sum to the whole. - No session-count recommendation appears anywhere. A peak is a reading of the data, not a number of sessions to run.
The stable link
The share page is refreshed hourly to the same URL by an in-app scheduled job
(AMC-work-output-share), which passes the stored share token back as updateToken. The address
never changes, so it is safe to bookmark: a refresh replaces the page in place rather than minting a
new link each run.
That job is never stood down when the box is busy, which is not the same thing as declining the
load hold: a busy box is exactly the reading that is wanted, and the box is only ever at 100+ live
sessions during the very conditions that put a scheduled job on hold. It declares liveness, the
one declaration that exempts a row from every term (measured 2026-10-02: the row was being deferred
wouldDefer: true while the page's last publish was over an hour old).
The hourly card in your inbox
Each successful hourly run also raises one inbox card, so the key stats arrive without opening anything:
- One card, not a pile — every raise lands on the same card, replacing the hour before it. Dismissing it does not silence the next hour, because a standing reading whose numbers have moved is not a repeat of the one you put away.
- Two windows, each labelled — the last hour and the last 24 hours, with branches landed (one land is one branch and its merge commit, so the branch and commit figures are the same number), ready-to-merge tags, inbox deliveries, finished agent turns, active session-hours, and the typical tool wait.
- Every measure carries its own coverage — a source that could not be read, or that starts inside the window, is said so on the card rather than printed as a low number, and a 24-hour window the tape does not cover says how much of it the figures cover.
- A failed publish says so and carries no numbers — the alternative would be the previous hour's figures presented as this hour's. The card names the last successful page as the last successful one.
- It has an off switch — the card is an alert type, so it can be muted like any other card in Settings, and it is ON by default.
The page states this only as strongly as the app can support it: it is published by hand until the
branch that ships the job has landed, so the script probes GET /ops/jobs at publish time and says
"once the job is registered" until the row is really there.
If the link is ever revoked, or its 30-day default lapses, the server answers 400 No such active share to update before publishing anything — an update deliberately keeps the target's own expiry.
The job then forgets the dead token and republishes without it, so the next run mints a fresh link
and the hourly refresh recovers on its own instead of failing forever. That means the address CAN
change after a revocation; the state file at ~/.amc/work-output-share.json always names the current
one.
Session lifecycle — how long a session lives, and how long its branch takes to land
The readout carries a second axis beneath the load tables: what happens BETWEEN a session starting and its work reaching master. Three intervals, each as p25 / p50 / p75 / p90 minutes:
- Start → first branch landed — how long a session's own branch takes to be merged.
- Start → session ended — split into sessions that landed a branch and sessions that did not.
- Land → session ended — the tail after a merge.
Each is shown over the whole window, by day, and by the owner's local clock hour of the
session's start. Reviewer sessions — a branch beginning review-snapshot/ — are on their own line
and out of every headline number, because a review snapshot ends as soon as it answers.
Three things about it worth knowing before you read a number:
land → endedis signed. It counts only sessions whose branch landed BEFORE they ended. A branch the landing queue finishes afterwards is still real work and stays in the lander count, but it is counted on its own line — around half of lands happen after the session that wrote them, so including them would make the median meaningless rather than merely low.- Every excluded group is named with its size — still running, no branch, landed after the end, and the reviewers — and the lander and non-lander rows sum back to the section's own count.
- The clock is yours, not UTC. The stored timestamps are UTC; grouping by them would file an evening session under the next day, so both groupings use the local clock.
Reader: scripts/perf/session-lifecycle.mjs. It reads
sessions, worktrees and auto_lander_events once each inside one window and joins them in
memory; the land source is the same auto-lander event the rest of this readout reads, so the two can
never disagree.
Session time — where every second of one session actually went
The lifecycle section says how LONG a session lives. This one says where that time WENT. One session's clock, from its start to its end (or to now while it runs), is partitioned into named phases, every second claimed by exactly one:
app-down · tools · model · turn-rest · gate · mint · queue · parked · idle · unaccounted
npm run perf:session-time -- <session-id>prints the phase table and a timeline of the intervals.- The share page carries the fleet's phase share of total session-seconds, per local clock hour.
Why it exists, measured on the session that built it: of a 25.2 h span, over half the clock went to the two phases nothing on the box had ever measured —
| phase | claimed | share |
|---|---|---|
| idle, by what ended each gap | 453.6 min | 30.0 % |
| gate / test wait | 361.5 min | 23.9 % |
| tools | 223.3 min | 14.8 % |
model (apiMs) |
58.2 min | 3.9 % |
| mint wait | 15.2 min | 1.0 % |
| turn-rest | 2.1 min | 0.1 % |
| queue · parked · app-down | 0.0 min | 0.0 % |
| unaccounted | 396.5 min | 26.3 % |
An agent that looks busy for 25 hours is idle a third of it and waiting on a gate a quarter of it. The named phases now cover 73.7 % of the clock; the remainder is dominated by the single largest unknown in the instrument — the tape's own rotation, which never saw this session's first 9.95 h. The gate, mint and queue ledgers reach back further than the tape, which is why coverage beats what the tape alone could name.
Five things to know before reading a number:
- The phases sum to the span exactly, and that sum is a DISPLAY check, never the proof. It is true by construction on every input, including a broken one — so the reader is pinned by a fixture whose every phase number is hand-computed, not by the sum.
- First claim wins, and the order is published. A tool call made while a gate is out charges to
tools, so thegatefigure is a residual — "gate time not otherwise occupied" — and not the total gate cost. Both readings are printed side by side:rawis what the phase was offered andclaimedis what it kept, so a gate a tool call shadowed still shows its full wait inrawand readsclaimed 0.0beside it. - A phase with no record is not a zero.
0.0 measuredmeans this session waited on no gates; a ledger that could not be read saysno record, a source that is genuinely dark saysnot measurable, and a tape that starts mid-session saystruncated. None of the four ever share a rendering. - The unaccounted remainder is its own named row — it is never folded into idle, so a gap is visible rather than absorbed.
- A green test suite is not proof the reader works. The unit spec feeds the pure fold a synthetic fixture, so it cannot see the read path at all. Every claim about the reader is positive-controlled by running it on a session whose real numbers are already known — see the postmortem for the four defects a green suite hid.
Reader: scripts/perf/session-time.mjs. It reuses the tape,
gap-ender and app-down readers work-output-report.mjs already had — one parser per source, so the
two surfaces cannot disagree — and adds readers only for the three ledgers nothing read before (the
gate-attempt ledger, the mint timings, the cloud slot waits), plus the read-only worktrees table
every one of those joins through.
For agents
Where the data lives
- Fold (pure, no I/O):
scripts/perf/work-rate.mjs, shared with the app's Performance Monitor block. - Reader and rendering:
scripts/perf/work-output-report.mjs. - Session time (pure fold + phase partition + the three new ledger readers):
scripts/perf/session-time.mjs. - Buckets and the time join:
scripts/perf/load-buckets.mjs. - All three measures: the live
mission-control.db, read-only, short windowed queries that each ride an index (worktree_events,auto_lander_events,inbox_alert_items). - The hourly publish job:
scripts/ops/ops-tasks/amc-work-output-share.mjs; its state (share token and last publish) lives at~/.amc/work-output-share.json. - Contract: work-rate-by-load-readout-contract.md, and for the lifecycle half session-lifecycle-readout-contract.md, and for the session-time half session-time-accounting-readout-contract.md.
Two things to know before changing a reader:
- Trim every value a
sqlite3child returns before parsing it. On Windows that child writes CRLF, and a single-column read swallows the\r, soDate.parseanswersNaNfor every row and the whole ledger reads as empty. This silently reported 0 ready tags while the identical predicate counted 573. - Ride
idx_worktree_events_recordedforworktree_events. The event-first plan took 11.7 s for 559 rows where the indexed plan took 0.66 s.
Related
- perf-load-curve.md — the wait half: how long each step takes at each load.
- perf-status.md — the live performance readout this is built from.
Last verified 2026-10-02