---
title: The load curve (what agent commands cost at different session counts)
---

# The load curve — how long agent commands take at different session loads

## What it is

## Where to find it

A command-line readout — it takes its own sample of nothing and simply joins records the app already wrote. The **Performance Monitor** panel is where you see the rest of the same picture.

## How it behaves

### What it answers

"Things feel slower when a lot of sessions are running — **which** things, and **how much**
slower?"

Run it:

```
npm run perf:load-curve
npm run perf:load-curve -- --hours=6
npm run perf:load-curve -- --json
npm run perf:load-curve -- --self-test
```

It prints, for each class of agent command, how long that command took at 0-25 live sessions,
25-50, 50-75, 75-100, 100-125 and 125+ — then names the class that grows fastest with load and the
bucket where it first takes twice as long as it did on a quiet box. Below that, a **WORK FINISHED**
section says how much work the fleet actually completed in each of those buckets.

### What it is NOT for

**It never tells you to run fewer sessions.** The hardware on this box is fine and it is meant to
run 100+ sessions at once. The curve exists to show which *per-session cost the app itself must
shrink*. Reading "N sessions is the knee" out of this data is the owner's call, not the tool's —
so the command prints the curve, the growth factor and the crossing point, and stops there.

### The two halves it joins

| Half | Who records it | What it holds |
| --- | --- | --- |
| **Load** — already existed | `governor-tape` (1 row/second) | live sessions, governor band, process births/s, event-loop lag p99 |
| **Cost** — new | `tool-call-timing-tape` (1 row per tool call) | the tool, its command class, how long the agent waited, and the live-session count at that moment |

Nothing previously recorded how long an ordinary agent command took *together with* the session
count at that moment, so the question could only be argued. This command is the join.

### Reading the output

```
WAIT ms — p50/p95 per class, per live-session bucket
  class     0-25           25-50          50-75          75-100         100-125        125+
  git       120/310 n=842  118/340 n=611  140/520 n=902  190/910 n=744  260/1840 n=511 n=12 (thin)
  test      ...
```

- Each cell is **p50/p95 and the sample count** for that class in that bucket.
- `n=12 (thin)` means fewer than 30 samples — the reader refuses to print a percentile it cannot
  stand behind, because at small n a p95 is just the slowest single call.
- A dash means no samples at all.

Below the curve, the same buckets carry the box context — governor tape ticks, lag p99, the
dominant band, caught main-thread stalls, and cloud gate durations — so a slow class can be lined
up against what the box was doing.

The last line names the **most load-sensitive class** — the one whose p95 grows fastest with session
count — and the bucket where it first doubles:

```
MOST LOAD-SENSITIVE CLASS: test — p95 410ms at 0-25 to 1620ms at 100-125 (4.0x), first doubles at 50-75.
  Ranked by p95 GROWTH RATIO, not by absolute cost — a class that starts fastest-growing wins this
  BY CONSTRUCTION, so read the p95 table above before calling it the bottleneck.
```

**This line is a ratio, so it is not a bottleneck ranking, and the distinction is not academic.**
Measured on the live tape 2026-09-20 it named `Edit` — while Edit was in fact *faster* than its
peers: across 1,686 paired moments at the same instant, Edit's median wait was 622 ms against the
other tools' 9,131 ms. It won on ratio alone, because it has the fastest quiet baseline (358 ms p95
at 0-25) and therefore the most room to grow. Read the absolute p95 to judge cost; read this line
to see which class is most sensitive to load.

### The command classes

`git` · `test` · `build` · `search` · `install` · `node` · `net` · `pkg` · `proc` · `wait` ·
`fs` · `shell` · `other`, assigned from the first word or two of a shell command. Tools that run no
command (`Read`, `Edit`, `Grep`, …) are listed under their own tool name instead.

**The full command is never stored.** Classification runs through the same redactor the friction
ledger uses, which keeps only a program basename from a strict character allow-list and stops
walking at the first flag — so a path, a search term, a URL or a flag's value can never reach the
file.

### When a number is missing

Every input is optional and independent. A missing timing tape, governor tape, heartbeat log or
cloud-runs ledger is reported as `UNMEASURED` **with its reason**, never as a zero and never by
silently dropping the section. An empty tape looks exactly like a quiet fleet, which is why
`--self-test` exists: it runs a synthetic tape with a known 4x growth through the real report code
and shows it being found. Run it before believing a flat curve.

### Two honest caveats

- **What "wait" means.** It is wall-clock from the tool call appearing on the session's output to
  its result coming back — hook time (including any governor hold), permission, and the tool
  itself. It is deliberately *not* Claude Code's own `duration_ms`, which excludes exactly the
  hook-and-hold component that grows with load.
- **Parallel calls.** When one assistant message dispatches several tools at once their waits
  overlap, so those rows are excluded from the percentiles and counted separately on the report's
  second line. Before 2026-09-26 the tape counted each streamed piece of a message as its own batch,
  so no row was ever excluded — percentiles from before that date cover every call, and a
  comparison across it mixes two populations.

### What the fleet finished — the WORK FINISHED section

Waits say how long each step took; they cannot say whether less work got done. Since 2026-09-22 the
report ends with a second table in the same live-session buckets: for each bucket, the minutes of
tape covered, the session-hours accrued, the mean live count, the turns, the tool calls (with the
failed ones counted beside them), and three REAL OUTPUT measures — commits landed, ready-to-merge
tags applied, and items brought to the owner's inbox — and then turns and calls per session-hour,
with each of the three per session-hour and per wall-clock hour.

- **Work is counted from records the app already writes** — a completed agent turn, a tool call that
  returned without error, a branch the lander recorded as landed. A failed call is counted beside the
  work, never inside it. Nothing is estimated from elapsed time or activity.
- **Session-hours accrue only over tape the governor actually wrote** — live sessions × the gap to
  the next tick, each gap capped at five seconds, so a silent tape adds no time. A bucket with under
  ten covered minutes prints its counts and `thin`, never a rate; a rate the fold withheld prints as
  `-`.
- **Two rates per bucket, on purpose.** Per session-hour says whether each session gets less done as
  load rises; lands per wall-clock hour says whether the fleet as a whole finishes more or less. A
  stretch covered at zero live sessions has no per-session denominator, so it shows lands per hour
  alone.
- **Each measure states its own coverage, and gates only itself.** Lands, tags and inbox deliveries
  each read from the app's own record and each carry `ok`, `partial — reaches back only to …`, or
  `unmeasured — <reason>`. One unreadable ledger withholds that measure alone; it never blanks the
  other two and never prints as a low number that looks like a slow fleet.
- **A card is a delivery when it REACHED the owner.** An inbox item counts at the moment it landed
  in the inbox, so one still held back is not yet counted — it lands in a later hour, or never.
- **Lands name their repository.** The lander lands for every repository on the machine, so the
  figure is shown with the split, this product first, rather than as a bare total.
- **An unreadable load is `unknown`, never quiet.** An event whose live-session count could not be
  read is counted on its own row, never dropped into the 0-25 bucket.
- **The two peaks are readings.** The last lines name the bucket with the most turns per
  session-hour and the bucket with the most lands per hour, and say that a falling rate at a higher
  bucket names where the app's per-session cost grows — never how many sessions to run.
- **Every bucket is time-correlated.** The minutes the box spent at one load may be one incident —
  the night that prompted this section had one deadlock in it — so a single heavy bucket is compared
  across days before it is read as a load effect. The section prints its blind spots beneath the
  table.

The same fold renders the **Work rate** block in the Performance Monitor, so a bucket can never read
differently on the two surfaces. `--self-test` runs a synthetic tape with a known 2x lower rate in one
bucket through the real fold and reports `WORK RATE OK` when it finds it.

## For agents

### Where the data lives

- Tape: `~/.amc/tool-call-timing.jsonl` (bounded at 8 MB, rotated; honours `$AMC_DATA_DIR`)
- Writer: `src/main/services/diagnostics/tool-call-timing-tape.ts`
- Reader: `scripts/perf/load-curve-report.mjs`
- The work-finished fold: `scripts/perf/work-rate.mjs` (buckets and the time join in `scripts/perf/load-buckets.mjs`), shared with the app's `work-rate-block.ts`
- Lands: read live from `GET /auto-lander/events` on the local CLI server with the CLI token; no token, no app or a bad answer → the lands column reads `unmeasured` with the reason, and the rest of the section still prints
- Ready tags: `worktree_events` rows with `event IN ('ops-ready','ops-ready-user-override')` in the live DB, read through `INDEXED BY idx_worktree_events_recorded`
- Inbox deliveries: `inbox_alert_items`, timed at `COALESCE(held_until, created_at)`
- The share page and its stable link: [perf-work-output.md](perf-work-output.md)
- Kill switch: `AMC_DISABLE_TOOL_CALL_TIMING_TAPE=1` (registry row `tool-call-timing-tape`)

Recording starts for sessions spawned **after** the app next restarts, so the first useful curve
needs a restart plus enough traffic to clear 30 samples in two or more buckets.

## Related

- [perf-status.md](perf-status.md) — the live performance readout this curve is built from.
- [perf-control.md](perf-control.md) — what is being paced right now, and which governor decided it.
- [slow-computer.md](slow-computer.md) — the plain-language guide to a machine that feels slow while agents run.

