Gauntlet Loop (build, blind-judge, repeat until a builder passes)
A developer panel that runs several builder agents at one goal, has fresh blind critics score every attempt against a concrete bar, and loops until an attempt clears your pass score or the run hits its round ceiling or cost cap.
What it is
A developer-tools panel inside Omniscio that runs Matt Shumer's "gauntlet loop" over a goal:
- Several builder agents each independently produce an attempt at the goal.
- Fresh, blind critic agents score every attempt (0-100) against a concrete bar — a critic never learns which builder wrote which attempt.
- It loops — a new round of builders, then critics — until a builder's attempt clears your pass score, or the run hits its round ceiling or its cost cap.
v1 handles TEXT goals — a written artifact (an announcement, a spec, a summary, a piece of copy). A code mode is a planned fast-follow. Every run spawns real, paid agents, so a run is cost-capped and lives under the Developer Tools sidebar group as an in-development / gated, desktop-only panel.
Where to find it
Enabling it
It ships hidden. Reveal it via Settings → Lab → Gauntlet Loop (gauntletLoopEnabled) or the env flag AMC_SHOW_GAUNTLET_LOOP=1. Once on, a Gauntlet Loop row appears in the Developer Tools sidebar group.
Starting a run
The panel opens on a new-run form:
- Project — which project the run operates on (real folder projects only; defaults to your active project).
- Goal — a textarea describing the written artifact you want.
Ctrl+Enterin the goal box runs the gauntlet. - Builders / Critics / Rounds / Pass score / Cost cap — compact numeric knobs, prefilled from sensible defaults (3 builders, 3 critics, 3 rounds, pass ≥ 80, $10 cap) and bounded (1-6 builders/critics/rounds, 1-100 score, $0.50-$100 cap).
- Run gauntlet — the primary action. It acknowledges instantly (a rapid double-click can't fire two paid runs) and the run then proceeds in the background.
How it behaves
Watching a run
Below the form is a history list of past runs — each row shows a status badge (running / paused / passed / failed / aborted), the goal, its round progress, its cost, and how long ago it started. Clicking a row opens the detail view:
- The run header — status,
done / max rounds,spent of cap, the pass threshold, the goal, and (while running) an Abort button. - The winning artifact — when a builder passed, its text is shown (with the winning round number).
- Each round — a collapsible card per round showing every builder's attempt (its produced text + cost) and the critic scores, with the winning attempt flagged.
Runs update live: as the builder/critic agents finish, the main process pushes an update and the panel re-fetches the list and any open detail (so a round completing appears without a manual refresh).
Paused runs — add budget & resume
If a run reaches its cost cap before a builder passes, it pauses (rather than failing) and its detail view shows an "Add budget & resume" control: raise the cap and the run continues from where it left off, spawning more paid agents (with a confirm, since it spends more money).
Deleting a run
A run row's hover action deletes it and its rounds from history (with a confirm — it can't be undone).
Safety & scope
- Desktop-only. Because a run spawns multiple paid agents, the whole request surface is refused over the mobile/web bridge — you drive the gauntlet from the desktop app. The sidebar row and panel are hidden on a phone.
- Cost-capped. Every run carries a hard USD ceiling; it pauses at the cap rather than running away.
- Non-spawnable virtual project. You don't start Claude sessions "inside" the Gauntlet Loop project from the sidebar — the builder/critic agents are launched by the run engine.
Related
Bake-Off (fan-out) is the lighter cousin: one round, several models, the same prompt, and you pick the winner yourself.
Last verified 2026-09-23