---
title: Gauntlet Loop (build, blind-judge, repeat until a builder passes)
---

# Gauntlet Loop (build, blind-judge, repeat until a builder passes)

## What it is

A developer-tools panel inside Omniscio that runs Matt Shumer's **"gauntlet loop"** over a goal:

1. Several **builder** agents each independently produce an attempt at the goal.
2. Fresh, **blind critic** agents score every attempt (0-100) against a concrete bar — a critic never learns which builder wrote which attempt.
3. It **loops** — a new round of builders, then critics — until a builder's attempt clears your pass score, or the run hits its round ceiling or its cost cap.

**v1 handles TEXT goals** — a written artifact (an announcement, a spec, a summary, a piece of copy). A `code` mode is a planned fast-follow. Every run **spawns real, paid agents**, so a run is **cost-capped** and lives under the **Developer Tools** sidebar group as an **in-development / gated**, **desktop-only** panel.

## Where to find it

### Enabling it

It ships hidden. Reveal it via **Settings → Lab → Gauntlet Loop** (`gauntletLoopEnabled`) or the env flag `AMC_SHOW_GAUNTLET_LOOP=1`. Once on, a **Gauntlet Loop** row appears in the Developer Tools sidebar group.

### Starting a run

The panel opens on a **new-run form**:

- **Project** — which project the run operates on (real folder projects only; defaults to your active project).
- **Goal** — a textarea describing the written artifact you want. `Ctrl+Enter` in the goal box runs the gauntlet.
- **Builders / Critics / Rounds / Pass score / Cost cap** — compact numeric knobs, prefilled from sensible defaults (3 builders, 3 critics, 3 rounds, pass ≥ 80, $10 cap) and bounded (1-6 builders/critics/rounds, 1-100 score, $0.50-$100 cap).
- **Run gauntlet** — the primary action. It acknowledges instantly (a rapid double-click can't fire two paid runs) and the run then proceeds in the background.

## How it behaves

### Watching a run

Below the form is a **history list** of past runs — each row shows a **status** badge (running / paused / passed / failed / aborted), the goal, its round progress, its cost, and how long ago it started. Clicking a row opens the **detail** view:

- The run **header** — status, `done / max rounds`, `spent of cap`, the pass threshold, the goal, and (while running) an **Abort** button.
- The **winning artifact** — when a builder passed, its text is shown (with the winning round number).
- **Each round** — a collapsible card per round showing every **builder's attempt** (its produced text + cost) and the **critic scores**, with the winning attempt flagged.

Runs update live: as the builder/critic agents finish, the main process pushes an update and the panel re-fetches the list and any open detail (so a round completing appears without a manual refresh).

### Paused runs — add budget & resume

If a run reaches its **cost cap** before a builder passes, it **pauses** (rather than failing) and its detail view shows an **"Add budget & resume"** control: raise the cap and the run continues from where it left off, spawning more paid agents (with a confirm, since it spends more money).

### Deleting a run

A run row's hover action **deletes** it and its rounds from history (with a confirm — it can't be undone).

### Safety & scope

- **Desktop-only.** Because a run spawns multiple paid agents, the whole request surface is refused over the mobile/web bridge — you drive the gauntlet from the desktop app. The sidebar row and panel are hidden on a phone.
- **Cost-capped.** Every run carries a hard USD ceiling; it pauses at the cap rather than running away.
- **Non-spawnable virtual project.** You don't start Claude sessions "inside" the Gauntlet Loop project from the sidebar — the builder/critic agents are launched by the run engine.

## Related

[Bake-Off (fan-out)](fanout.md) is the lighter cousin: one round, several models, the same prompt, and you pick the winner yourself.
