Omniscio documentation
Browse all documentation
  1. Getting Started13
  2. Sessions & Agents115
  3. Inbox & Notifications59
  4. Projects & Tasks95
  5. Automation & Scheduling75
  6. Knowledge & Memory26
  7. AI Features60
  8. Integrations100
  9. Plugins & Marketplace33
  10. Cloud & Teams56
  11. Settings & Customization58
  12. Account & Billing28
  13. Troubleshooting84
  14. CLI & API Reference22
  15. Legal & Policies4
  16. Uncategorised22

Gauntlet Loop (build, blind-judge, repeat until a builder passes)

A developer panel that runs several builder agents at one goal, has fresh blind critics score every attempt against a concrete bar, and loops until an attempt clears your pass score or the run hits its round ceiling or cost cap.

What it is

A developer-tools panel inside Omniscio that runs Matt Shumer's "gauntlet loop" over a goal:

  1. Several builder agents each independently produce an attempt at the goal.
  2. Fresh, blind critic agents score every attempt (0-100) against a concrete bar — a critic never learns which builder wrote which attempt.
  3. It loops — a new round of builders, then critics — until a builder's attempt clears your pass score, or the run hits its round ceiling or its cost cap.

v1 handles TEXT goals — a written artifact (an announcement, a spec, a summary, a piece of copy). A code mode is a planned fast-follow. Every run spawns real, paid agents, so a run is cost-capped and lives under the Developer Tools sidebar group as an in-development / gated, desktop-only panel.

Where to find it

Enabling it

It ships hidden. Reveal it via Settings → Lab → Gauntlet Loop (gauntletLoopEnabled) or the env flag AMC_SHOW_GAUNTLET_LOOP=1. Once on, a Gauntlet Loop row appears in the Developer Tools sidebar group.

Starting a run

The panel opens on a new-run form:

  • Project — which project the run operates on (real folder projects only; defaults to your active project).
  • Goal — a textarea describing the written artifact you want. Ctrl+Enter in the goal box runs the gauntlet.
  • Builders / Critics / Rounds / Pass score / Cost cap — compact numeric knobs, prefilled from sensible defaults (3 builders, 3 critics, 3 rounds, pass ≥ 80, $10 cap) and bounded (1-6 builders/critics/rounds, 1-100 score, $0.50-$100 cap).
  • Run gauntlet — the primary action. It acknowledges instantly (a rapid double-click can't fire two paid runs) and the run then proceeds in the background.

How it behaves

Watching a run

Below the form is a history list of past runs — each row shows a status badge (running / paused / passed / failed / aborted), the goal, its round progress, its cost, and how long ago it started. Clicking a row opens the detail view:

  • The run header — status, done / max rounds, spent of cap, the pass threshold, the goal, and (while running) an Abort button.
  • The winning artifact — when a builder passed, its text is shown (with the winning round number).
  • Each round — a collapsible card per round showing every builder's attempt (its produced text + cost) and the critic scores, with the winning attempt flagged.

Runs update live: as the builder/critic agents finish, the main process pushes an update and the panel re-fetches the list and any open detail (so a round completing appears without a manual refresh).

Paused runs — add budget & resume

If a run reaches its cost cap before a builder passes, it pauses (rather than failing) and its detail view shows an "Add budget & resume" control: raise the cap and the run continues from where it left off, spawning more paid agents (with a confirm, since it spends more money).

Deleting a run

A run row's hover action deletes it and its rounds from history (with a confirm — it can't be undone).

Safety & scope

  • Desktop-only. Because a run spawns multiple paid agents, the whole request surface is refused over the mobile/web bridge — you drive the gauntlet from the desktop app. The sidebar row and panel are hidden on a phone.
  • Cost-capped. Every run carries a hard USD ceiling; it pauses at the cap rather than running away.
  • Non-spawnable virtual project. You don't start Claude sessions "inside" the Gauntlet Loop project from the sidebar — the builder/critic agents are launched by the run engine.

Related

Bake-Off (fan-out) is the lighter cousin: one round, several models, the same prompt, and you pick the winner yourself.

Last verified 2026-09-23