---
title: Dev Pipeline
---

# Dev Pipeline

## What it is

**Dev Pipeline** is an opt-in bundled skill that runs a full software-development workflow end to end on one task — investigate, plan, red-team, build (which starts by creating the run's worktree), code-elegance pass, docs, git-prep — pausing at **five approval gates** (one per phase 1–5; phase 6 finishes on its own) so you stay in control.

**Off by default — except on a box set up with the team dev profile**, which enables it for you (`devPipelineSkillEnabled`, applied by `npm run setup`); on an Omniscio dev machine it is likely already on. Otherwise enable it under **Settings → Features → "Enable Dev Pipeline skill"**. Flipping it on installs the skill into your `~/.claude/skills/dev-pipeline/` (this requires Omniscio's global skill-sync, which is on by default); flipping it off removes it. No restart needed. Because it installs into your global skills dir, `/dev-pipeline` then works in **every** project, and — once the config-sync feature is on — is carried to your other AI CLIs too.

**It also installs `/ready-to-merge`** (into `~/.claude/skills/ready-to-merge/`, same toggle). That's the merge-prep gate on its own: rebase once onto the base branch, run the repo's real lint/typecheck/tests, classify each red as pre-existing debt vs. a regression this branch introduced, and — only if it's green for what the branch changed — stamp the SHA-bound `ready-to-merge` tag the git-guardrails push gate validates. The pipeline runs that same gate itself in Phase 6, so you never need `/ready-to-merge` mid-pipeline; it's there for the common case of a branch you built **without** the pipeline that still has to earn the tag before it can be pushed. Like the pipeline, it never pushes and never merges.

### The flow

Invoke it in a Claude Code session with `/dev-pipeline "<task>"` (or "run the dev pipeline on X"). It drives six phases in order; phases 1–5 each STOP at a hard gate where you reply **approve** / **feedback** / **abort**:

| Phase                                   | Ends at gate header                              |
| --------------------------------------- | ------------------------------------------------ |
| 1 — Investigate & plan                  | `## 🔍 🟢 Plan Ready`                            |
| 2 — Red team (10 lenses) + revised plan | `## ⚔️ 🟢 Red Team Complete`                     |
| 3 — Build & verify                      | `## 🔨 🟢 Build Complete`                        |
| 4 — Code elegance (mechanical only)     | `## 💎 🟢 Elegance Pass Complete`                |
| 5 — Docs                                | `## 📝 🟢 Docs Complete`                         |
| 6 — Git prep (branch ready)             | `## 🚀 🟢 Ready to Merge` _(no gate — finishes)_ |

Every gate report reads the same way from one phase to the next — a one-line headline, then a few `##` section headers (one sentence per bullet, like the final "Ready to Merge" report), then the standard **approve / feedback / abort** footer (the "Ready to Merge" report is the exception: no headline, and a one-line bottom summary instead of the footer) — so you never re-learn the layout mid-run. The header's **status dot triages how much input the gate needs from you**: 🟢 nothing (safe to just approve), 🟡 a question to answer, 🔴 a serious call or a blocker — and a non-green dot swaps in an honest title (e.g. `## 🔍 🟡 Plan Not Ready — One Question`), so a green header always means "nothing needed". In the chat, the little pipeline-stepper widget under each report lights the current step to match — green, amber, or red.

By default each gate is a real, observable stop — there is no hands-free seam. (Opt-in **autonomous mode** — see below — self-advances the gates and stops only when something genuinely needs you.) In Phase 6 the pipeline runs **one authoritative, self-contained verification pass** — it runs the project's full lint/typecheck/tests/build itself, in the same session (never trusting a prior run's "it passed" claim), and classifies each failure: a failure already red on the base branch is pre-existing debt and does **not** block, while a regression _this_ branch introduced does. When every check is green or classified pre-existing it stamps the branch's SHA-bound _ready-to-merge tag_; the pipeline emits `## 🚀 🟢 Ready to Merge` **only once that tag is confirmed on the current commit** — so the header always means a genuinely-tagged branch (if it can't be tagged, it reports a 🔴 not-ready status instead). The pipeline never pushes or merges on its own — even in autonomous mode.

**Right-sized rigor.** The Phase-2 red team always applies all 10 lenses, but verdict DEPTH scales to the task: a docs-only or tiny mechanical change gets terse one-line verdicts instead of manufactured paragraphs of hypothetical risk, while anything sizable, risky, or novel gets full depth — the revised plan is always full quality either way.

## Where to find it

### Running inside Omniscio

When you run it inside Omniscio, completing a gate auto-advances to the next phase and the run shows up on the agent status board — Omniscio reads the skill's `.claude/pipeline/state.md` and the byte-exact gate headers above. Outside Omniscio the workflow is identical; you approve each gate yourself. (Those markers are locked to Omniscio's parsers by a conformance test that runs at **every** merge gate, so the integration can't silently drift.)

Six robustness layers keep real-world runs from stalling or mis-advancing:

- **A bookkeeping slip can't cost you the whole run.** The run's `state.md` records which gate it is
  parked at, and the agent is supposed to move that label forward at every phase boundary. Agents
  forget. When the label is left BEHIND — say it still reads "waiting at the plan gate" while the
  agent has just posted its Red Team report — every later gate used to be checked against that stale
  label, match nothing, and quietly refuse to advance, for the rest of the run: one missed update and
  you hand-approve every remaining gate (observed live, 2026-08-25, and again 2026-09-20 at the 🚦
  Standards Check). Omniscio now advances the gate the agent ACTUALLY finished, provided the run's
  own history proves it genuinely reported every earlier gate — the same in-session lineage proof the
  engine-agnostic fallback uses. **Manual gates stay manual:** a gate you left manual — the plan gate
  by default, or any gate you switched off auto-approve — is only stepped past after a real person
  replied following that gate's report; Omniscio's own automatic messages never count. Advancing only
  ever moves FORWARD, never past a hand-back or any other deliberate stop, and every existing veto
  still applies. Both outcomes show in the log: a rescue says it advanced a gate the state file was
  still behind on, and a refusal names both the stale label and the real gate. See the repo's
  `pipeline-auto-advance-contract.md` (in `.claude/memory/contracts/`) invariant `stale-gate-token-backstop`.
- **Label forgiveness.** Agents hand-write `state.md` and sometimes invent label variants (`PHASE_5_VERIFY`, `GATE_3_LOCAL_VERIFY (…)`, `DONE — …` — all observed live). Omniscio normalizes number-prefixed variants to the canonical labels so auto-advance and the panels keep working; a genuinely unknown label stays raw and never advances anything. Advancing still requires the byte-exact gate evidence in the agent's message (two-factor), so forgiveness can't cause a wrong approval.
- **Question-widget veto.** A gate report whose **body** asks you a real question is NEVER auto-approved — even if the agent mistakenly printed an `AUTO-APPROVE: OK` marker beside it. Omniscio asks its canonical question-widget parser (every widget format it can render), fences excluded, so quoted examples don't false-hold. One thing it does NOT hold on: a routine "approve to proceed?" the agent tucks into its own Plain Speak overlay **card** (below the report) — that tail is overlay metadata Omniscio strips before the check, so a green gate that just repeated its approve/feedback/abort footer as a card widget still auto-advances instead of stranding on a question you were never really asked. A genuine decision rides the report body (or a 🟡/🔴 header), which still holds.
- **A HOLD always holds.** If a gate report says `AUTO-APPROVE: HOLD` for a gate anywhere in the message, Omniscio never auto-approves that gate — not even when an `AUTO-APPROVE: OK` for the same gate turns up further down (a recap, an example, the Plain Speak card). Only when every marker is an OK does the last one decide. This holds on every path, including the backstops that rescue a run whose `state.md` token is missing or left behind. Markers and headers inside a fenced code block are illustrations and never count, while a gate header the agent wrapped in backticks still counts as its real header — so a quoted example can't approve a run and a backticked 🔴 header still stops one. See the repo's `pipeline-auto-advance-invariants-1-contract.md` (in `.claude/memory/contracts/`) invariants `the-marker-is-primary` and `a-quoted-marker-is-never-the-verdict`.
- **Any AI engine — even one that skips the state file.** Auto-advance is not Claude-only: cursor/grok, codex, gemini and the other non-Claude engines advance their gates too, with the "approved" delivered through the engine-aware sender (a one-shot engine has no live stdin like Claude). And because some non-Claude agents don't reliably write `state.md`, Omniscio keeps a loop-safe **state.md-free fallback** — it advances a gate on the trailing `AUTO-APPROVE: OK` marker ONLY when the session's own history proves it genuinely reported every earlier gate in its own prior turns (the in-session lineage). A session merely quoting a marker has no such lineage, so it can never self-approve; a single injected turn can't invent prior turns. Kill-switch `AMC_DISABLE_EXTERNAL_PIPELINE_FALLBACK=1` disables just that fallback (never the strict `state.md` path). See the repo's `pipeline-auto-advance-contract.md` (in `.claude/memory/contracts/`) invariant `non-claude-engines-auto-advance`.
- **A non-Claude engine gets the workflow handed to it.** Claude (and the Claude-compatible vendors like DeepSeek / Kimi / GLM) runs the pipeline from the installed `/dev-pipeline` skill. A skill-less engine — Codex, Gemini, OpenCode, Cursor and the like — can't invoke that skill, so on a software-development session in an enabled repo Omniscio injects a compact **portable brief** of the whole workflow (the six phases, the five stop-and-wait gates, and the exact gate markers) straight into that engine's own instructions. It's best-effort — a leaner model follows it less precisely — but a Codex run now drives the same gated pipeline and lights up the panel like a Claude run. The brief also carries the skill's guidance for verifying under a per-command time limit — offload heavy checks (a `--cloud` wrapper) off the engine's clock, and treat a check that can't finish in time (a timeout / missing toolchain) as an environment caveat that keeps the gate green rather than a stop — so a time-limited engine like Codex no longer strands the Build gate on a lint that timed out (the fix for a live 2026-08-15 go-live wave). Claude is unaffected: the brief is only ever sent to a non-Claude engine.
- **A missing to-do tool never stalls a run.** Some harnesses don't expose TodoWrite; the run notes that and carries its six-step backbone in the state file instead.
- **A stranded gate gets picked back up.** Auto-advance normally happens the instant a phase reports done. If that instant is ever missed — a rare timing edge, or Omniscio restarting at exactly the wrong moment — a periodic background check finds any run still parked at a gate whose latest report carries a valid `AUTO-APPROVE: OK` and advances it for you. It's a true backstop, not a competitor to the instant path: it only touches gates you've left on auto-approve, only after one has sat parked a few minutes, and only when nothing newer from you is waiting — so it can never race the normal path, re-approve over you, or double-spend a turn. It also stays on your side when the computer is busy: it used to stand down completely under heavy load, which could leave a valid gate parked for the whole busy stretch — now it just raises the bar (rescuing a gate that's been stuck about fifteen minutes, one at a time) instead of standing down, so a genuinely stuck gate still gets cleared mid-crunch. Kill-switch `AMC_DISABLE_PIPELINE_GATE_RECONCILE=1`. See the repo's `pipeline-auto-advance-contract.md` (in `.claude/memory/contracts/`) invariant `level-triggered-reconcile-backstop`.

### Companion toggles (Omniscio)

Enabling the skill reveals five optional, independent companions under the same **Settings → Features → "Enable Dev Pipeline skill"** row (plus the plan-gate sub-option) — see the repo's `dev-pipeline-companions-contract.md` (in `.claude/memory/contracts/`):

- **Gate auto-approval (per phase)** — for EACH of the five gates (Plan · Red Team · Build · Elegance · Docs) a toggle for whether Omniscio clears it itself or always asks you, plus a **select-all** master. Backed by `pipelineGateAutoApprove`, resolved by `resolveGateAutoApprove` (falling back to the deprecated `pipelineAutoAdvanceEnabled` / `pipelineAutoApprovePlanGate`, so no prior choice is lost — zero migration). Defaults preserve the prior behavior: the Plan gate asks you, while Red Team / Build / Elegance / Docs auto-clear the moment a phase reports done. A gate that asks a question or reports a blocker is never auto-approved, whatever its toggle. The same controls appear in Settings → Features AND the Dev Pipeline panel — both bind the one resolved map.
- **Remind every session to use the pipeline** (`devPipelineReminderEnabled`, default OFF) — writes a short, self-gating note into each spawned session's `.claude/amc-instructions.md`: follow `/dev-pipeline` IF this session is doing software development, otherwise ignore it. Skips the in-app helper bots (Ask Omniscio / Ask-page / Automation Helper) and is independent of the Default Agent Instructions feature.
- **Add the pipeline quick replies** (`devPipelineQuickRepliesEnabled`, shipped default OFF but turned ON for every developer by `npm run setup`) — seeds the "Full Local Git", "Open a Pull Request", "Package My Work Into PRs", and "Worktree Cleanup Analysis" quick replies if they are missing, grouped inside a "Dev Pipeline" FOLDER placed at the top of your quick-reply list (so they stay tidy and, being a folder, out of the composer picker's first-pop-up view). All four are day-to-day git actions, in the order you meet them: "Full Local Git" (a local land-all-ready + worktree/branch cleanup + mobile-rebuild pass, unnumbered), then "Open a Pull Request" (the contributor path — finalize ONE worktree, run the gate, rebase, open the PR, stop), then "Package My Work Into PRs" (the whole-checkout path — read everything your local trunk has that the shared trunk does not, subtract what already landed or is already in an open PR, and slice the rest into PRs **sized on BOTH ends** — too big and a PR sits unmerged collecting conflicts, too small and you drown in merge cycles, since each one costs a branch + a check run + a review + a merge; so: one AREA of the codebase per PR, reviewable in about fifteen minutes, revertible on its own, split on the AREA boundary rather than the commit boundary, two same-area changes that add no extra risk kept in ONE PR rather than split just because they can be named separately, and past roughly eight PRs it goes back and combines same-area neighbours; every file lands in exactly one PR or on a stated left-out list, dependency / migration / security changes always get their own PR regardless of size, and it stops for your approval before opening anything — you can approve all of it, some of it, or send it back), then "Worktree Cleanup Analysis" (a bare `/worktree-cleanup`, which retires the worktrees the other three leave behind — one line on purpose, because that skill already owns every safety rule and restating them here would fork them). The two PR replies are the ones that do NOT auto-submit: opening PRs is outward-facing, so they land in the composer for you to read and edit before they send. One-way: turning it off never deletes quick replies you already have; a reply you already keep — including your own "Full Local Git" — is left exactly where it is (never pulled into the folder), and if you already have all four no folder is created.
  **When they arrive: on every launch, not only when you flip the toggle.** Omniscio checks the bundle at startup, so a reply ADDED to it in a later release still reaches you even though your toggle went on months ago — which it could not before 2026-09-07, when "Package My Work Into PRs" sat undelivered on installs that already had the companion on. Each reply is offered exactly ONCE (recorded in `devPipelineQuickRepliesSeeded`), so one you delete never comes back on the next launch. Switching the toggle off and on again is the deliberate way to ask for the whole set afresh.
  **"Package My Work Into PRs" was hardened after a red-team pass (2026-09-01) and its rules are load-bearing, not boilerplate.** It may never reset / rebase / amend your local trunk or touch a worktree it did not create (the unlanded work often lives ONLY there), it slices in NEW worktrees and proves your trunk still points where it did, it **bails out and says so** if the repo lands work locally instead of by PR (an owner box with an auto-lander — where opening PRs is against the house rules), it measures the delta and comes back with a strategy rather than grinding through an oversized one, it de-duplicates against open PRs by CONTENT rather than sha (a landed PR is a cherry-pick with a NEW sha, so sha matching silently re-publishes), it scans for secrets before any push, and it builds each cut branch standing alone to catch the classic split failure — two PRs that each pass on the author's machine, one of which will not compile on the trunk without its sibling. Each rule maps to a specific failure; read the block comment above the constant before trimming any of them.
  **The set is landing-only by design.** The "Red Team" (skill Phase 2) and "Full Documentation Prompt" (Phase 5) mirrors were seeded here until 2026-09-01 and were removed: the skill already runs those phases itself, so a second hand-fired copy in everyone's picker was duplicate surface, not a shortcut. Red Team is still a normal seeded default in `seed-defaults.ts` — still the single editable copy the Phase-2 parity test pins, so that test is unaffected — it is just no longer duplicated into this folder.
- **Add the workflow auto-replies** (`devPipelineAutoRepliesEnabled`, default OFF) — seeds a small set of preset auto-reply rules for the pipeline's common handoffs (e.g. answering a red-team completion with "proceed"). Two-way by stable id: turning it off removes exactly the presets it added, never rules you wrote.
- **Write a durable note into the repo** (`devPipelineFileNoteEnabled`, default OFF) — **this one edits a git-tracked file in your own repository**, which is why it is called out here: `reconcileDevPipelineFileNote` splices a marker-bracketed managed block (`<!-- AMC-DEV-PIPELINE:note:start … -->` … `<!-- AMC-DEV-PIPELINE:note:end -->`) into the repo's OWN `CLAUDE.md` — else `AGENTS.md`, else it **creates** `CLAUDE.md` — for every repo enabled in the per-repo selection. Unlike the ephemeral reminder above, the note persists in a file you see and commit, so it also reaches sessions Omniscio did not spawn. Written/removed at Claude session spawn: two-way (turning it off strips the block again), idempotent, atomic, fail-open, and independent of `devPipelineReminderEnabled`. This repo's own `CLAUDE.md` carries such a block.

## How it behaves

### The build checklist + definition of done

Before it builds, each run pins down what "done" means — and then tracks the work so nothing slips:

- **Definition of done** — at the plan gate the run lists the concrete "Done when …" outcomes for the task; approving the plan approves them. In Phase 3 each one must be met with real evidence (a test or an observed result) before the build gate can go green — a passing test suite isn't the same as the thing you asked for being finished.
- **Implementation checklist** — alongside those outcomes, Phase 1 breaks the plan into a granular, itemized checklist of every build task and shows it at the plan gate as a `## Build Checklist`, then mirrors it into the session's to-do list so it appears on the agent status board and ticks off live as the build runs. In Phase 3 every item must end up checked off — or explicitly deferred with a reason — before the build gate goes green; an item left unfinished blocks it, exactly like an unmet outcome. A long, multi-part build can't quietly forget a requirement.

### Autonomous mode (opt-in)

Ask for hands-free operation — "run autonomously", "no check-ins", "leave it overnight", or `/dev-pipeline --auto "<task>"` — and the run drives itself through every phase, doing the full work of each, self-approving the gates and **stopping only when something genuinely needs you**: a real decision (scope, a trade-off, anything irreversible), an ambiguous ask, a risky rewrite, or a build it can't get green. It still ends at "Ready to Merge" and still never merges or pushes on its own. The default stays the normal five-stop flow; autonomous mode is per-run, recorded in the run's `state.md` as `## Mode: autonomous` so it survives a mid-run compaction.

### Tiered-model execution (opt-in)

Ask for it — "run tiered", "smart planner, middle executor", "plan on Opus, build on Sonnet", or `/dev-pipeline --tiered "<task>"` — and the run splits the work across two model tiers: a **smart** model does the planning and judging, a **middle** model does the heavy execution. By default only the **Build** phase drops to the middle tier (plan, red-team, elegance, and docs stay smart) — which is exactly "smart plans and judges, middle executes" — but you can reassign any of phases 1–5 (e.g. "tiered, docs on middle too"). The default pair is the **smart** tier for planning and judging and the **worker** tier for execution — tier names, never a pinned model id, so the sentence does not go stale when a new model ships. ("Plan on Opus, build on Sonnet" is how you might *say* it; the pipeline stores the tiers.) You can name a different pair per run.

The smart main loop always keeps the judgment: on a middle phase it delegates only the generative work to a middle-model subagent, then reviews that output and presents the gate itself — the tiering never quietly moves a judgment call to the weaker model. If a middle-model subagent can't actually be run, the pipeline reports **"tiered execution unavailable"** rather than silently running everything on one model while implying a saving. Like autonomous mode it's per-run, recorded in `state.md` as `## Models: tiered` so it survives a mid-run compaction, and the two compose freely. The default stays a single model for the whole run.

**Even without tiered mode, the build already leans cheap by default.** In an ordinary single-model run, the Build phase now _encourages_ handing the actual code-writing to cheaper **Sonnet-level implementer subagents** while the orchestrator reviews and double-checks every diff — the same "cheap writes, smart reviews" economics as the tiered Build tier, but with no opt-in. Crucially this is **decoupled from parallelism** — even work that can't be fanned out in parallel (a shared-git audit fix-wave committed one finding at a time) still hands each edit to a cheaper implementer and reviews it serially (a _sequential cheap-writer_), so "can't parallelize" is never by itself a reason to keep the writing on the expensive model. It's a bias, not a rule: a tiny cohesive change, or work that genuinely needs the smart model's own hand, can stay on the main loop. Tiered mode above is just the formal, recorded, per-phase version of that same split. Details: the `dev-pipeline` skill `CONTRACT.md` invariant 31.

### Independent-review mode (opt-in)

By default the pipeline's reviews are self-review — the same agent that wrote the plan red-teams it, and the build checks its own work. Ask for a second, genuinely independent opinion — "peer review", "bring in other AIs", "independent review", or `/dev-pipeline --peer-review "<task>"` — and at two points the run brings in **blank-slate reviewer subagents** (fresh context, no memory of having written the work) to tear it apart: the **plan** (after the red team) and the **built code** (after it verifies). It's an addition, not a replacement — your existing red team still runs, and the independent pass rides on top.

The reviewers are read-only: each sees only the work plus an adversarial brief and returns findings, never editing anything. The orchestrator then **verifies each finding before acting on it** — aggregating them, discarding false positives, and folding only the real ones in — so a reviewer's mistake (even a bad "drop this check" suggestion) can never turn into a real change. Each affected gate report gains an `## Independent Review` summary of what was flagged, confirmed, discarded, and folded in; if reviewer subagents genuinely can't be spawned, the run says **"independent review unavailable"** rather than implying it happened.

Like the other modes it's per-run, recorded in `state.md` as `## Review: peer` so it survives a mid-run compaction, and it composes freely with autonomous and tiered mode (the reviewers can even run on a cheap model). The default stays self-review, byte-for-byte as before. Details: the `dev-pipeline` skill `CONTRACT.md` invariant 33.

**Want genuinely different AIs? — the cross-vendor tier (`peer-multi`).** Ask for "peer review with other vendors / GPT / Gemini", "review with different AIs", or `/dev-pipeline --peer-multi "<task>"`, and instead of in-session Claude reviewers the run spins up **real sessions on other vendors** (GPT, DeepSeek, Kimi, GLM…) to review the plan and the code, then collects and verifies their findings the same way — never Codex or Gemini specifically, since each of those runs its own separate program instead of the one the locked-down reviewer harness can confine. It's off by default and **costs real money each run** — a per-run **cost cap** (~**Want genuinely different AIs? — the cross-vendor tier (`peer-multi`).** Ask for "peer review with other vendors / GPT / Gemini", "review with different AIs", or `/dev-pipeline --peer-multi "<task>"`, and instead of in-session Claude reviewers the run spins up **real sessions on other vendors** (GPT, DeepSeek, Kimi, GLM…) to review the plan and the code, then collects and verifies their findings the same way — never Codex or Gemini specifically, since each of those runs its own separate program instead of the one the locked-down reviewer harness can confine. It's off by default and **costs real money each run** — a per-run **cost cap** (~$1 by default) plus a preference for cheap vendors like DeepSeek keeps it small; it auto-skips any vendor you haven't set up and says **"cross-vendor review unavailable"** if none are. Two things to know: the outside reviewers are only ever handed the work to review — never any token or secret — and because they run on other vendors, `peer-multi` inherently sends the work-under-review to those vendors' servers. An outside reviewer is **locked down for its whole life**: it can read the code in your repository (including the branch under review and the installed libraries), search the web and read public pages, and nothing else — it cannot run a command, change a file, message anyone, reach services on your own computer or local network, or read your keys or the secret files in the project folder such as `.env`. When it finishes, **the app itself** hands its answer to the agent that asked and archives the reviewer, so a finished reviewer never lands in your inbox — the reviewer is archived either way, so a failed hand-over never leaves it sitting there for you to clear, and its findings stay readable in its own history. A reviewer that can't be locked (an AI engine without the lock, or spawn approval switched on) is refused and never started. So is one whose session would be handed its model credential by Omniscio **at run time** from the session's own identity — a reviewer never holds, and is never given a way to obtain, the app's keys — which is why an engine that could only ever run on Omniscio credits is reported as **not run** before anything starts — with the reason, and with the exact model id that runs the same model on a key you already have, so switching is one edit. An engine your own key can serve simply runs on that key, whatever order your **Who pays & who serves** list puts things in, so naming Kimi gets you a Kimi reviewer as long as you have a Kimi key saved. And if one never answers — held up by its vendor's rate limit, say — the run drops it and reports that review as **not run**, never as done. Omniscio's usage cascade never switches a reviewer to a different AI: the AI you named is the AI that answers, because the review's title and the verdict it hands back both name that engine, and re-running it somewhere else would credit one AI with another's work. When your named AI runs out mid-review it waits for that AI to come back, or ends saying the engine was unavailable. Details: the `dev-pipeline` skill `CONTRACT.md` invariant 34 and the locked-reviewer contract (`.claude/memory/contracts/ai-code-review-locked-reviewer-contract.md`).
 by default) plus a preference for cheap vendors like DeepSeek keeps it small; it auto-skips any vendor you haven't set up and says **"cross-vendor review unavailable"** if none are. Two things to know: the outside reviewers are only ever handed the work to review — never any token or secret — and because they run on other vendors, `peer-multi` inherently sends the work-under-review to those vendors' servers. An outside reviewer is **locked down for its whole life**: it can read the code in your repository (including the branch under review and the installed libraries), search the web and read public pages, and nothing else — it cannot run a command, change a file, message anyone, reach services on your own computer or local network, or read your keys or the secret files in the project folder such as `.env`. **It reads a frozen copy of the change under review**: a code review names its branch's folder, and Omniscio freezes that branch's commit when the reviewer starts and runs the reviewer in a snapshot of it, so what it reads is exactly the change under review rather than whatever checkout happened to be open, and it cannot move mid-review. A code review that names no branch folder is refused instead of being started against the wrong tree, and a review of a change that has **already landed** names the project folder itself — because then the main checkout *is* the change under review. Either way Omniscio tells the reviewer exactly which commit it is reading, so you can always see what a verdict was based on. When it finishes, **the app itself** hands its answer to the agent that asked and archives the reviewer, so a finished reviewer never lands in your inbox — the reviewer is archived either way, so a failed hand-over never leaves it sitting there for you to clear, and its findings stay readable in its own history. A reviewer that can't be locked (an AI engine without the lock, or spawn approval switched on) is refused and never started. So is one whose session would be handed its model credential by Omniscio **at run time** from the session's own identity — a reviewer never holds, and is never given a way to obtain, the app's keys — which is why an engine that could only ever run on Omniscio credits is reported as **not run** before anything starts — with the reason, and with the exact model id that runs the same model on a key you already have, so switching is one edit. An engine your own key can serve simply runs on that key, whatever order your **Who pays & who serves** list puts things in, so naming Kimi gets you a Kimi reviewer as long as you have a Kimi key saved. And if one never answers — held up by its vendor's rate limit, say — the run drops it and reports that review as **not run**, never as done. Omniscio's usage cascade never switches a reviewer to a different AI: the AI you named is the AI that answers, because the review's title and the verdict it hands back both name that engine, and re-running it somewhere else would credit one AI with another's work. When your named AI runs out mid-review it waits for that AI to come back, or ends saying the engine was unavailable. Details: the `dev-pipeline` skill `CONTRACT.md` invariant 34 and the locked-reviewer contract (`.claude/memory/contracts/ai-code-review-locked-reviewer-contract.md`).

### AI Code Review phase (opt-in — a review step after ANY phase)

Independent-review mode above adds fresh-eyes reviewers at two **fixed** points. If you want that review as a proper step you can put **wherever you like**, turn on the **AI Code Review** phase: Dev Pipeline panel → **Setup** → **AI code review**, one switch per phase (after Plan, Red Team, Build, Elegance, or Docs). Switch on one, several, or all five. It's **off everywhere by default**, and with nothing switched on the pipeline runs exactly as it did before.

Each point you enable becomes a real phase with its own gate — `## 🧭 🟢 AI Code Review` — and **the review's own verdict decides whether it interrupts you**. A review that came back clean, or whose findings the run already confirmed and fixed, closes green and the pipeline moves straight on: nothing lands in your inbox. A review that turned up something you should actually weigh in on stops the run and waits for you, exactly like any other gate. That's why there's no auto-approve switch for it — the gate isn't a blanket stop you'd want to turn off, it's a stop that only happens when there's something for you.

In chat it reads as one of the family: its bubble draws the same row of phase discs as every other gate, with its own 🧭 disc lit **immediately after the phase it reviews**, so you can see at a glance where the review sat in the run.

What it does at each point is the same blank-slate mechanic peer-review mode uses: 2–3 fresh-context reviewers see only the work — the diff at a late point, the plan at an early one — plus a fixed review brief, never the orchestrator's reasoning. There are exactly two briefs, and neither is improvised on the spot: after Plan or Red Team every reviewer gets the **plan-review brief** (it attacks the plan from seven angles — the user who hits the edge case, the attacker, the operator mid-crash, the future maintainer, the bill payer, the tester, the skeptic), and after any other phase it gets the **code-review brief**. Both were tuned on measured runs to find real problems without inventing ones that aren't there — "no issues found" is a perfect answer on clean work — and every finding must name a concrete failure and the defense the reviewer checked. The skill keeps them as `review-briefs/plan-review.md` and `review-briefs/code-review.md`. They're read-only and return findings; the orchestrator verifies each one, discards false positives, fixes the real ones, and re-verifies. By default the reviewers are free in-session subagents, and a review switched on that way will **never** spin up paid other-vendor sessions on its own — that stays behind the per-run `peer-multi` opt-in above. Name another company in that phase's reviewer setting instead (Dev Pipeline panel → **Setup** → **AI code review**; the field's placeholder reads `worker, codex, gemini`) and it does start real sessions at that vendor on purpose: **naming one is the authorization**, so it spends money at their rates on every run and the work under review leaves this machine for them — the same trade `peer-multi` makes, chosen once when you set the phase up rather than per run, and under the same rules: the reviewer is locked to reading your code and the web, a reviewer that never answers is reported as **not run**, and when one finishes the app hands its findings to the run and archives it — it does not appear in your inbox.

Every one of those reviewer sessions — here and in `peer-multi` — is titled the same way: `[Plan Review] GLM: <what it reviews>` when it checks the plan, `[Code Review] GLM: <what it reviews>` when it checks the work, with the AI's name filled in by the app and the subject written by the run that asked for the review. More in [review-session-titles.md](review-session-titles.md).

**A reviewer never waits on you.** Each reviewer is archived the moment it answers. Once every reviewer on a review has answered — or the review's 45-minute limit passes — Omniscio hands you the WHOLE review's results in one go, waking the session that asked for it if it had gone to sleep (archived by the app, or its program stopped or crashed), but never one you paused or closed. A session you paused isn't skipped, though: the result is **held** and arrives the moment you next run it, rather than being quietly dropped. A reviewer that never finishes at all — because it was stopped, its program died, or it never got going — clears itself off the board about fifteen minutes after it goes quiet, without you removing it by hand; it deliberately leaves alone any reviewer it can see the app is already bringing back, such as one waiting out its AI company's rate limit. And a reviewer whose session ends without an answer is reported **with the reason** — the app repeats the same explanation it wrote on that session, such as the engine being one it cannot run on — instead of the review waiting out its 45-minute limit only to be told that time passed. Nothing is lost when a delivery doesn't land: Omniscio keeps the result and delivers it as soon as it safely can. A reviewer also opens in the same hub as the session that asked for it, whatever project the request named, so you find it next to the work it reviews and it can read that work's code.

**Where a reviewer reads is decided by the lock, not by luck.** The folder a review names is also the folder the reviewer is started in, so a code review reads its branch by default. A code review that names no branch folder — or names one that is not a checkout of that project — is refused rather than started against the wrong code, and a review of something already merged names the project folder itself. That refusal is the point: before this, a review whose branch was not named simply read whatever was open and reported on it confidently, and nothing in its answer said which tree it had read.

Two practical notes. **Each point you enable spends reviewer tokens on every run**, so switching all five on roughly triples the review work in a pipeline — start with one or two. And where a review lands next to the pre-merge **Standards Audit** (both can sit after Docs), the **review runs first** and the audit stays last. Your choice is stored alongside your custom phases so the pipeline picks it up on every run, and your own custom phases are never touched by these switches. Details: the `dev-pipeline` skill `CONTRACT.md` invariant 34a.

**A repo's shared review steps are the team default — your own settings win on your machine.** A repo can commit review steps in its shared `.claude/dev-pipeline/phases.json`, and they run for everyone who hasn't set that step up themselves. Once you change any review setting for a step in your own Setup tab — which AIs, how many, or whether it's required — your settings run for that step on your machine, and your Setup tab never copies them into the shared file. A step you switch on without changing anything keeps the shared settings, and your Setup tab shows them, so if you then change one setting you start from the team's values. Changing the team default itself means editing the shared file and committing it. **You can always switch a shared step off for yourself, though** — a step that comes from the repo is drawn in your Setup tab too, marked as one the repo sets for everyone, and its switch turns it off **for you alone**: your team keeps its reviews, and switching it back on restores the repo's step on the repo's settings — which is also how you go back to the team default after changing a shared step's settings. Details: `CONTRACT.md` invariant 24b.

**One switch turns every review step off at once — even in runs already under way.** At the top of the same card, **Run AI code review steps** (on by default) stops every review step on this computer, including a shared step from a repo's committed file. It reaches runs that are already going: the app stops listing review steps for every session, whatever that session's own copy of the pipeline files says; a reviewer asked for from a review step is refused and the run is told to skip it (never a failure, even for a required review); and a session already waiting at a review step still moves on. Your per-step choices stay saved, so switching it back on restores them. Reviewers asked for outside those steps — a foreman's review of merged work, or a run's own `peer-multi` review — are not affected. It is the live twin of the `DEV_PIPELINE_DISABLE_BUNDLED_AI_REVIEW=1` environment switch, which still works. Details: `dev-pipeline-panel-contract.md` (`ai-review-switch-reaches-running-work`) and `ai-code-review-locked-reviewer-contract.md` (`switched-off-reviewer-never-starts`).

**Keeping score — which AI finds the most, for the money.** Every reviewer ranks each finding **Blocker**, **Major** or **Minor**; the scorecard counts Blocker and Major as **serious** (it would break behavior, lose data, open a security hole, crash, or break the build) and Minor as **minor** (real but low-stakes). The pipeline checks every finding, and it — not the reviewer — decides what counts: serious and minor include only problems it confirmed, and anything that turned out not to be real is counted as a **false alarm**. Omniscio then logs one result per reviewer it tried, including one that never answered, and adds what that reviewer cost from the reviewer's own session — none of this needs the pipeline to ask for it. Two more things are logged with each review: **whose work was reviewed** (the model the requesting session runs on, unless it names the AI that actually wrote the work — for instance a cheaper builder it handed the coding to), and the pipeline's own **approve or disapprove** of each answered review (disapprove means mostly false alarms, a serious problem it missed that the pipeline then found, or an answer it couldn't use). See the totals in Dev Pipeline panel → **Setup** → **AI code review** → **How each AI has done**, one row per model: reviews, no-answers, serious, minor, false alarms, how often its reviews were approved (blank — never 0% — until one is rated), total cost and cost per review. Below it, **Whose work was reviewed** lists each AI whose work got reviewed, with how many reviews and how many serious problems per review were found in it — which AIs need more review, and which need less. A dash in the cost column means the cost isn't reported — a reviewer that runs inside the pipeline's own session is billed with it, and some AI companies (Gemini, for one) don't report cost to Omniscio yet. Nothing is logged while every review step is switched off, results are kept for a year, and only counts and names are stored — never code or the findings themselves. Agents read the same numbers with `GET /dev-pipeline/ai-review/scorecard`. The review instructions reach the pipeline from the app (`GET /dev-pipeline/phases`), so improving them takes effect on the next run. Details: `ai-review-scorecard-contract.md`.

### Per-phase custom instructions (persist across updates)

Append your own standing instructions to any phase and they're honored on every run — and they **survive updates to the skill**. Because Dev Pipeline is a bundled skill, an update overwrites its shipped files under `~/.claude/skills/dev-pipeline/`; your custom text therefore lives in a separate folder Omniscio never touches: `~/.claude/dev-pipeline/custom/` (override the base dir with `$DEV_PIPELINE_CUSTOM_DIR`). Drop `phase-1.md` … `phase-6.md` to target a single phase, or `all-phases.md` to apply to every phase. Each file is optional and invisible until you create it; at the start of each phase the pipeline reads the matching file(s) and honors them as ADDITIONAL instructions — your text, appended to the phase's built-in steps. They're additive and bounded: custom text can never skip a gate, change a gate header, or make the pipeline push or merge (Phase 6 still never does). Set one up by dropping a markdown file in the folder, or just ask an agent to "add `<instruction>` to phase 3 of the dev-pipeline custom instructions."

### Project-level phases (shared with your whole team)

Custom _phases_ — extra gated steps woven into the flow (distinct from the per-phase instructions above) — can be **committed to a repo** so they become the project's standard and travel to every developer who works on it. Two sources are read and merged, **project-first**:

- **Project manifest** — `<repo>/.claude/dev-pipeline/phases.json`, committed to the repo. It applies automatically to anyone running the pipeline in that repo — no per-dev setup, and independent of the personal "Enable custom phases" toggle — and never leaks to other repos or to people running the pipeline on their own projects, because it lives inside that repo's code.
- **Personal manifest** — `~/.claude/dev-pipeline/phases.json`, your own private phases, applied on whatever project you run the pipeline in (gated by your personal toggle).

On a same-name clash the **project** phase wins, so a shared standard can't be quietly overridden by one dev's personal config. With no project file and the personal toggle off, the pipeline is exactly the built-in six. Project gates are still manual by default (a real stop), so an auto-activated standard never advances or spends on its own. Because project wins, editing such a phase in the Setup-tab panel writes your change THROUGH to that committed project file (and the editor shows the effective, what-actually-runs value), so a personal edit to a shared phase takes effect instead of being silently overridden — you still commit the file yourself. **AI Code Review steps are the one exception:** your own review settings win on your machine, so the editor shows them and never writes them into the shared file (see the AI Code Review section above). Details: [dev-pipeline-panel.md](dev-pipeline-panel.md) and the `dev-pipeline` skill `CONTRACT.md` invariant 24b.

### Always-on standards phases (every run)

Beyond the six built-in phases, every run also gets **two always-on standards phases** — a **🚦 Standards Check** just after the red team (before building) and a **🔬 Standards Audit** just after docs (before git-prep). By default they're auto-approving — they self-advance unless they find a real problem — so they don't add stops to your flow. If you'd rather review one yourself, you can switch either gate to manual in the **Dev Pipeline panel → Setup → Gate auto-approval** (those two live there only; Settings → Features shows the built-in five), and Omniscio then holds it for your approval like any other gate.

Instead of hardcoding any one project's rules, each phase reads a **standards file**, resolved in this order:

- **Your repo's own** `<repo>/.claude/dev-pipeline/standards.md` — a plain-markdown file with two sections, `## Before you build` (read by the check) and `## Before you merge` (read by the audit). Commit it and it travels to everyone on the project; put the repo's real standards in it (point at its design system, its contracts, its docs). **This is the file you customize per repo.**
- **A bundled general fallback** — if a repo has no standards file of its own, the phases fall back to a general staff-engineer standards file shipped with the skill (`standards/general.md`), so the check and audit still do something useful in any repo.

A repo's own file wins over the fallback section-by-section; if a section is missing it uses the general one, and if neither file exists the phases apply general engineering judgment and say so — a missing standards file never blocks a run. The standards file is additive and bounded: it can't skip a gate, change a gate header, or make the pipeline push or merge. To turn the two phases off entirely, set `DEV_PIPELINE_DISABLE_BUNDLED_STANDARDS=1`. Details: the `dev-pipeline` skill `CONTRACT.md` invariant 29.

### The auto-approve marker (how a run can advance itself)

Every gate message ends with one machine-readable line — the last line, on its own:

- **Gates 1–5:** `[DEV-PIPELINE | GATE <n> | AUTO-APPROVE: OK]` when the phase is done and you have nothing to **decide or act on**, or `[DEV-PIPELINE | GATE <n> | AUTO-APPROVE: HOLD — <reason>]` when you genuinely do (a real decision, a question, or a blocker). An informational caveat the agent just wants on the record — an environment/test quirk, an acceptable coverage gap, a coordination note — is NOT a HOLD: it stays `OK` and rides in the report body, so the run flows on instead of stopping you.
- **Gate 6 (the finish):** `[DEV-PIPELINE | GATE 6 | READY-TO-MERGE]` once the branch is tagged ready, or `[DEV-PIPELINE | GATE 6 | NOT-READY — <reason>]` if it couldn't be.

It is the single line an automation reads to decide whether to advance a run for you: `OK` means "safe to approve", `HOLD` means "a human should look first". You never type these — they're for the machinery, and they ride _below_ your plain-English gate report. The header's status dot agrees with this marker: 🟢 pairs with `OK`, 🟡 / 🔴 with `HOLD`. A gate that asks you a real question always carries `HOLD`, and even if an `OK` ever slipped through, the question-widget veto (above) still holds it for you. Custom/standards gate ids are drift-tolerant: the canonical form is `GATE CUSTOM_<SLUG>` (e.g. `CUSTOM_STANDARDS_AUDIT`), but if an agent drops the `CUSTOM_` prefix Omniscio folds it back to canonical instead of stranding the run — the recognition widens, the safety proofs don't. Full rules: the repo's `pipeline-auto-advance-contract.md` (in `.claude/memory/contracts/`).

### Resuming

The skill is re-entrant: if a run's `state.md` exists, invoking it again resumes from the recorded phase instead of restarting — after a mid-run compaction, or when you reply at a gate. Resume is **owner-scoped**, which is what makes it safe to run many pipelines at once: each run stamps its `state.md` with an owner id (inside Omniscio, the session's id), and a resuming agent picks up ONLY its own run — another agent's in-flight run is invisible to it and is never touched. The state lives in the run's worktree (`<worktree>/.claude/pipeline/state.md`) for a normal repo task, or in a private per-session dir (`~/.claude/dev-pipeline-runs/<owner>/`) when there's no repo to isolate in — so two agents working the same project never share one file.

## For agents

### Self-contained worktrees (no setup, no pile-up)

Each run builds in a throwaway git worktree. Three things keep those from ever slowing your machine — automatically, with no external tooling, on any OS:

- **Created only when Build starts.** Investigating, planning and the red team change no code, so they get no worktree: the run keeps its notes in a private per-session folder (`~/.claude/dev-pipeline-runs/<session>/`) and reads the code straight from the base branch. The first step of Build creates the worktree and moves the notes into it. A worktree made at the start used to sit idle through both approval gates while still paying the machine-wide create wait — on 2026-09-23, 61 of 605 fresh worktrees never received a single commit. The Dev Pipeline panel and the status board show a run from its first phase either way.

- **Created OUTSIDE the repo.** By default the worktree goes in a repo-**sibling** `<repo>-worktrees/` folder (a repo that ships a `worktree-locations.json` may override the base). Because it lives outside the repo, search tools (Glob / ripgrep / `find`) never walk it — a big pile of worktrees can't drag down file search or the editor.
- **If that location won't do, you are told — and it moves.** Omniscio will not put a working folder on the system drive, on the same drive as a repo that needs room for its own history, on a drive that is nearly full, or behind a checkout queue so saturated it can't keep up. When any of those holds, new folders for that project are created in a fallback instead and you get an inbox card titled **"Working folders rerouted to a fallback drive"**, naming the fallback folder, the reason, and the location you had configured. It is a notice, not a failure and not something you mis-set: cleanup speeds up while the reroute holds, and **the card clears itself once normal placement resumes**. Nothing is lost and there is nothing to fix — but if you _want_ folders somewhere specific, the per-project **Worktree location** control ([edit-a-project.md](edit-a-project.md)) is the knob, and setting it to a roomy drive outside the repo is the way to keep the reroute from ever firing.
- **Auto-cleaned when you use the pipeline.** On a new run, before Phase 1, the pipeline reaps the FINISHED worktrees of prior runs — but ONLY its **own** (the branch matches the `<type>/<slug>-<timestamp>` naming) and only when they are fully merged into the base branch, clean, idle, and not an active run. It **never uses force** (git itself refuses to delete unsaved or unmerged work), never touches a worktree you made by hand or a run still in progress, logs what it removed (recording each removed commit's SHA, so a mistaken reap is recoverable from git's reflog), and is best-effort — it can never block or fail a run. Off-switch: set `DEV_PIPELINE_DISABLE_REAP=1`.

### Self-contained — no plugins to install

The pipeline is fully self-contained: each phase applies its own build discipline — planning, test-driven development, systematic debugging, and code review — inline, with no external plugin to install or manage. The one built-in helper it uses is Omniscio's own `omniscio-systematic-debugging` skill (Phase 3, when something goes red). There is nothing to set up — `/dev-pipeline` works as soon as the skill is enabled.

## Related

Every control and toggle for this workflow — which repos it watches, the gate auto-approval switch, and the panel's own setup — lives in the [Dev Pipeline panel](dev-pipeline-panel.md), never in Settings. The housekeeping jobs that ride alongside it are documented in [Dev Pipeline Maintenance](dev-pipeline-maintenance.md).
