---
title: Audit Framework (bundled audit skills)
---

# Audit Framework (bundled audit skills)

## What it is

A set of eight interlocking Claude Code skills that together run a **compaction-resilient codebase audit** of any Omniscio-managed repo, then turn the results into commits on a dedicated branch. The audit framework is bundled with Omniscio — when you launch Omniscio, the eight skills are installed into your `~/.claude/skills/` directory automatically (alongside the existing `omniscio-control` and `search-sessions` skills), so any Claude session you spawn from Omniscio can invoke them.

Audits run for hours and survive many context compactions because **disk is canonical**. Every phase writes its progress into `audit-reports/audit-<type>/state.md`, `plan.json`, per-lot findings files, an aggregate, verifications, red-team output, and a final report. If the agent's in-context summary loses an anchor mid-run, it re-reads disk and picks up exactly where it left off.

The eight skills are:

| Skill                          | Role                                                                                              |
| ------------------------------ | ------------------------------------------------------------------------------------------------- |
| `codebase-audit-orchestration` | Master orchestrator — routes by phase, JIT-loads the matching phase skill, owns `state.md`        |
| `audit-coverage-planning`      | Phase 0 — slices the repo into lots sized for one subagent each                                   |
| `audit-lot-discovery`          | Phase 1 — dispatches one discovery subagent per lot in batches of 3, writes `findings/L*.json`    |
| `audit-aggregation`            | Phase 2 — merges, deduplicates, and severity-sorts the per-lot findings into `aggregate.json`     |
| `audit-verification`           | Phase 3 — re-checks each finding against current source, writes `verified.json`                   |
| `audit-red-team`               | Phase 4 — orthogonal sweep for cross-cutting defect classes lot-by-lot would miss                 |
| `audit-final-report`           | Phase 5 — synthesizes a human-readable `report.md` + canonical `findings.json`                    |
| `applying-audit-findings`      | Companion — turns the report into commits on an `apply/<auditType>` branch, one finding at a time |
| `audit-system-decomposition`   | Phase 0 **override** — cuts the repo into whole systems instead of file lots (see below)          |
| `audit-system-assessment`      | Phase 1 **override** — one design verdict per system instead of file-level defects (see below)    |
| `audit-first-principles-review` | Phase 1 **override** — asks whether each unit should exist at all, not whether it is done well   |

The orchestrator + apply skills both **just-in-time load** the phase skills they need, which is why they have to ship together — installing only the orchestrator without the phase skills leaves a session that knows _what to do next_ but can't _do it_.

#### Phase overrides — auditing whole systems

Almost every audit type asks *"where in the codebase is defect X?"*. The default chain is built for that: Phase 0 slices the repo into equal-weight buckets of 30–60 files, and Phase 1 requires every finding be anchored to a file and line.

That shape cannot answer a different, occasionally more important question: **"does this subsystem's design need to change?"** A system can be made entirely of clean files and still be structurally wrong, and no file-level sweep will ever say so.

So a spec may declare a `## Phase overrides` section naming replacement skills for particular phases. The orchestrator reads it before loading each phase skill and routes accordingly; everything else — artifacts, exit checks, checkpoint commits — is unchanged. If a declared override skill is missing, the orchestrator stops rather than silently falling back, because a system audit run through file-level discovery produces confident, useless output.

Two audit types use this today, and they share the Phase 0 decomposition but bring different Phase 1 skills.

**`system-refactor-assessment`** differs from the default chain in three ways:

- **Its unit is a system, not a file** — a vertical capability slice spanning service, IPC, renderer, shared types, DB tables, and contracts.
- **It covers 100% of the source.** The default planner caps at 40 lots × ~60 files, which on a large repo is a single-digit-percent sample. This one maps every in-scope file to exactly one system, parking any genuine residue in an explicit `S00-UNASSIGNED` bucket that is then assessed like any other system.
- **Its output is a verdict, not a defect list** — each system gets `healthy`, `monitor`, `refactor-recommended`, or `refactor-urgent`, eight scored dimensions, and, for anything needing work, a thesis, a redesign sketch, a counted blast radius, and the cost of doing nothing.

#### Phase overrides — questioning the frame itself

**`first-principles-review`** goes one level further out. Every other audit — including the system one — takes the current design as the frame and grades the work inside it. This one interrogates the frame: *"should this be done at all, and does it have to be done this way?"*

It exists because of a specific, reliable failure mode. An AI shown a `SessionOrchestrator` reasons about orchestrating sessions better; it does not ask whether sessions need orchestrating. The code supplies the vocabulary, the vocabulary supplies the problem statement, and the problem statement rules out every answer outside it. The result is an endless supply of improvements to something that maybe should not exist. What it hunts above all is **a design that is a correct answer to a question nobody is asking any more** — the constraint that justified it expired, the design stayed, and nothing looks broken.

It asks that in **two layers, both first-class, both run on every unit**:

- **Layer A — the WHAT.** Should this capability exist? Its rare, high-value result is that something can go away.
- **Layer B — the HOW.** *Given that it should exist* — is the technology, mechanism or process fundamentally the right one? **This is the more common win and it removes nothing.** "Keep this capability entirely, but we are delivering it with the wrong mechanism" is a complete finding. Layer B covers **process as much as technology** — how work is built, tested, gated, released and operated is as load-bearing as any storage choice, and usually less examined.

A `sound` verdict means both layers passed; answering "yes we need this" and stopping is half the job, which the required `layerA`/`layerB` fields make visible rather than invisible.

Three things make it work rather than generate architecture astronomy:

- **Seven mechanical frame-breaking techniques**, applied in a fixed order. The job is derived from user-visible outcomes *before* any implementation is read (reading code first means reading the job off the code). The current approach must then be restated using **no identifier from the code** — a checkable test for whether the reviewer escaped the frame. The most productive of the seven is the **mechanism inventory**: list three or four *fundamentally different* ways the job could be done, where the test of a good list is that the alternatives are **not reachable from the current design by refactoring**. That single rule is what separates "a different approach" from "an ordinary refactor".
- **A mandatory Chesterton's Fence gate.** Nothing may be proposed for removal until why it exists has been established from history, comments, contracts and tests. Searching and finding no reason is a valid answer that *lowers* confidence — never raises it. Phase 3 attacks this answer first.
- **`sound` is a first-class verdict and expected to be the majority**, alongside a required cost-paid-today, a falsifier, a reversibility label, and CONFIRMED-vs-HYPOTHESIS on every claim. Phase 1 runs a three-way calibration check: verdict inflation, frame inheritance (if most findings merely propose a better version of what exists, the audit failed at its own job and re-dispatches), and layer balance (all-`what` is a deletion hunt; all-`how` may mean nobody asked whether anything could go).

It runs in two modes: **unscoped** over the whole codebase (one verdict per system), or **scoped** to a named subject via an automation's scope prompt, where it goes deeper instead of narrower — one verdict per load-bearing decision inside that subject.

Both audits are **advisory-only and must never be auto-applied** (`autoApplyMode: off`). Their findings are whole-subsystem restructures, whole-mechanism swaps, and deletion recommendations — and a Layer B swap is if anything the more dangerous to automate, because it touches the entire unit while looking constructive. An unattended per-finding commit loop pointed at either would act on a classification.

**`systematization` is advisory for the same reason.** It runs on the stock phase chain — no overrides — but each finding proposes building a new system (one entry point, one complete list, one check) and moving several working implementations onto it, so it is a design decision, never a per-finding commit. Its spec carries the same declaration, and the gate refuses it too.

**Four more are advisory on the same terms:** `guardrail-integrity` (does every rule have a check that can fail), `parallel-change-friction` (which file shapes make parallel changes collide), `unfinished-migrations` (which jobs are done two ways after a stalled move) and `domain-language` (which concepts go by several names). Each runs on the stock phase chain and ends in a choice a person must make — changing what the project enforces, reshaping a shared file, deciding a migration's direction, or renaming — so each spec carries the declaration, and the gate refuses all four.

**That is now enforced, not honour-system.** `applying-audit-findings` asks `advisory-audit-gate.mjs` in its Phase 0 pre-flight — before a worktree is created — and an advisory audit type is REFUSED (exit 10) with nothing written. A spec it cannot read is exit 3, INCONCLUSIVE, which means say so and ask rather than treat an unreadable spec as permission. There is deliberately no override flag: the sanctioned path is a human reading the report, picking ONE item, and running it through the normal development pipeline with its own plan and review — which is not this skill. Until 2026-09-10 the rule lived only as prose inside the two specs and the apply skill never opened a spec at all, so an overnight wave was handed 18 `system-refactor-assessment` findings and told to land them unattended. The gate's detector is pinned by `tests/unit/lint/advisory-audit-gate.test.ts`, which executes it over every shipped spec and asserts the exact partition, so a reworded declaration fails the build instead of silently disarming the gate.

## Where to find it

### How to use it

1. **Spawn a session in the repo you want to audit.** Open Omniscio, pick the project (e.g. a real codebase, not a virtual project), click **+ New Session**. The eight audit skills are already installed for you; no extra setup needed.
2. **Ask the session to start an audit by name.** Phrasings the orchestrator recognises: "run a bug-hunt audit", "start a security sweep", "begin a UI/UX audit", "do a test-coverage audit", "audit my error handling". The orchestrator supports 80+ named audit shapes (bug hunt, security sweep, dependency health, test coverage, test hardening, documentation, API design, contract drift, codebase cleanup, race condition audit, performance, efficiency, frontend quality, state management, observability, backup & disaster recovery, product polish, strategic opportunities, and more).
3. **Approve the worktree creation.** The orchestrator creates an isolated worktree at `.worktrees/audit-<type>/` (so the audit doesn't disturb your main checkout) and stamps it with a "do not delete" tag: `git config branch.audit/<type>.description "status: audit-in-progress | <type> audit started <date> — do not delete this worktree"`. Omniscio's worktree-prune sweep skips any worktree whose branch description starts with `status: audit-in-progress` or `status: audit-apply-in-progress`, so the audit can run for hours / days without getting cleaned up underneath it.
4. **Let the audit run.** Phase 0 plans the slicing, Phase 1 dispatches discovery subagents in batches of 3, Phase 2 aggregates, Phase 3 verifies each finding, Phase 4 red-teams cross-cutting classes, Phase 5 writes the report. Each batch and each finding is checkpointed to disk. You can close the session, restart Omniscio, change accounts, hit rate limits — the next turn re-reads `state.md` and resumes.
5. **Read the report.** When Phase 5 finishes, the orchestrator points you at `audit-reports/audit-<type>/report.md`. The structure varies by audit type (a bug-hunt report looks different from a UI/UX one) because each type has its own spec at `~/.claude/audit-types/<type>.md`. At the same close-out the audit **auto-lands its report to master**: after a fail-closed check that the branch changed nothing but its own `audit-reports/audit-<type>/` files (plus the `.gitignore` exception), Phase 5 flips the branch tag to `ready-to-merge` and Omniscio's merge auto-lander lands it — so the report reaches your repo's history with no manual step. If that check finds anything else, the branch stays tagged `audit-complete` for a human to land by hand (the audit path can never carry source changes into master).
6. **Apply the findings.** Tell the orchestrator (or a fresh session in the same repo) "apply this audit" / "fix the audit findings" / "land the audit". That hands control to the `applying-audit-findings` skill, which creates a second worktree at `.worktrees/apply-audit-<type>/` (also tag-protected), checks out an `apply/<auditType>` branch, and walks the findings one-by-one — each finding becomes its own commit. Each lot is dispatched to a fix subagent via the `Agent` tool (invoking `omniscio-systematic-debugging` on a verify-RED), with the subagent's per-commit verify plus the orchestrator's per-lot review as the gate. A manual apply stops at "branch ready for inspection"; an **unattended (autonomous) apply — including a Nighty Tidy auto-apply run — hands a fully-green branch to `/ready-to-merge` so the merge auto-lander lands it** (a red gate, a failed finding, or a dependency change is held for review with an inbox alert instead).

> **Read it as a web page (not raw markdown).** A finished report's `report.md` can run 400 KB+; for a far nicer read, `npm run render:audit-report -- <report folder>` turns that run's `findings.json` into **one self-contained, offline HTML page** — severity-grouped finding cards (code evidence + verifier notes), in-page filter / group-by / search, and the fix-wave plan. One template renders every audit type (it reads the uniform `findings.json`, not the per-type markdown). Add `--publish` to open it on your phone via Omniscio Shares — opt-in, and it prints a sensitivity warning first, since a report maps the codebase's security-sensitive weak spots. Source: [/scripts/render-audit-report.ts](/scripts/render-audit-report.ts).

## How it behaves

### How it works

#### Bundled-skills installer

On every launch, Omniscio's main process calls `ensureBundledSkills()` from [/src/main/services/bundled-skills-installer.ts](/src/main/services/bundled-skills-installer.ts). The installer iterates over `BUNDLED_SKILLS`, which is derived from the central [/src/shared/integration-registry.ts](/src/shared/integration-registry.ts) — every entry's `cliSkillIds` array contributes to the set. The audit-framework entry contributes all eight skills:

```ts
{
  id: 'audit-framework',
  displayName: 'Audit Framework',
  cliSkillIds: [
    'codebase-audit-orchestration',
    'audit-coverage-planning',
    'audit-lot-discovery',
    'audit-aggregation',
    'audit-verification',
    'audit-red-team',
    'audit-final-report',
    'applying-audit-findings'
  ],
  llmDocPath: 'docs/llm-library/audit-framework.md'
}
```

For each skill, the installer:

1. Reads every top-level shippable file in the source dir (sorted): `*.md` skill content PLUS any top-level helper script (`.mjs`/`.cjs`/`.js`) sitting next to `SKILL.md` — e.g. `audit-final-report/compute-waves.mjs`, the Phase-5 wave planner. Subdirectories like `applying-audit-findings/smoke-test-fixture/` are intentionally skipped — they're fixtures, not skill surface — unless the skill's manifest sets `bundleIncludesSubdirs`.
2. Computes a SHA-256 hash over `<filename>\n<bytes>\n` concatenated across all files, so a change to any companion `*-prompt.md` triggers reinstall even when `SKILL.md` is byte-identical.
3. Compares against `~/.claude/skills/<name>/.amc-managed.json` — if the marker file exists with a matching `sourceHash`, skip; if the hash differs, copy the new files and update the marker; if there's no marker but the directory exists, leave it alone (the user installed it themselves).
4. Writes the marker as `{ installedBy: 'agent-mission-control', version: <app version>, installedAt: <ISO>, sourceHash: <hex> }`.

Foreign-marker skip is the safety net — if you've hand-edited a bundled skill, Omniscio won't clobber your changes on the next launch. The bundled-skills opt-out is `Settings → Workflow → Sync skills to Claude Code` (defaults to on).

#### Phase chain (orchestrator)

The orchestrator skill (`codebase-audit-orchestration`) reads `state.md`'s `## Phase` line on every turn and dispatches to the phase skill responsible for that phase. The phase skills don't call each other — they each finish their work, write their output to disk, advance `state.md`, and the orchestrator's next turn picks up the new phase. Compactions are survived because every transition lands on disk before the next subagent dispatch.

Per-audit-type knowledge (what to look for, finding categories, severity scale, report layout, lot sizing) lives at `~/.claude/audit-types/<type>.md`. The audit-prompt library at `https://docs.google.com/document/d/1tB5W85b34nA2QJjW0rFYTvQj27cL3b4A_FWMd-MW0Sw/edit` is the authoritative list of ~44 audit shapes; new shapes get a spec file added under `~/.claude/audit-types/` without changing the orchestrator's generic backbone.

#### Worktree-tag protection

The orchestrator stamps both audit worktrees and apply worktrees with branch descriptions that prefix `status: audit-in-progress` / `status: audit-apply-in-progress` followed by a timestamp. Omniscio's worktree-prune sweep treats any worktree whose branch has one of those descriptions as "in flight — leave alone". This is purely advisory (a human can still delete the worktree by hand if needed), but it stops well-meaning multi-agent cleanup loops from deleting an audit that's two hours into Phase 1.

#### Packaging

For development builds (`npm run dev`) the installer reads from `<app path>/.claude/skills/`. For packaged builds (`npm run package`), electron-builder's `extraResources` config in [package.json](/package.json) copies `.claude/skills/` → `bundled-skills/` in the final resources folder, and the installer reads from `process.resourcesPath/bundled-skills/`. End users never see the bundled-skills directory — Omniscio just makes sure `~/.claude/skills/` is in sync on launch.

## Related

- [use-skills.md](use-skills.md) — managing skills generally in the Omniscio Skills view
- The audit-prompt library lives in a Google Doc on the owner's Google account; the `gog` skill is the authenticated reader for it
