---
title: Nighty Tidy — how it works and its limits (part 2)
---

# Nighty Tidy — how it works and its limits (part 2)

## What it is

This is part 2 of the [Nighty Tidy](nighty-tidy-2.md) page. It covers how a run actually happens, and where it stops.

## Where to find it

The same **Nighty Tidy** panel as [part 1](nighty-tidy-2.md).

## How it behaves

### Auto-tagged sessions

Every Nighty Tidy session is tagged automatically so you can find and group them from
the tag filter and the sidebar chips:

- An **audit** session gets two [library tags](tag-manager.md): `nighty tidy` and the
  audit's name (e.g. `Test Coverage`).
- A spawned **apply / implementation "wave"** session gets those two plus `implementation`.

The tags are global (they filter across every project), created once and reused on later
runs (no duplicates), and created without a color — recolor or rename them anytime in
**Settings → Tags**. Tagging is best-effort: if the tag system ever hiccups it is skipped
silently and never disrupts the audit or the wave.

When the feature first ships it also runs a one-time pass over your EXISTING Nighty Tidy
sessions and tags them the same way. That pass runs in the background after startup (it
never slows launch) and only touches sessions that aren't already tagged, so it is safe to
re-run.

Implementation: [nighty-tidy-tags.ts](../../src/main/services/nighty/nighty-tidy-tags.ts);
locked by invariant I48 in the contract. To manage the tags from outside the app, see
[cli-tags.md](cli-tags.md).

### Operator sessions (drive Nighty Tidy with AI)

The panel sidebar carries an **Operator sessions** list with "+ New session".
Each spawns an AI **operator** — a normal Claude session, briefed by a standing-
context primer on what NT2 is and how to drive it — so you can say _"go set up
Nighty Tidy"_ and it does it. The operator acts through a dedicated, gated CLI
control API (`/nighty-tidy-2/*` on the local control server; the panel is in-app React
with no bridge, so a spawned CLI session drives NT2 through this API instead):

- **reads** (bearer + read-budgeted): `GET /nighty-tidy-2/audits | projects |
subscriptions | runs`, `…/runs/:runId/report | findings`.
- **manage automations** (applies immediately): `POST /nighty-tidy-2/subscriptions`
  (body carries an optional `id` to edit, plus `name`, an optional `customPrompt` scope,
  and an optional Claude-family `engine`/`model`; honors `X-Client-Request-Id` — a retried
  save replays the remembered response instead of double-creating), `DELETE
/nighty-tidy-2/subscriptions/:id` — deletes ONE automation by its id (the id rides in the
  URL path; a DELETE request body is dropped by the control server).
- **run an audit now** (**approval-gated** — it spends a run): `POST
/nighty-tidy-2/audits/:slug/run-now` (body may carry an optional `customPrompt` scope +
  an optional Claude-family `engine`/`model`) enqueues `nighty_tidy_2.run_now`, which the
  dispatcher runs only after you approve it in the inbox.

Operator sessions live in their own session-host project (`__nightytidy2_agent__`),
NOT the plugin project — so a session that needs you surfaces in the global **Needs
You** like SMS / Mind Map sessions, instead of being `__plugin__`-hidden. Clicking
one swaps the panel for its chat (the sidebar stays); a nav tab brings the panel
back. Operator sessions are desktop-only; the mobile React panel ships the four
audit screens (Run / Automate / History / Findings), not the operator-sessions dock.

### Known boundaries

- The offered LIST is independent, but two audits sharing a slug share the on-disk
  spec — use a NT2-specific slug to diverge a prompt.
- Concurrent same-slug runs from both panels on one project collide on the
  `audit-<slug>` worktree (one fails gracefully).
- A scheduled audit spanning an app restart loses its in-memory watcher; at next launch
  the startup cleanup RESCUES a run whose audit finished before the restart (its report on
  disk → History records its real success/findings result) and reaps only report-less rows
  to `failed` (the 24h watchdog is the steady-state backstop). Parked
  sessions stay in your inbox either way; only a restart landing BEFORE the report was
  written still records `failed` for work that might have finished. A crash that wrote NO
  new report but left a PRIOR run's report on disk is recorded `failed`, **not** a false
  `findings`: the run's report is kept only when it is genuinely this run's output (a fresh
  file time OR a `_Completed:` stamp at/after the run's start), so a weeks-old leftover can't
  be mis-recorded as this run's result (fixed 2026-08-27).
- **A restart never re-runs an audit that already has a session (fixed 2026-08-25).** The
  cleanup can auto-restart an interrupted run, but only when the run never spawned a session
  at all. Three checks now stand between a restart and a duplicate: the run is left alone if
  an audit session for it exists in ANY state (a restart parks live sessions into a terminal
  status and a recovery sweep re-arms them a minute LATER — that gap spawned duplicates);
  completion is proved from the run's own report folder rather than a file timestamp (on
  Windows a copied report keeps the SOURCE file's older timestamp, so a finished audit read
  as unfinished); and a re-dispatched run whose automation was deleted or disabled in the
  meantime is cancelled, not spawned — it shows in History with that reason.
- Auto-fix (apply) sessions are a separate cap and out of scope for the inbox-stay change
  above — the reported issue + fix were the read-only audit schedule.

## For agents

### How it works

- **Run mode.** The AUDIT always runs read-only (`buildAuditPrompt` forces
  `audit_mode=read-only` and requires the placeholder). Read-write means: after a
  findings audit, `maybeSpawnApply` calls the shared apply engine at the severity —
  a separate, visible fix session that runs autonomously and, at close-out, hands a fully-green
  branch to `/ready-to-merge` so the auto-lander merges the fixes (a red gate / failed finding /
  dependency change is held for review + an inbox alert; the skill's close-out Step 6 owns this).
  Gated on findings + a global apply cap.
- **Cadence scheduler** (`nighty-tidy-2-scheduler.ts`, in `STARTUP_TASKS`). A
  **stateless** once-a-minute tick. For each enabled automation whose current cadence
  window has opened (and opened after the automation was created), it fires the audit that
  automation has gone **longest without running** (`orderByLeastRecentlyRun` scoped to the
  automation — never-run first; ties → canonical catalog order, NOT the saved slug order, so a
  fresh automation runs its audits in catalog order however its slug list was serialized), and
  at most one per automation per tick. Because it always advances that automation's longest-waiting
  audit, successive nights cover its WHOLE list before any audit repeats, then loop — and two
  automations sharing a slug rotate independently. There is no queue table or cursor: the
  `nighty_tidy_2_runs` table (with a nullable `automation_id`) IS the source of truth.
- **Reposition the rotation — "Start from" moves the queue.** The **"Start from"** audit (in
  Advanced) isn't just for new automations: change it any time and the rotation **restarts at that
  audit** and cycles through the rest — so if an automation is up to audit 13, you can move it to 25
  and it continues 25 → 26 → … → wraps back through. It works by stamping a **rotation anchor**
  (a timestamp) whenever you change the start audit; the scheduler then ignores runs from before the
  anchor, so the chosen audit becomes "longest-waiting" and leads. Changing an unrelated setting
  (schedule, model, audits) never moves the queue — only changing the start audit does. The picker
  shows **"Up next: &lt;audit&gt;"** so you can see where the rotation is before moving it. Setting it
  back to "From the top" clears the anchor and returns to plain least-recently-run order. (This only
  reorders which audit runs next — it never resets a "Stop after N runs" cap or your nightly
  count/spend limits.)
  **Per-repo pacing (2026-08-08):** each automation is gated by its OWN three caps — a concurrency
  cap (`concurrency_cap`; the scheduler skips it once its running-run count reaches it), a nightly
  audit-count limit, and a nightly $ spend cap (both `0 = off`, counted since THAT automation's own
  cadence fire).
  **Pacing modes + drip + one-shot manual batch (2026-08-25; concurrency restyled 2026-08-27):** the UI
  presents concurrency as a **"Max simultaneous audits"** number stepper (direct `concurrency_cap`, 1–20)
  and the per-day limit as a pick-a-mode segmented control (`nightly_count_limit` + a `throughput_mode`
  column, `'cap'|'drip'`). `throughput_mode='drip'`
  spaces an automation's fires by `window ÷ N` (a drip gate in the fire loop, INERT unless drip AND
  N>0 — every 'cap' automation stays byte-identical to before). A **paced manual "Run now"** creates a
  one-shot `manual_batch` subscription (`max_cycles=1`) the SAME scheduler drives — bypasses the cron
  window, uses a rolling-24h rate window, fires one-per-tick, and goes dormant after one full pass; the
  upsert handler fires an immediate tick so it starts NOW. "All at once" with no per-day limit keeps
  the immediate per-audit `runAuditNow` path, so manual pacing is strictly OPT-IN (I6). Contract: I11a.
  **"All at once" is an explicit boolean, never the number 20 (2026-09-02, I11b).** When the one/few/many
  control was retired the run screen kept asking `cap === 20`, which is also the stepper's legal maximum —
  so dialling the cap all the way UP was the one value that turned pacing OFF, and a user who set 20
  started all 56 selected audits inside three seconds. Every number 1–20 now paces; only the labelled
  "All" position fires immediately.
  **Cancel is backend work, and the Run screen watches the rows (2026-09-02, I11c).** The footer's Cancel
  stops the dispatch loop, then asks Main to settle every in-flight MANUAL run for that project (queued →
  failed, running → skipped, both reading "Cancelled by you.") and end those audit sessions — rows first,
  sessions second, `automation_id IS NULL` only, so a scheduled automation is never swept up and no F013
  "scheduled audit failed" alert fires. The audit watcher also exits as soon as its run row leaves
  `running`, which releases the in-flight key so the same audits can be re-run at once. The Run screen
  re-reads the run rows on a self-stopping 10s poll while anything is unsettled, so a killed or ended
  session stops reading "Running" without a push and without a UI reload.
  On top sit **global ceilings** — the three former global settings
  (`nightyTidyConcurrencyCap` now the concurrency ceiling, default 8; `nightyTidyCostCapUsd`;
  `nightyTidyNightlyAuditCountLimit`, each `0 = off` for spend/count) — that bound the SUM across
  every automation: the run-engine's limiter hard-caps total concurrency at the ceiling, and the
  cycle count/spend ceiling STOPS all firing for the window when crossed (the $ ceiling also raises
  one dismissible inbox warning; count gates silently; per-repo gates are silent). Counted since the
  most-recent SCHEDULED cadence fire — NOT the calendar day — so an evening batch's post-midnight
  tail can't starve the next run; one fire per automation per tick; kill switch
  `AMC_DISABLE_NIGHTY_TIDY_2_SCHEDULER`. Setting the global USD spend
  cap to `0` (off) disables **every** Nighty Tidy cost cap — the nightly limit above AND a hidden
  per-audit-session runaway ceiling (`AUDIT_SESSION_SPEND_CEILING_USD`, default $50, metered
  API-key runs only, resolved by `nt2SpendCeilingUsd`) that would otherwise still stop a single
  audit at ~$50 — so turning the one visible cap off fully removes cost limits on audits; the 24h
  wall-clock runaway backstop always remains regardless. It also purges runs
  older than 30 days each tick. Each fire is **durably claimed** (`nighty_tidy_2_fire_claims`)
  synchronously BEFORE the fire-and-forget dispatch, so a crash between the fire and the run's
  durable `running` row can't re-fire a second paid audit on restart — at-most-once per cadence
  window (the in-memory attempted-set is just the fast twin; RT-F022). The tick runs as a
  **registered service** (`createRegisteredService`), so a silently-stopped timer is flagged by the
  overdue-liveness auditor + an inbox alert instead of every nightly audit going dark unnoticed (F032).
- **Lifetime cycle cap ("Stop after N runs", I36).** An automation's optional `max_cycles`
  (nullable; NULL = unlimited) caps it at **N full cycles** through its audit list. It is
  enforced **statelessly** from the runs table (no counter column, in keeping with the
  stateless design): the scheduler skips a slug once it has that automation's N runs
  (`countRunsPerSlugForAutomation`, only queried when a cap is set); least-recently-run
  ordering makes this exactly N even passes. The count is **status-blind** (a failed run
  still advances the cycle — a failing slug can't wedge it) AND **`is_deleted`-blind**
  (clearing/deleting History can't reset a cap and re-spend); the ONLY reset it can't survive
  is the 30-day retention purge. **Manual "Run now" never counts** (its run row has no
  `automation_id`). The Automate list enriches each capped automation with a read-only
  `completedCycles` (min per-slug run count) for the "ran X/N" + **Complete** badge; the
  scheduler never reads it. The cap threads the same upsert path as the UI + operator CLI.
- **Unattended audits advance the list (don't stall).** The scheduler walks the
  subscription's audits in least-recently-run order (one fire per tick; the next starts the
  moment a slot frees). Each spawned audit carries `UNATTENDED_AUDIT_CONTRACT` (appended in
  `buildAuditPrompt`): it never asks for a decision; if its type already has a
  `done — complete` audit at the current HEAD it finishes cleanly, re-auditing only when
  HEAD moved since.
- **Run completion — finished audits STAY in your inbox.** An autonomous audit doesn't
  end on its own — it finishes its turn and parks at "Needs You" so you can read it. NT2
  does NOT close it: when an audit finishes AND has written a fresh report, NT2 records the
  run's result and **leaves the session parked in your inbox**. The schedule keeps moving
  because the scheduler paces by in-flight RUNS, not live sessions — the moment the run is
  recorded its slot frees and the next queued audit starts, even though the finished session
  is still in your inbox. If a run instead parks asking a human-blocking question without a
  report, it's recorded as needs-input and also left in your inbox so you can answer it. If an
  audit genuinely gets STUCK, Omniscio's own stall detector — which watches for real silence in the
  output, not a clock — flags it; NT2 records that run and leaves the stuck session in your
  inbox to pick back up. An audit that's still working, or just waiting out a rate limit, is
  **never** force-stopped for it. The ONLY hard stop is a 24-hour runaway backstop, for a
  session that never goes quiet. A startup cleanup reaps run rows orphaned by an app restart
  but does NOT touch the parked sessions. See the stay-in-inbox + never-stall contract (I18).
- **Custom scope prompt.** A run's / automation's optional custom instructions are
  appended by `buildAuditPrompt` as a fenced "ADDITIONAL SCOPE" block — after the audit's
  own body (so pasted `{{vars.*}}` text can't be re-substituted) and before the unattended
  safety contract (which stays the final word). Trimmed and capped at 4000 characters.
- **The session's opening message is a short stand-in, not the full prompt (I51).** Each
  audit session's first chat bubble — and its sidebar / "new session" preview — reads
  **`Run the <Audit Name> audit. (automation: <Automation Name>)`** plus any custom
  instructions you added during setup, never the multi-page audit prompt. The complete prompt
  still drives the agent behind the scenes: it's delivered to the CLI while the visible bubble
  uses `createSessionWithPrompt`'s `displayText` (`buildAuditDisplayText` builds the short
  line). Recovery is safe — the full prompt lives in the agent's own transcript, so a restart
  re-nudges with the short line rather than losing instructions.
- **Which automation fired this run (I51).** A repo can hold several automations, and each
  one spawns sessions named only after the audit — so two automations running the same sweep
  used to look identical. The opening message now carries the firing automation's name, taken
  from the automation's **Name** field, so a glance tells you which run a session belongs to.
  The label is for you, not the agent — it never changes the audit's instructions. A manual
  **Run now** shows no label (no automation fired it), and neither does an automation you
  haven't named yet — name it in the Automate screen and every future run is labelled.
- **Every run gets a UNIQUE name (I62).** An AUDIT session reads **`[NT Audit] <Audit>[ (area)]
  #<N>[ — <label>] · <date-time>`** and a fix WAVE reads **`[NT Apply] <Audit>[ (area)] #<N>:
  Wave <M> — <Severity> (<count> findings) · <date-time>`** — e.g. `[NT Audit] Test Coverage #
  4 — checkout-flow · Aug 26, 8:15 PM`, or a wave `[NT Apply] Error Recovery #7: Wave 2 — High
  (17 findings) · Aug 26, 11:40 PM`. Rather than pile up as identical `<Audit>` rows, the
  **run ordinal `#<N>`** says how many times that audit has been run, and it LEADS the part of
  the title that varies — a sidebar truncates after a few words and every run of one audit
  shares the same audit head, so a number placed any later was cut off on exactly those rows.
  (This supersedes the earlier rule that the WAVE number lead the title.)
  The **Label** is OPTIONAL: type it in the Run screen's "Label" box or set a per-automation
  Label; leave it blank and the name falls back to the automation's name (for a scheduled run),
  then to a short snippet of your custom instructions, then to just the audit + ordinal + date/
  time (always unique via the timestamp). A numbered wave drops the label — the ordinal and the
  wave number already say which slice it is. The label is capped to one clean line, stored on the
  run (History shows it beside the audit name) and on the automation, and it never changes which
  audit a session belongs to — the dedupe guard keys on hidden metadata, not the name.
- **A custom audit can name the AREA it covers.** A custom (non-repo-wide) audit may carry a
  short **Area** — "auth module", "the payments path" — set on its create/edit form. It shows in
  parentheses right after the audit name in BOTH titles, and it is the counter's second axis, so
  each area keeps its own `#1, #2, …` instead of one running total across every area. It is a
  LABEL only: it never changes what the audit looks for (that stays its instructions) and it never
  changes its slug, so saved automations keep working.
- **An audit's findings live in the DATABASE, not in files.** When a run finishes, its findings
  and its fix-wave plan are stored as rows, one per finding, tracking severity, file/line, the
  verification verdict, and how far its fix has got — pending, applied, verified, skipped or
  failed. Two things follow from that. **"What is still open for this audit" becomes a question
  the app can answer** across every past run, instead of something you work out by reading files.
  And **an apply wave reads its own work list from the app** rather than re-parsing a findings
  file that can be several megabytes. A finding the audit itself refuted is kept as history and
  shown in the report's appendix, but never counted as outstanding work. Nothing about how you
  use Nighty Tidy changes — the same runs produce the same reports, from a sturdier store.
- **Engine + model (I35).** A run (`engine`/`model` on the run-now IPC) or automation
  (persisted `nighty_tidy_2_subscriptions.engine`/`.model`) can override which harness +
  model the audit spawns on; the scheduler threads an automation's choice like the scope
  prompt. Both reach the shared `createSessionWithPrompt({ provider, model })`. **Scope is
  the Claude family only** — native Claude + the anthropic-compat vendors (DeepSeek / Kimi /
  GLM / MiniMax / Meta), the `usesClaudeBinary` set: it's exactly what that spawn path
  launches AND what runs the Claude-Code audit skill, so Codex/Gemini (separate session
  managers) are rejected. **Omitted/null inherits today's default** (existing runs +
  automations unchanged). Three guards: the IPC boundary rejects a non-Claude-family engine
  or engine/model mismatch; `resolveEngineForSpawn` re-guards at spawn (a stale/removed
  engine drops to inherit — never mis-spawns codex through the Claude path); and a no-key
  vendor still fails _closed_ at the process spawn (a visible failed run, never a silent
  Claude mis-bill). The shared `EngineModelPicker` (Run + Automate, desktop + mobile) offers
  the family gated by the `showAlternateProviders` master toggle.
- **Fresh report on re-run — per automation.** Completion prefers THIS run's freshly-written
  report (by modified-time) and overwrites the saved copy. A manual run saves to the shared
  `audit-reports/audit-<slug>/`; an **automation** saves to its own
  `audit-reports/audit-<slug>/auto-<automationId>/`, so two automations sweeping the same
  audit keep separate reports instead of clobbering one another.
- **Structured findings.** `getRunFindings` reads the run's `findings.json` and
  normalizes it — severity casing drifts across audits (`{critical,..}` vs
  `{High,Medium,Low}`), so keys + per-finding severity are lowercased; the list is
  capped for render-perf; null falls back to the raw markdown. The canonical
  `findings.json` shape audits should emit is documented in the bundled
  `codebase-audit-orchestration` skill. Report reads are containment-checked: a
  DB-sourced `report_path` outside any `audit-reports` tree is refused (defense in
  depth against a tampered row).
- **Automations.** Multiple named automations per project (each a `nighty_tidy_2_subscriptions`
  row with a `name` + optional `custom_prompt`); saving with an id edits in place and
  preserves `last_run_at` so editing config never re-arms an instant run; deleting targets
  one automation by id.
- **Feature telemetry.** NT2 emits its own `nighty_tidy_2` feature-usage events — audit
  fire / findings / failed and automation create / update / delete — from the service
  chokepoints, so NT2 activity now shows in the Feature Usage dashboard (NT1 was already
  instrumented; NT2 previously emitted nothing). Metadata is counts / slugs / ids / enums
  only: report and finding text never cross the boundary, and a run's custom-scope prompt
  is recorded only as a `hasCustomScope` boolean.

The contract [nighty-tidy-2-contract.md](../../.claude/memory/contracts/nighty-tidy-2-contract.md)
locks the invariants (each by a named test) — including I5 (sha-provenance refresh),
I26 (visible pre-flight failures), I27 (the subscriptions push), I28 (the push-driven run
queue), I31 (the editor's unsaved-changes guard), I35 (the per-run/automation engine +
model choice, Claude-family only), I36 (the "Stop after N runs" lifetime cycle cap), and
I51 (the short synthetic first message that hides the full audit prompt).

## Related

- [Nighty Tidy](nighty-tidy-2.md) — part 1, the owned-versus-shared boundary and how to run one.
- [overseers.md](overseers.md) — the standing coordinator that hands out worker sessions.
- [night-shift.md](night-shift.md) — queuing a plan of work to run while you sleep.

