---
title: Nighty Tidy (an overnight audit that only touches its own work)
---

# Nighty Tidy (the full-vision audit panel)

## What it is

> **Shipped label / internal id:** users see this panel as **Nighty Tidy** — the "2" was dropped once the legacy original (`nightytidy`) was retired from view. Its internal id, DB tables, and IPC channels stay `nightytidy2` / `nighty_tidy_2_*` / `NIGHTY_TIDY_2_*` unchanged; this page uses **NT2** for that internal identity.

### What it is

**Nighty Tidy** (`nightytidy2`) is THE Nighty Tidy going forward — the original
(`nightytidy`) is left untouched as legacy. It is the full code-maintenance audit
panel, with four screens:

- **Run an audit** — pick a project (the repo picker shows each repo's icon — its
  real logo or a generated tile), pick a **run mode** (Read-only = report only, or
  Read-write + a severity = auto-fix the findings), tick audits from its own offered
  list (a user-edited spec shows an **"Edited"** badge; the four Omniscio-infrastructure
  audits are badged **"Omniscio-specific"** and are offered ONLY on your Omniscio checkout —
  they describe the app's own internals, so they are hidden on every other repo rather than
  producing Omniscio-flavoured findings there; if you have no Omniscio checkout registered
  they stay offered everywhere, as before), narrow what gets audited with the **Scope** section
  (always visible): quick-scope **preset buttons** (Frontend / Backend / Tests / Shared /
  Recent changes), a **folder browser** modal to pick specific directories, selected
  folders shown as removable **chips**, and a free-text **custom instructions** textarea
  (all composed into the scope prompt appended to every audit; char counter shows
  combined remaining vs the 4000-char limit). Optionally pick an **engine + model**
  (a Claude-family harness — default inherits your usual engine/model; I35), **Start Run**
  (one regular session per audit, billed normally).
  The panel is **compact**: Repository + Run mode sit together, the engine + model section
  tucks behind a small toggle that opens in place — and auto-opens when a value is already
  set, so nothing set is ever hidden (I39). The queue
  below shows REAL progress — each row is driven by run-updated pushes (queued →
  running → completed/failed; never invented client-side), and a **"Done — new run"**
  button appears only once every audit settled.
- **Automate** — the **automations list**: automations **grouped under each repo**. A repo's
  header shows its name + one **outcome signal** (last night's worst severity "1 critical" in red /
  "Clean" / blank if never run); beneath it sits **one row per automation** — its **label** (or
  "Automation 1 / 2 …"), a one-line summary (**schedule · audits · auto-fix**), and its OWN
  **On/Off toggle** (enables/disables just that automation — never deletes). A quiet footer reads
  **"Last night: N audits · $X · Limits"** (Limits opens the global **ceilings** editor); only repos
  that already have an automation show here, with an empty state prompting the first.
  **"+ New automation"** (page header) opens a **step-by-step wizard** — one page each:
  **Repo → Audits → Instructions → Model → Schedule → Auto-fix → (optional) Advanced → Name & review
  → Create**. It gathers every choice and saves ONE automation at the end (the same upsert the
  quick-edit pickers use); you can't finish without a repo and ≥1 audit, and a brand-new automation
  is **read-only** (auto-fix OFF) unless you pick a fix level on the Auto-fix page. The wizard reuses
  the very same controls as the edit pickers, so the two never diverge.
  **Click an automation row → its quick-edit detail**: a findings summary on top, then settings rows —
  **Name** · **Schedule** · **Audits** · **Instructions** · **Auto-fix** · **Pacing** · **Advanced** —
  each shows its value inline and opens one small picker; changes apply the moment you make them (no
  Save button). The **Pacing** picker: **Max simultaneous audits** — a −/+ number stepper (1–20; replaced
  the old one/few/many mode control 2026-08-27). An automation always paces — only the **Run now**
  screen offers one more position above 20, labelled **"All"**, and it says on screen how many of your
  selected audits that will start at once — and **Limit how
  many per day?** — _no limit_ / _cap N/day_ / _spread N/day evenly_ (the **drip**: starts ~N spaced
  across the day instead of all at the cadence time). The nightly **$ spend cap** moves under an
  Advanced disclosure (each `0` = off; the $ is an estimate, not real charges on a managed plan).
  **Advanced** holds a cheaper **engine +
  model**, a **"Start from"** audit (which audit the rotation begins at — shown when 2+ audits are
  selected), and a **"Stop after N runs"**
  lifetime cap; once a capped automation has run N full passes it reads **"Finished N of N runs"**
  (dormant — raise the cap to re-arm). Remove one with **"Delete this automation"** (a confirm, plus
  Undo). A repo can hold as **many automations** as you like (e.g. one per feature) — each an
  independent row with its own schedule, audits, and toggle.
  Each automation runs paced by its OWN Pacing caps, all bounded by **global
  ceilings** (the concurrency / nightly-count / nightly-$ settings in the global editor, capping the
  SUM across every automation) — advancing through each automation's list before repeating, then
  looping, optionally auto-fixing at a severity.
- **History** — run rows with status, finding-count badges, and a "fixes applied"
  marker. A run that couldn't even START (recipe missing / repo gone) records a
  visible **failed** row with humanized copy — a broken automation is never silent.
  A run that FINISHES is recorded by its **result** (findings, or a clean pass) —
  including one that finishes and then parks in your Needs-You inbox on a trailing
  question; a completed audit is never mislabeled **failed** (and no false "scheduled
  audit failed" alert fires for it). A **clean pass is only ever claimed when the audit's
  findings list was read and is empty**: an audit whose session ended or was archived
  before it wrote its report records **failed** ("stopped before it wrote its report"), and
  one whose report came without its findings list records **failed** too, with the report
  still there to open. **Delete a run** (per-row trash) or **Clear** the
  whole list — both soft-delete, behind a quick confirm.
- **Findings** — an **intelligent, structured display** (severity summary + filter
  pills + group-by + expandable finding cards) read from the audit's `findings.json`,
  instead of the giant raw-markdown report (which stays as a fallback). The
  no-findings state is status-aware: still-running, failed (see History for the
  reason), or **"Clean pass"**.

It has its own run history (`nighty_tidy_2_runs`), its own subscriptions
(`nighty_tidy_2_subscriptions`), its own on/off switch (Settings → Plugins →
**Nighty Tidy**), and a moon-icon sidebar row in the sidebar's **Plugins** section, which is
expanded by default. It ships **on by default** — a one-time backfill enables it, so the row is
visible out of the box; you can turn the plugin off afterward. (Until 2026-09-04 the row nested
under **Developer Tools**, which ships collapsed — that hid the row entirely and left the plugin
unreachable, since its row is its only entry point. You can also reach it from Settings → Plugins
→ **Open**.) Its panel sidebar is a **native
column** — brand → the screens as vertical nav → an **Operator sessions** list
(see below). **The original Nighty Tidy is touched zero times.**

**Desktop is a marketplace WEBVIEW plugin; mobile is native React.** On desktop the panel
renders through the generic plugin webview (`src/plugins/nightytidy2/ui/`) so it can
update without shipping a whole app release; the webview **consumes the app's injected
theme tokens** — surfaces, accent, semantic status colours, fonts, and the effective font
size (plugin-webview-theme-contract `optional-look-and-feel-tokens`) — so it matches the app's look, custom themes,
and text size, and draws its own 5-screen nav rail (Run · Automate · History · Findings ·
Custom audits).

**The webview carries the app's UI standards by hand (contract I61).** `src/plugins` is
excluded from the root tsconfig, eslint, and every renderer-globbed i18n / a11y /
design-system guard, so NOTHING in the host toolchain reads this panel's code — which is
how its number dropdowns drifted until they could not commit a value at all. It therefore
ships its own equivalents of the app's primitives, and its own tests are the only thing
holding them up.

_What holds, and is guard-locked:_ `mountDropdown` (themed, owns a LIVE selection — it
repaints its trigger on every pick and compares each pick against its own current value,
so returning to the value it opened with still commits; Enter/Space commit on their own
branch rather than stepping the highlight) replaces the native `<select>`, which is
build-banned here exactly as in the renderer, alongside native date/time inputs.
`attachModalA11y` gives every modal Escape, a Tab trap, focus return, **Ctrl/Cmd+W**,
**Ctrl/Cmd+Enter to submit**, and a labelled close (X). The open menu is **portaled to the
document body**, so a modal's `overflow: hidden` can no longer clip a long list to the first
handful of rows — and because a portaled menu no longer dies with the container that owns its
trigger, it watches for that trigger leaving the page and closes itself (the wizard empties its
container, and a modal's close removes only its overlay; neither fires the outside-click handler,
which listens for a mouse press a keyboard activation never sends). An open list owns **Escape**
and **Ctrl/Cmd+Enter** ahead of the modal, found by its mount rather than its menu.
`mountTimeField` — hour / minute / AM-PM over the same `"HH:MM"` string — replaces the native
time input, whose OS-painted chrome rendered light-on-dark and which could be cleared to an empty
value the schedule commit never validated. Every user-facing string resolves through a
`rendererT` seam (`ui/lib/i18n.js`, a global classic script because an `export` blanks a panel
loaded as plain `<script>` tags); it ships **English-only**, so the rendered copy is unchanged —
proven by extracting every rendered string before and after the sweep.

_What is NOT built, stated so this page does not overstate it:_ there is no named
`nt2CloseOpenDropdown` helper (the Escape arbitration above is what it was for), and the React
desktop panel is still on disk with no mount path.

On mobile the panel is native React (`NightyTidy2MobilePanel`) because a
mobile-web surface can't host an Electron `<webview>`. Both talk to the SAME
`NIGHTY_TIDY_2_*` IPC/service + live pushes (`…RUN_UPDATED` / `…SUBSCRIPTIONS_UPDATED` —
forwarded to mobile web clients via `INBOX_ALLOWED_CHANNELS`), so runs, history, and
automations stay in sync (I21). Operator sessions are a desktop feature; the mobile panel
is the audit screens only.

**On a phone, every screen's buttons stay put.** Each mobile screen is a bounded column
(the shared `ScreenScaffold`): one scrolling body with its header and its action bar pinned
above and below it, so the thing you are reaching for never scrolls off. Concretely — the
new-automation wizard's **Back / Next**, the Run screen's **Start Run** bar, Automate's
**New automation** and its nightly-spend line, a repo's back link and its add/delete
buttons, History's **Clear / Refresh** and status chips, Findings' way back to History, and
Custom Audits' **New custom audit** — all sit outside the scroll. The audit list itself no
longer scrolls separately either: it is part of the page, so a drag moves the screen instead
of being swallowed by an inner box, and each audit row is a full 44px touch target. Before
this, picking audits for a new automation put **Next** below every audit row, under a scroller
you had to reach the bottom of first, and the Automate tab's **New automation** button ran
off the right-hand side of a phone screen entirely (I33). (The desktop webview rebuild is COMPLETE — all five screens are
at full parity with the native mobile ones: Run (engine/model, run-mode + severity, custom
scope, live progress), Automate (the multi-automation manager — grouped-by-repo list + cross-repo
overview + an in-panel wizard & editor, live-reload on `…SUBSCRIPTIONS_UPDATED`), History + the
markdown Report, Findings (a run-list → per-run structured detail: severity bar,
severity/search/group-by filters, collapsible finding cards, and the applied-fixes card, with
a markdown fallback), and Custom audits (a global list + in-panel editor that never sends a
slug — the service derives the immutable one — live-reload on `…CUSTOM_AUDITS_UPDATED`).
Editors guard unsaved changes and confirm deletes with two-tap inline confirms in place of a
dialog.)

## Where to find it

The **Nighty Tidy** panel in the app, where a run is started and its findings are read. The sessions it starts are tagged so you can tell them apart from your own.

## How it behaves

### The owned-vs-shared boundary

The audit-EXECUTION + apply layers are **shared framework**: the bundled
`codebase-audit` recipe defers to the `codebase-audit-orchestration` skill, and the
**apply engine** (`spawnAndWatchApplySession` + the `applying-audit-findings` skill)
is shared. NT2 REUSES them by import — never editing an NT1 source. Provisioning keeps
the installed copies in `~/.claude` **current via a sha-provenance refresh**: missing
files install, stale-but-untouched files refresh atomically when a newer bundle ships,
and anything the user edited is NEVER overwritten (kept + badged "Edited" in the
picker; ledger at `~/.claude/.nighty-tidy-2-provenance.json`).

| NT2 owns                                                                                                         | NT2 reuses (shared)                                                                         |
| ---------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------- |
| Panel (native React, desktop + mobile) `features/nightytidy2/`                                                   | `~/.claude/audit-types/<slug>.md` specs + the `codebase-audit` recipe + orchestration skill |
| IPC `NIGHTY_TIDY_2_*` + handlers + service `nightyTidy2Service`                                                  | the apply engine `nighty-tidy-apply-spawn.ts` (`applying-audit-findings` skill)             |
| `nighty_tidy_2_runs` + `nighty_tidy_2_subscriptions` + queries                                                   | `<repo>/audit-reports/audit-<slug>/` report layout + resolver                               |
| The stateless cadence scheduler + the structured-findings reader                                                 | cron parse (`getNextRun`/`getPreviousRun`, `presetToCron`)                                  |
| Offered-audit list `resources/nighty-tidy-2/manifest.json`                                                       | the 24h runs-watchdog pattern                                                               |
| Operator session-host project `__nightytidy2_agent__` (sentinel + spawnable + resolver workdir + hide + seed)    | the session-host shell (`useProjectSessionHost` + `SessionHostSidebar` + `SessionRow`)      |
| Operator CLI control API `/nighty-tidy-2/*` + `nighty_tidy_2.run_now` approval handler + standing-context primer | the run-now approval family (`requireApprovalForCliNightyTidyRun`, shared with NT1)         |

### How to use it

1. **It's on by default** — Nighty Tidy ships enabled in the sidebar's **Plugins** section; toggle it any time at Settings → Plugins → **Nighty Tidy**, where an **Open** button also takes you straight to it.
2. **Run an audit** — pick a project; choose **Read-only** (just the report) or
   **Read-write** + a severity (Critical / Critical+High / +Medium / All); tick
   audits — or use the **Select all / Deselect all** toggle by the Audits heading to
   pick (or clear) the whole list in one tap; optionally type **custom instructions** to
   scope every audit in the run, and optionally choose a non-default **engine + model**;
   **Start Run**. Each audit is a regular
   session (`source: 'nighty_tidy_2'`, `Nighty Tidy: <Title>`), visible +
   interruptible — and when it finishes it **stays in your Needs-You inbox** so you can
   read the result. A read-write run, after a
   findings audit, spawns a **separate** fix session (`applying-audit-findings`) at
   the chosen severity; a fully-green fix run **auto-lands** into your main branch (a failed or
   dependency-touching run is held on its branch for review with an inbox alert).

   Choose **Abuse & Adversarial Misuse** when you want Nighty Tidy to think like a
   malicious but legitimate user: it traces how valid features, identities, incentives,
   lifecycle transitions, automation, and feature combinations could create unfair value,
   shifted cost, harassment, manipulation, resource exhaustion, or hard-to-detect harm. It
   is intentionally separate from **Security Sweep**, which looks for technical
   vulnerabilities such as authentication, authorization, injection, and secret exposure.
   The abuse audit is always read-only and requires a reachable, evidence-backed path before
   reporting a finding; it does not perform live exploitation or external actions.

   Choose **Efficiency** when you want to find the work your software repeats that nobody
   needed. It prices every finding as the cost of one run × how often it runs × how many
   copies run at once, so it catches what a line-by-line review cannot: a cheap call made every
   few seconds in every open session, a new process started for every small question, polling
   where a signal already exists, work redone although nothing changed, capacity left idle while
   work queues, and waste in build pipelines and AI calls. It works on any kind of software and
   measures with the repository's own tools where it can. It deliberately leaves slow
   algorithms, how long database queries take, what a database stores and the work it does, rendering, cloud sizing, test speed,
   perceived speed and memory leaks to **Performance**, **Database Waste**, **Cost & Resource
   Optimization**, **Test Efficiency**, **Perceived Performance** and **Memory Leak** — run those
   alongside it for full coverage. Every Efficiency report lists all 35 optimization categories,
   with the audit that owns each.

   Choose **Database Waste** when you want to find what your databases keep, write and are made
   to do that nobody needs: data no code ever reads, the same facts stored twice, oversized
   values like whole transcripts or files kept in a column, history kept forever, deleted rows
   that are never purged, indexes that slow every save but serve no query, saves that change
   nothing, backups or copies that pile up — and the work the database burns for nothing: reads
   whose result no one uses, and statements that sift through hundreds of thousands of rows to
   hand back a handful. It measures storage and work read-only first and prices every finding in
   space kept, growth per day, disk reads, disk writes and CPU per day, and it only calls
   something unused after proving nothing reads it. Its fixes never delete data a person
   created — only rebuildable, duplicate, expired, diagnostic or already-deleted data; anything
   else comes back as a decision for you, with what keeping it costs.

   It deliberately overlaps **Performance**: a query can finish in twenty milliseconds and still
   be wasteful, so Performance judges how long queries take and Database Waste judges what they
   made the database do — rows examined, disk reads, CPU burned. Either may report such a
   statement and each names the other, so running both gives you both halves rather than a
   duplicate.

   Choose **Systematization** (Architecture) when you want to know which jobs your codebase
   does piecemeal in many places with no owner — switches, timers, process launches, retries,
   logs, registries, permission checks — and which of them should become one named system: a
   single entry point, a complete list, and a check that keeps new code on it. It needs three or
   more independent copies and evidence the scattering has already cost something before it
   recommends anything, and "leave it scattered" is a normal answer. It is advice only: it never
   changes code, and its findings are never applied automatically.

   Choose **Guardrail Integrity** (Architecture) when you want to know whether the rules your
   project relies on are really enforced. It pairs every written rule with the check meant to
   enforce it, and flags rules no check covers plus checks that cannot fail — a pattern that
   matches no files, a block that only warns or is switched off, a skip that reports as a pass.
   Advice only, like every audit in this group.

   Choose **Parallel Change Friction** (Architecture) when branches or agents working at the same
   time keep colliding. It reads the repository's merge history, finds the files and data shapes
   every change must edit — shared lists, numbering that every insert shifts, one file holding
   everything — and recommends the shape that would let each change merge cleanly. A repository
   with no parallel work gets "no parallel friction".

   Choose **Unfinished Migrations** (Architecture) when the code seems to do some jobs two ways
   because a move from an old library, API, pattern or storage format stalled. It counts each
   side, checks from history whether the move is still happening, and recommends finishing it,
   freezing it with a check, or rolling it back.

   Choose **Domain Language** (Architecture) when the same thing goes by different names on
   screen, in the docs and in the code, or one name means two things. It maps every concept to its
   names and recommends one name per concept plus a glossary — and it never renames stored data.

3. **Automate** (optional) — click **"+ New automation"** on the Automate screen to open the
   **step-by-step wizard**: pick the **repo**, choose the **audits** (search + Select-all / Clear),
   add optional **instructions** (free-text scope added to every audit — e.g. "only the checkout
   flow"), pick a cheaper **model**, set the **schedule** (a **day-of-week picker** — tap the days — plus a
   time, with an honest next-run preview), choose the **auto-fix** level (off / critical / critical +
   high / everything — **off by default**), tune the optional **Advanced** page (per-automation
   pacing caps, a "Start from" audit, a "Stop after N runs" lifetime cap), give it a **Name**, and
   **Create**. Each automation then shows as its **own row under its repo** with its own On/Off
   toggle; to tweak one, click its row and open any picker — changes apply the moment you make them,
   no Save button. Add as **many automations per repo** as you like (one per feature), and remove one
   with **"Delete this automation"** (a confirm, plus Undo). The global
   **concurrency cap** and the nightly spend/count warning limits live in the **"Global pacing &
   limits"** editor (opened from the page header or the footer's _Limits_ link). Everything
   live-reloads when ANY writer saves/deletes an automation — including an operator session over
   the CLI API (`NIGHTY_TIDY_2_SUBSCRIPTIONS_UPDATED` push). Over that CLI an operator can change
   just ONE field of an existing automation — e.g. move the nightly audit to 8 PM — with a partial
   `PATCH /nighty-tidy-2/subscriptions/:id`, no need to resend the whole config (omitted fields are
   kept, run history untouched, no restart).
4. **Read findings** — History → a run → the **Findings** screen: severity-grouped,
   filterable, expandable cards. Reports also persist to
   `<project>/audit-reports/audit-<slug>/report.md`.
   A **`subsystem-map`** finding additionally offers **"Create an automation…"**, which opens
   the automation wizard PRE-FILLED with that subsystem's name, scope paths, recommended
   audits and cadence — that audit names one subsystem per finding and states the scope it
   should be run against, so the action just saves retyping it. It only opens a draft: nothing
   is created or enabled until you finish the wizard, because a scheduled run costs money on
   every fire. Findings from other audit types carry no recipe and show no button.
5. **Write your own audit** (optional) — on the **Custom audits** screen, **New custom
   audit** opens an editor: give it a name, a category, an optional one-line description,
   and a plain-language **"what should this audit look for?"**. Save it and it appears in
   the Run and Automate pickers next to the built-ins (badged **Custom**), running
   exactly like them — read-only or auto-fix, one-off or scheduled. You only describe WHAT
   to look for; the audit engine supplies the multi-phase sweep + the findings list, so a
   good custom audit can be a few sentences. A custom audit is offered for **every repo** by
   default, like the built-ins — or set **Runs on** to a single project and it is offered for
   that one only, hidden from the Run and Automate pickers everywhere else (and refused if
   something still tries to run it elsewhere, so an automation saved before you scoped it
   cannot quietly fire on the wrong repo). Scoped or not, it always stays on the **Custom
   audits** screen so you can edit or delete it — including after you remove the project it
   was scoped to. Edit or delete your own any time (the built-ins stay read-only). The list
   live-reloads when any writer changes it — including an operator over the CLI
   (`NIGHTY_TIDY_2_CUSTOM_AUDITS_UPDATED` push).

6. **Send a run to a cloud machine** (optional) — an audit can run on a rented cloud machine
   instead of your own computer. Open **Run location** on the Run screen (or **Run location** in
   an automation's **Advanced** panel and its wizard) and pick **Cloud machine**; leave it alone
   and nothing changes, because **Local is the default**. A cloud run goes through the same one
   gate every other cloud launch goes through, so the cloud kill switch, the one-time spend
   consent, the billing check, the clock check and the engine check all still apply — and if the
   cloud cannot take the run it **stops with a plain reason on its row rather than quietly running
   here**. When it finishes, its report and findings come back to the same place a local run's do,
   and History reads the same. Two things worth knowing: a cloud automation **skips its auto-fix
   step** (the fix session would be a code-editing session on your own computer, which is not what
   "run in the cloud" should mean), and a cloud audit's session is ended once its report is home so
   the machine is not woken a second time to collect work that is already here. Cloud runs are
   offered only when cloud sessions are switched on for the install; the choice is hidden
   otherwise, never shown and then refused.

## Related

- [Nighty Tidy part 2](nighty-tidy-2-part-2.md) — how it works underneath, the auto-tagged sessions, driving it with an AI operator, and the boundaries worth knowing.
- [night-shift.md](night-shift.md) — the general overnight sequencer this panel runs sit on.
- [audit-framework.md](audit-framework.md) — the audit machinery an overnight audit run uses.

### Related

- The original (legacy) Nighty Tidy (`nightytidy`) has been retired and removed. NT2 reuses
  the shared audit framework + apply engine it left behind; everything user-facing is forked.
