---
title: Crash recovery (auto-resume sessions after a crash or restart) (part 2)
---

# Crash recovery (auto-resume sessions after a crash or restart) (part 2)

## What it is

This is part 2 of the [Crash recovery (auto-resume sessions after a crash or restart)](crash-recovery.md) page. It covers what happens when recovery itself goes wrong — the breaker that gives up on a session which keeps crashing the app, the watchdog that tells you the instant the app is killed, the states that are deliberately never resumed, and why sessions come back one at a time — plus where all of it lives in the code.

## Where to find it

The surfaces are the ones on the [parent page](crash-recovery.md): the sidebar, the resume toast and the Manage sessions modal. Nothing described here has a screen of its own — the alert it describes arrives as an email and, when a phone is paired, as a page on that phone.

## How it behaves

### When Omniscio gives up on a session (the crash-loop circuit breaker)

There's a safety limit on auto-resume itself. If a session's revival keeps **crashing the app** — it gets resumed on boot, the app dies before that session ever finishes a turn, and the next boot tries it again — Omniscio stops trying after **3 consecutive failed boots** (within a ~15-minute window). Instead of resuming it a fourth time it **quarantines** the session: parks it in the orange **recovery-failed** state in your inbox with a message — _"Auto-resume gave up after N attempts — this session kept crashing on resume. Open it and send a message to resume manually."_ — and lets the app boot cleanly.

This exists because one badly-behaved session (a huge history that locks the renderer, a resume that stalls the main thread until the OS kills the app) used to be able to **brick every launch**: each boot re-resumed it, each resume re-crashed the app, forever, with no backoff. The breaker turns that fatal loop into a one-session graceful degrade — the rest of your sessions come back normally and the poison one waits for you. The count only climbs while a session keeps getting resumed and re-crashing without ever finishing a turn. Two things clear it: it self-heals after ~15 minutes with no further failed resume, **and a clean quit clears it immediately** — a normal shutdown (even the interrupted _quit → relaunch → quit → relaunch_ dance) is proof the app wasn't crash-looping, so those sessions simply come back instead of being mistaken for a poison one. So a transient hiccup, or a burst of quick relaunches, never strands a session — only a genuinely poison one that keeps _crashing the app_ trips it.

**The app-crasher is the _only_ recovery-failed state held for a manual restart.** Every _other_ "recovery-failed" session — one that gave up for a non-crashing reason (all your logins were out of capacity, an aborted reply hit its daily retry cap, a stuck turn exhausted its retries, **a transient API error — a brief 5xx server hiccup — outlasted its ~7-minute silent-retry window**) — is **no longer a dead end**: a background safety net keeps re-driving it **forever** on a gently-growing interval (seconds, then minutes, capped around half an hour) until it comes back, so nothing is ever silently abandoned. Only the session that keeps _crashing the app_ is excluded from that automatic retry (re-driving it would just re-brick the launch) — it waits in the inbox for you to restart it by hand, or for a one-click **Restart stuck sessions** ([bulk restart](bulk-stop-restart-sessions.md)). The full re-arm rules live in the [session self-heal contract](../../.claude/memory/contracts/transient-spawn-retry-contract.md).

**A boot that dies before it gets to your sessions never costs a strike either.** The three
strikes are meant to catch a session whose _own_ revival kills the app — so a strike only
counts once Omniscio has actually started bringing that session back. Revivals are
deliberately gentle (one at a time, seconds apart), so if the app dies early in startup for
some unrelated reason, most of the queue was never even reached. Those sessions get their
strike handed back on the next launch and come back normally. Before this, three quick
failed startups in a row could quarantine every session that had been running — sessions
that were perfectly healthy and had never been tried once. (2026-09-06: 48 of them.)

**If any session doesn't get restarted, you get told.** After every launch Omniscio checks
whether it started back up everything it meant to. If some sessions were never retried, it
raises a single inbox card — _"Some sessions have not been resumed yet"_ — with a **Resume
them now** button. Clicking it asks you to confirm first (each resume uses a turn of that
session's budget) and shows the live count, since the sweep may have already brought some
back on its own. If everything came back, no card appears.

**A missing project folder never costs a strike.** If a session's folder isn't there when recovery reaches it — the drive is unplugged, the folder was moved or renamed — a resume can only fail, so Omniscio doesn't attempt it and doesn't count it: the session is skipped without burning any of its 3 quarantine attempts, and the red "Project folder is missing" banner tells you why (see [project folder missing](project-folder-missing.md)). Recovery picks the session up again by itself once the folder exists — the background safety net re-checks roughly every 90 seconds, and the next app start re-checks too — so reconnecting the drive is all it takes.

### Instant alert when the app itself is killed (the silent-death watchdog)

Session recovery only runs at the **next** launch — so if the whole app is killed without a clean shutdown (a stray terminal-window close, an external `taskkill`, anything that terminates the process tree instantly), nothing above can tell you until you happen to relaunch. Since 2026-07-29 a tiny detached **silent-death watchdog** closes that gap: the app arms it at startup, it watches the app's 5-second heartbeat file from outside, and if the heartbeat goes silent for ~60 seconds AND the app's process is genuinely gone, it immediately sends the crash email and pages your phone (if one is paired) — typically within 60–75 seconds of the death, even though the app itself is dead. It is **alert-only by design**: it never restarts the app for you. A frozen-but-alive app is deliberately NOT treated as dead (no false pages during a system-wide freeze), a clean quit makes the watchdog exit silently, and the post-mortem email at your next launch notes when the instant alert already fired so you're never double-alarmed. Kill switch for power users: the `AMC_DISABLE_SILENT_DEATH_WATCHDOG=1` environment variable. The alert works in an installed build as well as one run from source, so a silent death is paged within about a minute wherever Omniscio is running. Invariants: [silent-death-watchdog contract](../../.claude/memory/contracts/silent-death-watchdog-contract.md).

### What does NOT get auto-resumed

- **Sessions you explicitly archived or paused** — those terminal states are respected. Omniscio won't unarchive on your behalf.
- **Sessions that errored on their own** (a turn failure, a dead CLI process, a _non-transient_ API error such as a context-window 400) — these are NOT crash victims (`crash_reconciled=0`), so they stay `error` (red) across restart, exactly as they ended, no matter how recently they erred. (A _transient_ API 5xx that outlasts its silent retries is the exception — it lands in **recovery-failed** (auto-recovering, the never-fail family above), not dead red.) The crash path (`listCrashRecoverableSessions`) only revives genuine app-crash victims; a session preserved in the graceful-shutdown snapshot — including one force-closed into `error` by a busy shutdown (above) — still resumes on the next launch regardless of how long the app was down.
- **Sessions whose CLI process exited cleanly** (status `ended` with `ended_at` populated) — they finished. There's nothing to resume.
- **Worktrees stuck in intermediate states** — these are reconciled separately by `reconcileWorktreeStates` at startup (any `merging` / `active` worktrees that didn't finish get marked `failed`), but no session re-launch comes out of that path.
- **Sessions Omniscio has quarantined after repeated resume crashes** — see the crash-loop circuit breaker above. They wait in `recovery_failed` (red — a failure — in your inbox) for a manual message instead of being resumed a fourth time.
- **Everything, when a second copy of Omniscio is already running on the same data** — if, at startup, Omniscio detects that **another live instance already owns this database** (for example a test/sandbox build that got pointed at your real data, or a second launch that didn't fully take over), it **skips auto-resume entirely** rather than re-parking and re-spawning sessions the other copy is already running. Your sessions are left exactly as they are — open one and send a message to continue it. This is a safety backstop against two copies fighting over the same sessions; a normal single-instance restart is unaffected (the previous copy is gone, so nothing is "already running").

### Why the trickle

The design exists because the alternative was bad. Before the queue (incident 2026-05-24), Omniscio tried to revive every recoverable session **in parallel, before the window was created**. With 13 sessions queued, the user's machine froze for 30+ seconds at every launch — no paint, no input, while every CLI child started spawning simultaneously and every keep-alive panel cold-mounted its full conversation history at the same time. The fix splits recovery into two phases: a **fast** DB-only phase that paints the sidebar immediately, and a **slow** spawn phase that runs **after** the window is interactive, draining one session at a time with a gentle, load-aware gap between each (a 5-second floor that stretches while the machine is still busy — its CPU, its memory, or the disk — capped at 20 seconds). The result: you never see a frozen launch, no matter how many sessions are queued. A later refinement (2026-06-02) makes that gap **weight-aware**: a session revived by replaying its full conversation history — the heaviest kind of restart — gets a larger gap held to a calmer CPU level, so a cluster of those can't stack up and peg the machine. That stacking was the cause of a ~29-second main-thread stall observed during a 20-session restart recovery. A further refinement (2026-07-12) makes the gap **disk-aware** too: on a machine where the database and the working folders share one physical drive, each revival's reload can saturate that disk while the CPU sits idle — so the gap now also holds while the disk is storming, which a CPU-and-memory-only check couldn't see. That blind spot was the cause of a ~50-second whole-app freeze during a 56-session restart.

## For agents

### Where it lives in code

For agents working in the repo:

- **Startup entry points:** `performCrashRecovery()`, `performRestartRecovery()`, and `performSuspendedSessionResume()` (the interrupted-on-close path) in [src/main/index.ts](../../src/main/index.ts) — all gated by `autoResumeOnCrash`, all feed the same `recoveryQueue`. The **surfacing** half (the close-time flip to `needs_you` + pendingAction `'suspended'`) lives in [shutdown-disposition.ts](../../src/main/app/shutdown-disposition.ts) + [shutdown-session-surface.ts](../../src/main/app/shutdown-session-surface.ts) (called from [shutdown.ts](../../src/main/app/shutdown.ts)); the **resume** half + its `listResumableSuspendedSessions()` gate (the Claude-binary providers — claude + the anthropic-compat vendors deepseek/kimi, via `usesClaudeBinary`; has a `cli_session_id`; re-checks `agentFinishedLastTurn`) are in [suspended-resume.ts](../../src/main/app/crash-recovery/suspended-resume.ts). Invariants: the [needs-you-visibility contract](../../.claude/memory/contracts/needs-you-visibility-contract.md) (new producer) + [bulk-session-actions contract](../../.claude/memory/contracts/bulk-session-actions-contract.md) (Restart target).
- **DB-ownership guard (before every door):** all five doors additionally gate on `runtime.foreignDbOwner`, captured once in the single-instance gate by `captureAndClaimDbOwnership()` ([db-owner-marker.ts](../../src/main/app/db-owner-marker.ts)). It stamps a `db-owner.json` marker beside the DB — written in ALL modes, unlike the packaged-only `running.pid` — and reads the PRIOR marker: a still-alive foreign pid means a second instance owns this database, so recovery is skipped rather than trampling its live sessions (the 2026-07-16 case where a sandbox that inherited `DATA_DIR` re-parked 58 real sessions; the launcher fix in [sandbox-shared.cjs](../../scripts/sandbox-shared.cjs) `buildSandboxChildEnv` sets the sandbox's own isolated `DATA_DIR` so it never opens the live DB). Keyed on a live foreign pid at OUR data dir, never on `DATA_DIR` presence; fail-open. Invariants in the [crash-recovery give-up contract](../../.claude/memory/contracts/crash-recovery-giveup-contract.md) § Instance-ownership guard.
- **External mirror (codex/gemini/pi/opencode/hermes + cursor/antigravity):** the SAME interrupted-on-close path for the external engines. Surfacing is `disposeExternalShutdownSessions()` (same [shutdown-session-surface.ts](../../src/main/app/shutdown-session-surface.ts)) — it classifies each session like the Claude path, then for a mid-work one flips it to `needs_you`+`'suspended'` AND **releases** it from its manager (`releaseSession` on whichever external base it extends — `BaseExternalSessionManager` for persistent (codex/gemini/pi/opencode/hermes/kimicode), `OneShotExternalSessionManager` for one-shot (cursor/antigravity), which sets a `released` flag so the aborting in-flight turn's conclusion can't re-write status — stop the turn through the engine's own abort, finalize the partial, charge the spend the engine already counted but had not written (codex's usage past its watermarks, pi's pending totals, opencode's completed steps — exactly once, [money contract](../../.claude/memory/contracts/monetary-integer-twin-contract.md) M3c) + drop from the map, no status write) so the shutdown `terminateAll()` can't clobber the surface back to `ended`/gray. Resume is `performExternalSuspendedResume()` + its `listResumableSuspendedExternalSessions()` gate (the registry-derived `SUSPEND_RESUMABLE_EXTERNAL_PROVIDER_IDS` = (persistent-external ∨ one-shot-cli) ∧ restart-resumable, i.e. crash-resumable MINUS openclaw; NO `cli_session_id` gate — external replays from the DB), resuming via `getCrashResumableExternalManager(provider).resumeAfterCrash` and marking `'starting'` before enqueue so a resume-crash degrades to `performExternalCrashRecovery`'s give-up (no boot-loop). Invariants **I7/I8** in the [needs-you-visibility contract](../../.claude/memory/contracts/needs-you-visibility-contract.md).
- **Recovery queue:** [src/main/services/recovery-queue.ts](../../src/main/services/recovery/recovery-queue.ts) — a dispatch-per-gap drip-feed with an adaptive load-aware gap (`INTER_SPAWN_MIN_DELAY_MS = 5000` floor, extended in `CPU_SETTLE_POLL_MS` steps while whole-system busy% ≥ `CPU_BUSY_HIGH_WATERMARK`, RAM ≥ `MEMORY_HIGH_WATERMARK_PERCENT` (85%), OR the shared disk is under a page-in storm (`isBoxUnderDiskPressure()` — the 2026-07-12 disk#5-freeze fix, `disk-pressure-holds-the-gap-capped`; the floor warms its rolling baseline via `rollingPageReadRate()` so the first settle read isn't a cold 0), capped at `INTER_SPAWN_MAX_DELAY_MS = 20000`; the gap is chosen at dispatch, for a launch that was ATTEMPTED — `the-settle-is-chosen-at-dispatch`, since the async probe means the outcome is not known yet). Any gap adjacent to a **heavy** transcript-replay launch uses a stricter profile — `HEAVY_INTER_SPAWN_MIN_DELAY_MS = 10000` floor, `HEAVY_CPU_BUSY_WATERMARK = 50`, `HEAVY_INTER_SPAWN_MAX_DELAY_MS = 45000` — with weight classified by `classifyLaunchWeight()` from `getSessionHistoryCharCount()` (restart items only; crash items are always light). a 30-second per-launch **liveness probe on the child**, run ASYNC so the next launch never waits on it (concurrency is bounded by the spawn gate's own recovery lane — `concurrency-is-the-spawn-gates-recovery-lane`; a probe that finds no registered child re-queues its session once, WITH a line, then parks it recoverably — `a-probe-that-finds-no-child-is-bounded`), a progress-aware `STALL_FLUSH_MS` (~2 min no-forward-progress) **stall backstop** that replaced the old fixed 15-minute ceiling — re-armed on every item the worker reaches, so a healthy slow drain is never guillotined and only a genuine wedge trips it (a no-forward-progress that a main-loop **freeze** explains DEFERS instead of flushing, so a heavy startup that briefly freezes doesn't strand its healthy tail — `a-freeze-explained-stall-defers`) — and a 10-minute "queue never started" safety net (gated on `isNeverStarted()` — fires ONLY when `start()` never ran; an actively-draining queue is left to the stall backstop). A flushed item is **parked recoverable** — `needs_you`+`recovery_failed` with a `SESSION_STATUS_CHANGED` push — so it shows the amber "your turn" dot, never dead red nor stale gray. The pacing + lifecycle invariants (heavy-launch `heavy-gap-on-either-side`, the dispatch-time settle `the-settle-is-chosen-at-dispatch`, the outcome ledger `one-outcome-per-dequeue`, never-started gating `is-never-started-is-pre-start-only`, recoverable-parking `a-flushed-item-is-parked-recoverable`, the stall backstop `progress-aware-stall-backstop`, the freeze-aware defer `a-freeze-explained-stall-defers`, disk-aware pacing `disk-pressure-holds-the-gap-capped`) are locked in the [recovery-queue contract](../../.claude/memory/contracts/recovery-queue-contract.md) and its [pacing shard](../../.claude/memory/contracts/recovery-queue-invariants-pacing-contract.md). **Queue-jump:** `expediteRecovery(sid)` synchronously splices a still-queued item out (find+splice can't race the single worker) and runs the worker's skip-path cleanup, returning `true` so the caller launches it now — `false` when the worker is already mid-launch on it. The send guard ([session-service.ts](../../src/main/services/session/session-service.ts)) and restart ([session-relaunch-service.ts](../../src/main/services/session/session-relaunch-service.ts)) call it so an explicit user action jumps the line; the behavior is locked by invariant **I9** in the [optimistic-running-dot contract](../../.claude/memory/contracts/optimistic-running-dot-contract.md). **Mobile seed:** [web-bootstrap-payload.ts](../../src/main/services/web/web-bootstrap-payload.ts) seeds a phone's recovery-pending set from `getQueuedRecoverySessionIds()` (the queue's still-waiting items — durable) **unioned** with the TTL-limited visual mirror, so the "Restart / Continue" button + the Starting row still show on a phone that opens the app more than 2 minutes into a restart (a big restart drains far slower than the 2-minute optimistic-dot TTL that would otherwise leave the phone's seed empty) — invariant **I5**.
- **Crash-victim gate:** `listCrashRecoverableSessions()` requires `crash_reconciled = 1` — set ONLY by `reconcileStaleSessionStatuses()` on the `running`/`starting`/`stalled` → `error` flip, and cleared by every other status write (`updateSessionStatus` + `unarchiveSession`) — so a genuinely-errored session (marker 0) stays red across restart. The 120-second `status_changed_at` window is kept as a freshness bound, not the discriminator. Invariants in the [crash-recovery restart-persistence contract](../../.claude/memory/contracts/crash-recovery-restart-persistence-contract.md).
- **Per-session give-up:** `performCrashRecovery` bumps a durable `crash_recovery_attempts` counter (the `crash_recovery_*` columns) before each enqueue and quarantines a session at `CRASH_RECOVERY_ATTEMPT_LIMIT = 3` (within `CRASH_RECOVERY_ATTEMPT_WINDOW_MS = 15 min`). **Two things reset the counter:** the 15-min stale-gap window, and a **clean shutdown** — `resetCrashRecoveryAttemptsIfLastShutdownClean` at startup zeroes it when the `lastShutdownWasClean` flag (set in `gracefulShutdown`) is true, so an interrupted quit/relaunch burst can't accumulate the counter and wrongly quarantine healthy sessions. It is NOT reset from a turn handler (an early `handleResultEvent` reset draft was reverted — it broke 100+ under-mocked turn tests). Invariants are locked in the [crash-recovery give-up contract](../../.claude/memory/contracts/crash-recovery-giveup-contract.md).
- **Two-phase split:** Fast phase runs synchronously at startup (DB only, no spawns); slow phase gates `recoveryQueue.start(onDrained)` on the main window's renderer being **hydrated** — `awaitRendererHydrated(120s ceiling)` ([renderer-hydrated-signal.ts](../../src/main/app/renderer-hydrated-signal.ts), signalled from the main window's `'almost-ready'` boot beat when its four boot fetches complete; sender-scoped to the main window; kill switch `AMC_DISABLE_BOOT_HYDRATED_GATE=1`), composed with the orphan-reaper gate — so the UI is genuinely usable before respawns begin. (History: replaced the 2026-07-06→08 first-paint gate, whose 3 s ceiling always fell through on a contended boot — paint ≠ usable; that in turn replaced the former AGG-021 `setTimeout(500)` approximation. See [startup-order-contract](../../.claude/memory/contracts/startup-order-contract.md) `I9`.)
- **Toast hook:** `handleSessionRecoveryComplete` in [src/renderer/src/hooks/useAppSessionAlerts.ts](../../src/renderer/src/hooks/useAppSessionAlerts.ts) — on mount, PULLS the buffered count via the `SESSION_RECOVERY_CONSUME_PENDING` invoke (read-and-clear of a main-side buffer, [pending-recovery-summary.ts](../../src/main/app/pending-recovery-summary.ts)) and fires the info toast. Pull, not a timed push: recovery runs pre-window, so a `setTimeout(2000)` push raced the renderer's mount and was lost on a slow machine (the former AGG-021 gate).
- **`pendingSessionIds` watchdog exemption:** The liveness watchdog's Phase 2 reconciler marks any session that's `starting` in the DB but absent from both the in-memory `sessions` Map and the `pendingSessionIds` Set as `error`. The recovery queue registers every enqueued item via `markSessionPending(sid)` and releases it via `unmarkSessionPending(sid)` in the worker's `finally` block — without this, the watchdog would torch the queue mid-drain.
- **The strike is provisional until dispatch:** `incrementCrashRecoveryAttempt` claims it,
  `settleCrashRecoveryAttempt` earns it in the recovery-queue worker immediately before
  `launchFn` (plus the cloud re-wake that bypasses the queue), and
  `refundUndispatchedCrashRecoveryAttempts` hands back every un-dispatched claim at boot,
  before any door. `strike-settled-on-dispatch` in the [give-up contract](../../.claude/memory/contracts/crash-recovery-giveup-contract.md).
- **The per-boot census:** each door calls `noteResumeDoorPass(considered, dispatched)`;
  the drain logs one `resume owed … still owed …` line via
  [resume-census.ts](../../src/main/services/recovery/resume-census.ts) and raises the
  still-owed card. `stillOwed` is a subtraction, never a count of failure rows. G12.
- **`markRecoveryComplete()` timing:** Lives inside the queue's `onDrained` callback, NOT inline after enqueue. The watchdog's Phase 2 reconciler stays gated until drain.

See the [crash-recovery-startup-freeze-queue postmortem](../../.claude/memory/postmortems/crash-recovery-startup-freeze-queue-postmortem.md) for the full incident write-up, and [process-management-lifecycle.md](../../.claude/memory/process-management-lifecycle.md) for the surrounding lifecycle context (graceful shutdown sequence, the three race-guard fixes, stall watchdog, liveness watchdog).

## Related

The [parent page](crash-recovery.md) covers what recovery looks like from your side — which sessions come back, the trickle, and how to stop it. Dead logins and what to do when everything is out of capacity are on the [account pool](account-pool.md) page; a session the app refuses to revive can be restarted by hand from [bulk stop and restart](bulk-stop-restart-sessions.md); and a session whose project folder has gone missing is covered in [project folder missing](project-folder-missing.md).
