Omniscio documentation
Browse all documentation
  1. Getting Started13
  2. Sessions & Agents115
  3. Inbox & Notifications59
  4. Projects & Tasks95
  5. Automation & Scheduling75
  6. Knowledge & Memory26
  7. AI Features60
  8. Integrations100
  9. Plugins & Marketplace33
  10. Cloud & Teams56
  11. Settings & Customization58
  12. Account & Billing28
  13. Troubleshooting84
  14. CLI & API Reference22
  15. Legal & Policies4
  16. Uncategorised22

Crash recovery (auto-resume sessions after a crash or restart) (part 2)

Part 2 of the Crash recovery page: the crash-loop circuit breaker that quarantines a session, the silent-death watchdog that alerts you the moment the app is killed, the states that are never auto-resumed, why revivals trickle, and where each piece lives in the code.

What it is

This is part 2 of the Crash recovery (auto-resume sessions after a crash or restart) page. It covers what happens when recovery itself goes wrong — the breaker that gives up on a session which keeps crashing the app, the watchdog that tells you the instant the app is killed, the states that are deliberately never resumed, and why sessions come back one at a time — plus where all of it lives in the code.

Where to find it

The surfaces are the ones on the parent page: the sidebar, the resume toast and the Manage sessions modal. Nothing described here has a screen of its own — the alert it describes arrives as an email and, when a phone is paired, as a page on that phone.

How it behaves

When Omniscio gives up on a session (the crash-loop circuit breaker)

There's a safety limit on auto-resume itself. If a session's revival keeps crashing the app — it gets resumed on boot, the app dies before that session ever finishes a turn, and the next boot tries it again — Omniscio stops trying after 3 consecutive failed boots (within a ~15-minute window). Instead of resuming it a fourth time it quarantines the session: parks it in the orange recovery-failed state in your inbox with a message — "Auto-resume gave up after N attempts — this session kept crashing on resume. Open it and send a message to resume manually." — and lets the app boot cleanly.

This exists because one badly-behaved session (a huge history that locks the renderer, a resume that stalls the main thread until the OS kills the app) used to be able to brick every launch: each boot re-resumed it, each resume re-crashed the app, forever, with no backoff. The breaker turns that fatal loop into a one-session graceful degrade — the rest of your sessions come back normally and the poison one waits for you. The count only climbs while a session keeps getting resumed and re-crashing without ever finishing a turn. Two things clear it: it self-heals after ~15 minutes with no further failed resume, and a clean quit clears it immediately — a normal shutdown (even the interrupted quit → relaunch → quit → relaunch dance) is proof the app wasn't crash-looping, so those sessions simply come back instead of being mistaken for a poison one. So a transient hiccup, or a burst of quick relaunches, never strands a session — only a genuinely poison one that keeps crashing the app trips it.

The app-crasher is the only recovery-failed state held for a manual restart. Every other "recovery-failed" session — one that gave up for a non-crashing reason (all your logins were out of capacity, an aborted reply hit its daily retry cap, a stuck turn exhausted its retries, a transient API error — a brief 5xx server hiccup — outlasted its ~7-minute silent-retry window) — is no longer a dead end: a background safety net keeps re-driving it forever on a gently-growing interval (seconds, then minutes, capped around half an hour) until it comes back, so nothing is ever silently abandoned. Only the session that keeps crashing the app is excluded from that automatic retry (re-driving it would just re-brick the launch) — it waits in the inbox for you to restart it by hand, or for a one-click Restart stuck sessions (bulk restart). The full re-arm rules live in the session self-heal contract.

A boot that dies before it gets to your sessions never costs a strike either. The three strikes are meant to catch a session whose own revival kills the app — so a strike only counts once Omniscio has actually started bringing that session back. Revivals are deliberately gentle (one at a time, seconds apart), so if the app dies early in startup for some unrelated reason, most of the queue was never even reached. Those sessions get their strike handed back on the next launch and come back normally. Before this, three quick failed startups in a row could quarantine every session that had been running — sessions that were perfectly healthy and had never been tried once. (2026-09-06: 48 of them.)

If any session doesn't get restarted, you get told. After every launch Omniscio checks whether it started back up everything it meant to. If some sessions were never retried, it raises a single inbox card — "Some sessions have not been resumed yet" — with a Resume them now button. Clicking it asks you to confirm first (each resume uses a turn of that session's budget) and shows the live count, since the sweep may have already brought some back on its own. If everything came back, no card appears.

A missing project folder never costs a strike. If a session's folder isn't there when recovery reaches it — the drive is unplugged, the folder was moved or renamed — a resume can only fail, so Omniscio doesn't attempt it and doesn't count it: the session is skipped without burning any of its 3 quarantine attempts, and the red "Project folder is missing" banner tells you why (see project folder missing). Recovery picks the session up again by itself once the folder exists — the background safety net re-checks roughly every 90 seconds, and the next app start re-checks too — so reconnecting the drive is all it takes.

Instant alert when the app itself is killed (the silent-death watchdog)

Session recovery only runs at the next launch — so if the whole app is killed without a clean shutdown (a stray terminal-window close, an external taskkill, anything that terminates the process tree instantly), nothing above can tell you until you happen to relaunch. Since 2026-07-29 a tiny detached silent-death watchdog closes that gap: the app arms it at startup, it watches the app's 5-second heartbeat file from outside, and if the heartbeat goes silent for ~60 seconds AND the app's process is genuinely gone, it immediately sends the crash email and pages your phone (if one is paired) — typically within 60–75 seconds of the death, even though the app itself is dead. It is alert-only by design: it never restarts the app for you. A frozen-but-alive app is deliberately NOT treated as dead (no false pages during a system-wide freeze), a clean quit makes the watchdog exit silently, and the post-mortem email at your next launch notes when the instant alert already fired so you're never double-alarmed. Kill switch for power users: the AMC_DISABLE_SILENT_DEATH_WATCHDOG=1 environment variable. The alert works in an installed build as well as one run from source, so a silent death is paged within about a minute wherever Omniscio is running. Invariants: silent-death-watchdog contract.

What does NOT get auto-resumed

  • Sessions you explicitly archived or paused — those terminal states are respected. Omniscio won't unarchive on your behalf.
  • Sessions that errored on their own (a turn failure, a dead CLI process, a non-transient API error such as a context-window 400) — these are NOT crash victims (crash_reconciled=0), so they stay error (red) across restart, exactly as they ended, no matter how recently they erred. (A transient API 5xx that outlasts its silent retries is the exception — it lands in recovery-failed (auto-recovering, the never-fail family above), not dead red.) The crash path (listCrashRecoverableSessions) only revives genuine app-crash victims; a session preserved in the graceful-shutdown snapshot — including one force-closed into error by a busy shutdown (above) — still resumes on the next launch regardless of how long the app was down.
  • Sessions whose CLI process exited cleanly (status ended with ended_at populated) — they finished. There's nothing to resume.
  • Worktrees stuck in intermediate states — these are reconciled separately by reconcileWorktreeStates at startup (any merging / active worktrees that didn't finish get marked failed), but no session re-launch comes out of that path.
  • Sessions Omniscio has quarantined after repeated resume crashes — see the crash-loop circuit breaker above. They wait in recovery_failed (red — a failure — in your inbox) for a manual message instead of being resumed a fourth time.
  • Everything, when a second copy of Omniscio is already running on the same data — if, at startup, Omniscio detects that another live instance already owns this database (for example a test/sandbox build that got pointed at your real data, or a second launch that didn't fully take over), it skips auto-resume entirely rather than re-parking and re-spawning sessions the other copy is already running. Your sessions are left exactly as they are — open one and send a message to continue it. This is a safety backstop against two copies fighting over the same sessions; a normal single-instance restart is unaffected (the previous copy is gone, so nothing is "already running").

Why the trickle

The design exists because the alternative was bad. Before the queue (incident 2026-05-24), Omniscio tried to revive every recoverable session in parallel, before the window was created. With 13 sessions queued, the user's machine froze for 30+ seconds at every launch — no paint, no input, while every CLI child started spawning simultaneously and every keep-alive panel cold-mounted its full conversation history at the same time. The fix splits recovery into two phases: a fast DB-only phase that paints the sidebar immediately, and a slow spawn phase that runs after the window is interactive, draining one session at a time with a gentle, load-aware gap between each (a 5-second floor that stretches while the machine is still busy — its CPU, its memory, or the disk — capped at 20 seconds). The result: you never see a frozen launch, no matter how many sessions are queued. A later refinement (2026-06-02) makes that gap weight-aware: a session revived by replaying its full conversation history — the heaviest kind of restart — gets a larger gap held to a calmer CPU level, so a cluster of those can't stack up and peg the machine. That stacking was the cause of a ~29-second main-thread stall observed during a 20-session restart recovery. A further refinement (2026-07-12) makes the gap disk-aware too: on a machine where the database and the working folders share one physical drive, each revival's reload can saturate that disk while the CPU sits idle — so the gap now also holds while the disk is storming, which a CPU-and-memory-only check couldn't see. That blind spot was the cause of a ~50-second whole-app freeze during a 56-session restart.

For agents

Where it lives in code

For agents working in the repo:

  • Startup entry points: performCrashRecovery(), performRestartRecovery(), and performSuspendedSessionResume() (the interrupted-on-close path) in src/main/index.ts — all gated by autoResumeOnCrash, all feed the same recoveryQueue. The surfacing half (the close-time flip to needs_you + pendingAction 'suspended') lives in shutdown-disposition.ts + shutdown-session-surface.ts (called from shutdown.ts); the resume half + its listResumableSuspendedSessions() gate (the Claude-binary providers — claude + the anthropic-compat vendors deepseek/kimi, via usesClaudeBinary; has a cli_session_id; re-checks agentFinishedLastTurn) are in suspended-resume.ts. Invariants: the needs-you-visibility contract (new producer) + bulk-session-actions contract (Restart target).
  • DB-ownership guard (before every door): all five doors additionally gate on runtime.foreignDbOwner, captured once in the single-instance gate by captureAndClaimDbOwnership() (db-owner-marker.ts). It stamps a db-owner.json marker beside the DB — written in ALL modes, unlike the packaged-only running.pid — and reads the PRIOR marker: a still-alive foreign pid means a second instance owns this database, so recovery is skipped rather than trampling its live sessions (the 2026-07-16 case where a sandbox that inherited DATA_DIR re-parked 58 real sessions; the launcher fix in sandbox-shared.cjs buildSandboxChildEnv sets the sandbox's own isolated DATA_DIR so it never opens the live DB). Keyed on a live foreign pid at OUR data dir, never on DATA_DIR presence; fail-open. Invariants in the crash-recovery give-up contract § Instance-ownership guard.
  • External mirror (codex/gemini/pi/opencode/hermes + cursor/antigravity): the SAME interrupted-on-close path for the external engines. Surfacing is disposeExternalShutdownSessions() (same shutdown-session-surface.ts) — it classifies each session like the Claude path, then for a mid-work one flips it to needs_you+'suspended' AND releases it from its manager (releaseSession on whichever external base it extends — BaseExternalSessionManager for persistent (codex/gemini/pi/opencode/hermes/kimicode), OneShotExternalSessionManager for one-shot (cursor/antigravity), which sets a released flag so the aborting in-flight turn's conclusion can't re-write status — stop the turn through the engine's own abort, finalize the partial, charge the spend the engine already counted but had not written (codex's usage past its watermarks, pi's pending totals, opencode's completed steps — exactly once, money contract M3c) + drop from the map, no status write) so the shutdown terminateAll() can't clobber the surface back to ended/gray. Resume is performExternalSuspendedResume() + its listResumableSuspendedExternalSessions() gate (the registry-derived SUSPEND_RESUMABLE_EXTERNAL_PROVIDER_IDS = (persistent-external ∨ one-shot-cli) ∧ restart-resumable, i.e. crash-resumable MINUS openclaw; NO cli_session_id gate — external replays from the DB), resuming via getCrashResumableExternalManager(provider).resumeAfterCrash and marking 'starting' before enqueue so a resume-crash degrades to performExternalCrashRecovery's give-up (no boot-loop). Invariants I7/I8 in the needs-you-visibility contract.
  • Recovery queue: src/main/services/recovery-queue.ts — a dispatch-per-gap drip-feed with an adaptive load-aware gap (INTER_SPAWN_MIN_DELAY_MS = 5000 floor, extended in CPU_SETTLE_POLL_MS steps while whole-system busy% ≥ CPU_BUSY_HIGH_WATERMARK, RAM ≥ MEMORY_HIGH_WATERMARK_PERCENT (85%), OR the shared disk is under a page-in storm (isBoxUnderDiskPressure() — the 2026-07-12 disk#5-freeze fix, disk-pressure-holds-the-gap-capped; the floor warms its rolling baseline via rollingPageReadRate() so the first settle read isn't a cold 0), capped at INTER_SPAWN_MAX_DELAY_MS = 20000; the gap is chosen at dispatch, for a launch that was ATTEMPTED — the-settle-is-chosen-at-dispatch, since the async probe means the outcome is not known yet). Any gap adjacent to a heavy transcript-replay launch uses a stricter profile — HEAVY_INTER_SPAWN_MIN_DELAY_MS = 10000 floor, HEAVY_CPU_BUSY_WATERMARK = 50, HEAVY_INTER_SPAWN_MAX_DELAY_MS = 45000 — with weight classified by classifyLaunchWeight() from getSessionHistoryCharCount() (restart items only; crash items are always light). a 30-second per-launch liveness probe on the child, run ASYNC so the next launch never waits on it (concurrency is bounded by the spawn gate's own recovery lane — concurrency-is-the-spawn-gates-recovery-lane; a probe that finds no registered child re-queues its session once, WITH a line, then parks it recoverably — a-probe-that-finds-no-child-is-bounded), a progress-aware STALL_FLUSH_MS (~2 min no-forward-progress) stall backstop that replaced the old fixed 15-minute ceiling — re-armed on every item the worker reaches, so a healthy slow drain is never guillotined and only a genuine wedge trips it (a no-forward-progress that a main-loop freeze explains DEFERS instead of flushing, so a heavy startup that briefly freezes doesn't strand its healthy tail — a-freeze-explained-stall-defers) — and a 10-minute "queue never started" safety net (gated on isNeverStarted() — fires ONLY when start() never ran; an actively-draining queue is left to the stall backstop). A flushed item is parked recoverable — needs_you+recovery_failed with a SESSION_STATUS_CHANGED push — so it shows the amber "your turn" dot, never dead red nor stale gray. The pacing + lifecycle invariants (heavy-launch heavy-gap-on-either-side, the dispatch-time settle the-settle-is-chosen-at-dispatch, the outcome ledger one-outcome-per-dequeue, never-started gating is-never-started-is-pre-start-only, recoverable-parking a-flushed-item-is-parked-recoverable, the stall backstop progress-aware-stall-backstop, the freeze-aware defer a-freeze-explained-stall-defers, disk-aware pacing disk-pressure-holds-the-gap-capped) are locked in the recovery-queue contract and its pacing shard. Queue-jump: expediteRecovery(sid) synchronously splices a still-queued item out (find+splice can't race the single worker) and runs the worker's skip-path cleanup, returning true so the caller launches it now — false when the worker is already mid-launch on it. The send guard (session-service.ts) and restart (session-relaunch-service.ts) call it so an explicit user action jumps the line; the behavior is locked by invariant I9 in the optimistic-running-dot contract. Mobile seed: web-bootstrap-payload.ts seeds a phone's recovery-pending set from getQueuedRecoverySessionIds() (the queue's still-waiting items — durable) unioned with the TTL-limited visual mirror, so the "Restart / Continue" button + the Starting row still show on a phone that opens the app more than 2 minutes into a restart (a big restart drains far slower than the 2-minute optimistic-dot TTL that would otherwise leave the phone's seed empty) — invariant I5.
  • Crash-victim gate: listCrashRecoverableSessions() requires crash_reconciled = 1 — set ONLY by reconcileStaleSessionStatuses() on the running/starting/stalled → error flip, and cleared by every other status write (updateSessionStatus + unarchiveSession) — so a genuinely-errored session (marker 0) stays red across restart. The 120-second status_changed_at window is kept as a freshness bound, not the discriminator. Invariants in the crash-recovery restart-persistence contract.
  • Per-session give-up: performCrashRecovery bumps a durable crash_recovery_attempts counter (the crash_recovery_* columns) before each enqueue and quarantines a session at CRASH_RECOVERY_ATTEMPT_LIMIT = 3 (within CRASH_RECOVERY_ATTEMPT_WINDOW_MS = 15 min). Two things reset the counter: the 15-min stale-gap window, and a clean shutdown — resetCrashRecoveryAttemptsIfLastShutdownClean at startup zeroes it when the lastShutdownWasClean flag (set in gracefulShutdown) is true, so an interrupted quit/relaunch burst can't accumulate the counter and wrongly quarantine healthy sessions. It is NOT reset from a turn handler (an early handleResultEvent reset draft was reverted — it broke 100+ under-mocked turn tests). Invariants are locked in the crash-recovery give-up contract.
  • Two-phase split: Fast phase runs synchronously at startup (DB only, no spawns); slow phase gates recoveryQueue.start(onDrained) on the main window's renderer being hydrated — awaitRendererHydrated(120s ceiling) (renderer-hydrated-signal.ts, signalled from the main window's 'almost-ready' boot beat when its four boot fetches complete; sender-scoped to the main window; kill switch AMC_DISABLE_BOOT_HYDRATED_GATE=1), composed with the orphan-reaper gate — so the UI is genuinely usable before respawns begin. (History: replaced the 2026-07-06→08 first-paint gate, whose 3 s ceiling always fell through on a contended boot — paint ≠ usable; that in turn replaced the former AGG-021 setTimeout(500) approximation. See startup-order-contract I9.)
  • Toast hook: handleSessionRecoveryComplete in src/renderer/src/hooks/useAppSessionAlerts.ts — on mount, PULLS the buffered count via the SESSION_RECOVERY_CONSUME_PENDING invoke (read-and-clear of a main-side buffer, pending-recovery-summary.ts) and fires the info toast. Pull, not a timed push: recovery runs pre-window, so a setTimeout(2000) push raced the renderer's mount and was lost on a slow machine (the former AGG-021 gate).
  • pendingSessionIds watchdog exemption: The liveness watchdog's Phase 2 reconciler marks any session that's starting in the DB but absent from both the in-memory sessions Map and the pendingSessionIds Set as error. The recovery queue registers every enqueued item via markSessionPending(sid) and releases it via unmarkSessionPending(sid) in the worker's finally block — without this, the watchdog would torch the queue mid-drain.
  • The strike is provisional until dispatch: incrementCrashRecoveryAttempt claims it, settleCrashRecoveryAttempt earns it in the recovery-queue worker immediately before launchFn (plus the cloud re-wake that bypasses the queue), and refundUndispatchedCrashRecoveryAttempts hands back every un-dispatched claim at boot, before any door. strike-settled-on-dispatch in the give-up contract.
  • The per-boot census: each door calls noteResumeDoorPass(considered, dispatched); the drain logs one resume owed … still owed … line via resume-census.ts and raises the still-owed card. stillOwed is a subtraction, never a count of failure rows. G12.
  • markRecoveryComplete() timing: Lives inside the queue's onDrained callback, NOT inline after enqueue. The watchdog's Phase 2 reconciler stays gated until drain.

See the crash-recovery-startup-freeze-queue postmortem for the full incident write-up, and process-management-lifecycle.md for the surrounding lifecycle context (graceful shutdown sequence, the three race-guard fixes, stall watchdog, liveness watchdog).

Related

The parent page covers what recovery looks like from your side — which sessions come back, the trickle, and how to stop it. Dead logins and what to do when everything is out of capacity are on the account pool page; a session the app refuses to revive can be restarted by hand from bulk stop and restart; and a session whose project folder has gone missing is covered in project folder missing.

Last verified 2026-10-06