Omniscio documentation
Browse all documentation
  1. Getting Started13
  2. Sessions & Agents115
  3. Inbox & Notifications59
  4. Projects & Tasks95
  5. Automation & Scheduling75
  6. Knowledge & Memory26
  7. AI Features60
  8. Integrations100
  9. Plugins & Marketplace33
  10. Cloud & Teams56
  11. Settings & Customization58
  12. Account & Billing28
  13. Troubleshooting84
  14. CLI & API Reference22
  15. Legal & Policies4
  16. Uncategorised22

Auto-restart stuck sessions

Auto-restart stuck sessions — the self-healer for a session whose underlying process freezes right after you send a message and spins on "thinking" forever: how the wedge is detected, the setting that turns the automatic restart on or off, and what you see in each case.

What it is

Very rarely, the agent's underlying process freezes right after you send a message. It stops producing any output but still looks like it's working — the session shows the "thinking…" spinner and never moves. Under the hood the CLI is alive but wedged: it silently queued your message and never actually ran the turn. To you it looks like the session is thinking forever.

Auto-restart stuck sessions makes that self-healing. When a session has produced no output for several minutes while it is the agent's turn to reply, Omniscio treats it as wedged, stops the fake "thinking," and cleanly restarts the stuck turn so it continues — you get a real reply instead of an endless spinner. When the setting is off, Omniscio still stops the fake "thinking," but hands the session to you as Needs You to restart yourself instead of restarting it automatically.

Before this existed, a wedged "thinking" session had only one backstop — a 4-hour absolute turn cap — so a hang could sit for the better part of an hour (or until you restarted the app). This catches it in minutes.

Where to find it

How to turn it on / off

On by default. Open Settings → Sessions → the When sessions stall group (its advanced controls) and look for Auto-Restart Stuck Sessions (autoRestartUnresponsiveSessions, default on).

  • On — a wedged "thinking" session is force-recovered and its turn restarts automatically, with a short "stopped responding — automatically restarting it" note.
  • Off — the wedge is still stopped (no more fake "thinking"), but the session lands in your inbox as Needs You with a "send a message to restart it" note; your next message respawns it.

How it behaves

Why the fast stall detectors missed it

Omniscio already force-kills a session that produces no output within its first-output window (2 minutes, or up to 6 on a busy machine). But to be patient with a genuinely slow first token, the watchdog folds the CLI's transcript-file activity into its "last output" clock. A wedged-but-alive CLI keeps its transcript file warm (it writes your queued message, flushes keepalives), so that clock keeps getting reset and the silence-based timeout never fires — the session shows "thinking" indefinitely with no detection. This feature adds a second clock the transcript can't reset.

How detection works

The per-session stall watchdog anchors an absolute ceiling on when the turn's watchdog started — a clock that transcript activity can never reset. A session is treated as wedged only when all of these line up at once:

  1. It's the agent's turn — the session is running and has not yet produced its first output for this turn.
  2. It's been older than the 6-minute ceiling since the turn's watchdog armed.
  3. Yet its silence timer is still under budget — i.e. something (the warm transcript) has been resetting it, which is the exact wedge signature. A genuinely silent session trips the ordinary first-output timeout first and never reaches here.

Because the trigger is the absolute age plus the "silence-timer-was-reset" signature, a genuinely slow-starting CLI (which trips the ordinary timeout on real silence) is never cut off early, and a session that is legitimately waiting on you (its last message is the agent's) or parked in a declared wait is never touched — those aren't the agent's turn.

Behaviour summary

Situation What happens
Alive CLI, no output for 6 min while it's the agent's turn, transcript warm Wedge detected → force-recovered
…with the setting on (default) Turn is restarted automatically (error + response_aborted → aborted-response recovery)
…with the setting off Session surfaced as Needs You; your next message respawns it
Genuinely slow first token (real silence, cold transcript) Handled by the ordinary first-output timeout, not this — never cut off early
Session is waiting on you (last message is the agent's) Not touched — it isn't the agent's turn
Session parked in a declared wait (ScheduleWakeup / background job) Not touched — the deliberate wait is respected
Setting off, wedge detected Fake "thinking" still stopped; lands in the inbox as Needs You

A session that hangs inside its own compaction

A related case, and not governed by this page's setting: a long-running session can hang inside the conversation compaction its CLI does when a chat gets very large. You see the "Compacting conversation…" note appear (usually a few times) and then nothing — no "Conversation compacted", no output, and the session does not move. Underneath, the CLI process is alive but doing nothing at all.

Omniscio now recognises that shape specifically — the CLI said it was compacting and never said it finished — and restarts the session for you (a clean restart that keeps the conversation). It will do that at most twice, and if the conversation still will not compact it stops trying, parks the session with a "send it a message to continue" note, and tells you it gave up rather than looping.

  • Separate from the setting above — this recovery is not switched off by Auto-Restart Stuck Sessions, which watches for a different shape (a session wedged right after you sent a message, in its first few minutes).
  • Why it restarted rather than nudge — a process hung inside a compaction is not reading its turn input, so nudging it does nothing; the measured cure is a restart (in practice the retried compaction then finishes in about two minutes).
  • Anything already waiting for that session (a queued message, a cloud test result) delivers once it is back, in the normal way.

Out of scope (carry forward)

  • Tunable ceiling from the UI. The 6-minute ceiling is a code constant (the load-extended first-output wait's own top), exposed only as the on/off setting.
  • Recovering the exact queued input. Recovery restarts the turn (a clean --resume); the CLI's own dropped queue is not replayed byte-for-byte — the user's message is re-delivered on the respawn.

For agents

Where it lives in code

  • The detection + recovery — handleFirstOutputWindow and recoverUnresponsiveFirstOutput in src/main/process/session-stall-watchdog.ts. The ceiling reuses FIRST_OUTPUT_LOAD_CEILING_MS (6 min) from pty-output-constants.ts. Kill switch: AMC_DISABLE_FIRST_OUTPUT_ABSOLUTE_CEILING=1.
  • The compaction-wedge case — decideCompactWedgeRecovery (ndjson-context-decisions.ts) reads the turn's own turnHadCompactEvent / turnCompactionFinished pair, and the stall region relaunches through relaunchSessionInPlace(id, { trigger: 'auto', paced: true }), bounded by COMPACT_WEDGE_RECLAUNCH_MAX (2) inside a 30-minute decay window. Its contract clause is I11 in context-stall-recovery-contract.md.
  • Setting — autoRestartUnresponsiveSessions (default true) on AppSettings (session-behavior-settings.ts), Zod-validated (ipc-schemas), accessor getAutoRestartUnresponsiveSessions (accessors-settings.ts), UI toggle in SessionWaiting-definitions.ts.
  • Reproduction + regression tests — tests/integration/process-stall-detection.test.ts (the warm-transcript wedge, the off-path surface, and the "not cut off early" boundary).

Related

  • aborted-response-recovery.md — the sibling self-healer for when the CLI drops a turn (process exit / stream end); this one handles a turn that never started producing output.
  • context-stall-recovery.md — the out-of-room self-healer for a turn that completes but had no space to reply.
  • stalled-session.md — the orange "Stalled?" state + Auto-Continue Mid-Task (the broader post-first-output nudge).

Last verified 2026-10-06