---
title: Auto-restart stuck sessions
---

# Auto-restart stuck sessions

## What it is

Very rarely, the agent's underlying process **freezes right after you send a message**. It stops producing any output but still _looks_ like it's working — the session shows the "thinking…" spinner and never moves. Under the hood the CLI is alive but wedged: it silently queued your message and never actually ran the turn. To you it looks like the session is thinking forever.

**Auto-restart stuck sessions makes that self-healing.** When a session has produced no output for several minutes _while it is the agent's turn to reply_, Omniscio treats it as wedged, stops the fake "thinking," and cleanly restarts the stuck turn so it continues — you get a real reply instead of an endless spinner. When the setting is off, Omniscio still stops the fake "thinking," but hands the session to you as **Needs You** to restart yourself instead of restarting it automatically.

> Before this existed, a wedged "thinking" session had only one backstop — a 4-hour absolute turn cap — so a hang could sit for the better part of an hour (or until you restarted the app). This catches it in minutes.

## Where to find it

### How to turn it on / off

On by default. Open **Settings → Sessions** → the **Waiting & keeping sessions going** group and look for **Auto-Restart Stuck Sessions** (`autoRestartUnresponsiveSessions`, default on).

- **On** — a wedged "thinking" session is force-recovered and its turn restarts automatically, with a short "stopped responding — automatically restarting it" note.
- **Off** — the wedge is still stopped (no more fake "thinking"), but the session lands in your inbox as **Needs You** with a "send a message to restart it" note; your next message respawns it.

## How it behaves

### Why the fast stall detectors missed it

Omniscio already force-kills a session that produces **no output** within its first-output window (2 minutes, or up to 6 on a busy machine). But to be patient with a genuinely slow first token, the watchdog folds the CLI's **transcript-file activity** into its "last output" clock. A wedged-but-alive CLI keeps its transcript file _warm_ (it writes your queued message, flushes keepalives), so that clock keeps getting reset and the silence-based timeout **never fires** — the session shows "thinking" indefinitely with no detection. This feature adds a second clock the transcript can't reset.

### How detection works

The per-session stall watchdog anchors an **absolute ceiling** on when the turn's watchdog started — a clock that transcript activity can never reset. A session is treated as wedged only when **all** of these line up at once:

1. It's the **agent's turn** — the session is `running` and has not yet produced its first output for this turn.
2. It's been **older than the 6-minute ceiling** since the turn's watchdog armed.
3. Yet its **silence timer is still under budget** — i.e. something (the warm transcript) has been resetting it, which is the exact wedge signature. A genuinely silent session trips the ordinary first-output timeout first and never reaches here.

Because the trigger is the absolute age plus the "silence-timer-was-reset" signature, a genuinely slow-starting CLI (which trips the ordinary timeout on real silence) is never cut off early, and a session that is legitimately **waiting on you** (its last message is the agent's) or **parked in a declared wait** is never touched — those aren't the agent's turn.

### Behaviour summary

| Situation                                                                    | What happens                                                                             |
| ---------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------- |
| Alive CLI, no output for 6 min while it's the agent's turn, transcript warm   | Wedge detected → force-recovered                                                         |
| …with the setting **on** (default)                                           | Turn is restarted automatically (`error` + `response_aborted` → aborted-response recovery) |
| …with the setting **off**                                                    | Session surfaced as **Needs You**; your next message respawns it                          |
| Genuinely slow first token (real silence, cold transcript)                   | Handled by the ordinary first-output timeout, not this — never cut off early             |
| Session is waiting on **you** (last message is the agent's)                  | Not touched — it isn't the agent's turn                                                   |
| Session parked in a declared wait (ScheduleWakeup / background job)          | Not touched — the deliberate wait is respected                                            |
| Setting off, wedge detected                                                  | Fake "thinking" still stopped; lands in the inbox as Needs You                           |

### A session that hangs inside its own compaction

A related case, and **not** governed by this page's setting: a long-running session can hang *inside* the conversation compaction its CLI does when a chat gets very large. You see the **"Compacting conversation…"** note appear (usually a few times) and then nothing — no "Conversation compacted", no output, and the session does not move. Underneath, the CLI process is alive but doing nothing at all.

Omniscio now recognises that shape specifically — the CLI said it was compacting and never said it finished — and restarts the session for you (a clean restart that keeps the conversation). It will do that **at most twice**, and if the conversation still will not compact it stops trying, parks the session with a "send it a message to continue" note, and tells you it gave up rather than looping.

- **Separate from the setting above** — this recovery is not switched off by Auto-Restart Stuck Sessions, which watches for a *different* shape (a session wedged right after you sent a message, in its first few minutes).
- **Why it restarted rather than nudge** — a process hung inside a compaction is not reading its turn input, so nudging it does nothing; the measured cure is a restart (in practice the retried compaction then finishes in about two minutes).
- Anything already waiting for that session (a queued message, a cloud test result) delivers once it is back, in the normal way.

### Out of scope (carry forward)

- **Tunable ceiling from the UI.** The 6-minute ceiling is a code constant (the load-extended first-output wait's own top), exposed only as the on/off setting.
- **Recovering the exact queued input.** Recovery restarts the turn (a clean `--resume`); the CLI's own dropped queue is not replayed byte-for-byte — the user's message is re-delivered on the respawn.

## For agents

### Where it lives in code

- **The detection + recovery** — `handleFirstOutputWindow` and `recoverUnresponsiveFirstOutput` in [src/main/process/session-stall-watchdog.ts](../../src/main/process/session-stall-watchdog.ts). The ceiling reuses `FIRST_OUTPUT_LOAD_CEILING_MS` (6 min) from [pty-output-constants.ts](../../src/main/process/pty-output-constants.ts). Kill switch: `AMC_DISABLE_FIRST_OUTPUT_ABSOLUTE_CEILING=1`.
- **The compaction-wedge case** — `decideCompactWedgeRecovery` ([ndjson-context-decisions.ts](../../src/main/process/ndjson-context-decisions.ts)) reads the turn's own `turnHadCompactEvent` / `turnCompactionFinished` pair, and the stall region relaunches through `relaunchSessionInPlace(id, { trigger: 'auto', paced: true })`, bounded by `COMPACT_WEDGE_RECLAUNCH_MAX` (2) inside a 30-minute decay window. Its contract clause is I11 in [context-stall-recovery-contract.md](../../.claude/memory/contracts/context-stall-recovery-contract.md).
- **Setting** — `autoRestartUnresponsiveSessions` (default `true`) on `AppSettings` ([session-behavior-settings.ts](../../src/shared/types/settings/session-behavior-settings.ts)), Zod-validated ([ipc-schemas](../../src/shared/ipc-schemas/settings/session-behavior-settings.ts)), accessor `getAutoRestartUnresponsiveSessions` ([accessors-settings.ts](../../src/main/services/config-store/accessors-settings.ts)), UI toggle in [SessionWaiting-definitions.ts](../../src/renderer/src/features/settings/sections/session/SessionWaiting-definitions.ts).
- **Reproduction + regression tests** — [tests/integration/process-stall-detection.test.ts](../../tests/integration/process-stall-detection.test.ts) (the warm-transcript wedge, the off-path surface, and the "not cut off early" boundary).

## Related

- [aborted-response-recovery.md](aborted-response-recovery.md) — the sibling self-healer for when the CLI _drops_ a turn (process exit / stream end); this one handles a turn that never _started_ producing output.
- [context-stall-recovery.md](context-stall-recovery.md) — the out-of-room self-healer for a turn that _completes_ but had no space to reply.
- [stalled-session.md](stalled-session.md) — the orange "Stalled?" state + Auto-Continue Mid-Task (the broader post-first-output nudge).
