---
title: Aborted-response recovery
---

# Aborted-response recovery

## What it is

Sometimes the Claude CLI abandons its response mid-flight — the child process exits without a result, the output stream ends abruptly, or the stall watchdog times out waiting for the next token. When that happens there is no question to answer and no error the user typed something wrong; the turn just stopped. Before May 2026 these failures landed the session in the inbox as **`needs_you` with no pending action** — visually identical to a real question. The only way forward was for the user to notice the silent stall, click into the session, type "please continue", and hope.

**Aborted-response recovery makes those silent failures self-healing.** Every place in the app that announces an aborted response now stamps the session with a `response_aborted` marker. A background scan runs every 60 seconds, finds the stamped rows, and replays the user's last message through the same recovery machinery that rate-limit recovery uses — it kills the stuck child, forces a fresh transcript, and re-sends the last operator turn. Each session gets **2 automatic retries per day**; if both fail, the session is promoted to a durable **`recovery_failed`** state — a distinct **red** "Recovery failed" inbox row that the scan never retries again and that stays put until you act on it. While a session is still recovering it stays **invisible** (it looks like a normal running session); only once it lands in the red "Recovery failed" state does a **Retry** button appear on the session panel, forcing a restart immediately without waiting for the next scan tick.

The user-facing expectation: a CLI that hangs mid-response usually recovers within a minute, silently, with no notification — and you see nothing while it does. **While recovery is still going, the session is fully invisible** — it looks like a normal running session (no dot, no Retry button, no Inbox row, no chime), because the background scan is quietly handling it for you. **The invisibility is time-capped at 5 minutes** (2026-08-11): if recovery hasn't succeeded by then — a saturated machine deferring the scan, whatever the cause — the session stops hiding and surfaces in Needs You with an orange "Paused" dot and a **Retry** pill, while the background scan keeps retrying. A session can never again sit camouflaged as "running" indefinitely (before this cap, one finished session hid for 53 minutes). Separately, if it **exhausts its daily attempts** it surfaces as a **red "Recovery failed"** dot that persists until you send a message or click **Retry**. **The give-up is never the amber of a session merely waiting on you** ("red is for failures"), so a genuinely-broken session never hides among normal needs-you rows (a mid-2026 fix; before it, an exhausted session reverted to plain amber and looked identical to a real question). (Earlier builds showed an orange "Paused" dot while recovering; as of 2026-08-04 that in-progress state is invisible instead — only a real give-up, or now the 5-minute cap, is shown.)

> **Trade-off you accept:** recovery replays the _last operator message_, not a perfect resume of the half-finished response. For the common case (a transient network drop, a single connection slip) that is exactly right — the agent picks up where the user left off. For a CLI that is genuinely wedged and cannot make forward progress, the 2-attempts-per-day cap stops the retry loop from burning your account budget; after the cap the session waits quietly for you.

## Where to find it

### How to turn it on / off

Recovery is **on by default**. There are two independent kill switches, layered:

#### 1. Settings toggle (the everyday control)

Open **Settings → Sessions** and look for **Aborted-response recovery** (`abortedResponseRecoveryEnabled`, default on). Turn it off and the 60-second background scan never starts — sessions that abort still get the `response_aborted` label (so you can see what happened) and the **Retry** button still works, you just don't get the automatic retries.

#### 2. Environment kill switch (for diagnostics)

Set `AMC_DISABLE_ABORTED_RESPONSE_RECOVERY=1` in the environment before launching Omniscio. This disables the scan at boot exactly like the settings toggle, but can't be flipped from the UI — it's the "make the background machinery hold completely still while I debug" lever. Either switch being set is enough to stop the scan; both off is required to run it.

Note what is **not** gated by either switch: the emit-site wiring (sessions still get labeled when they abort) and the Retry button (you can always force one manually). Disabling recovery disables the _automatic_ retries, nothing else.

## How it behaves

### How recovery works (the scan loop)

The scan runs one cheap indexed query every 60 seconds (deferred 30 seconds past app boot so the database, auth, and process manager are all settled first). For each session it finds stamped `response_aborted`:

1. **Re-read the row.** A user message may have raced in since the scan started — if the session now carries a real user-facing action (a genuine question, a plan approval, a permission request, etc.) the scan defers to the user and leaves it alone.
2. **Check the status.** Only `needs_you` and `error` sessions are eligible. A `running` session is mid-turn (the abort signal is stale); `ended` / `archived` / `paused` are deliberate states the scan must never stomp on.
3. **Count the attempt.** A per-session daily counter is incremented atomically. If this would be the 3rd attempt today, the scan promotes the row to the durable `recovery_failed` state and skips the restart — the row collects in the sidebar's **Interrupted** section as a distinct **red** "Recovery failed" and the user takes over (as of 2026-08-21 a give-up is Interrupted, not Needs You — see [interrupted-sessions-section.md](interrupted-sessions-section.md)). (Before mid-2026 it _cleared_ the marker, which reverted the row to plain amber, indistinguishable from a normal question — that's the bug this replaced.)
4. **Replay.** The scan hands off to the shared replay machinery with a 90-second debounce. The debounce is deliberately longer than the 60-second scan cadence so two consecutive scan ticks can never both restart the same session.

The per-session daily counter resets the moment you send the session a real message of your own — if you're actively driving a session, the cap won't silence it.

#### Load-aware patience (the false-"Recovery failed" fix)

When your machine is busy (many sessions, worktrees, and background gates all at once), a brand-new session can take more than two minutes just to produce its _first_ word — not because it's broken, but because the CPU is saturated. The old behavior force-killed it, re-sent the message into the same busy machine, and after two failed retries marked it the durable red **"Recovery failed."** Two guards now stop that:

- **The stall watchdog waits longer for the first response.** As long as the session's process is still alive and the whole machine is genuinely busy, it keeps waiting (up to a 6-minute ceiling) instead of killing at the flat 2 minutes. An idle machine, a dead process, or hitting the ceiling all still give up — so a genuinely hung session still surfaces.
- **The scan skips a session entirely while the machine is saturated** — no replay (which would pile more load on), no attempt counted, no promotion to "Recovery failed." A session that aborted during a busy spell quietly recovers on its next turn once the load eases; only a genuine failure on a _calm_ machine reaches the durable terminal.

A related patience fix: before a session produces its first real word, the watchdog also counts the CLI's **stderr startup chatter** — MCP servers starting up, tool banners — as proof of life, so a CLI that's visibly busy getting its tools ready is never killed as "no output." Once real output starts flowing, stall detection goes back to watching the response itself.

#### A missing project folder pauses recovery instead of burning it

If the project's folder has vanished — moved, renamed, or on a drive that isn't plugged in — no retry can possibly succeed until the folder is back, so recovery doesn't spend any. The scan probes the folder **before** counting an attempt: while it's missing, the session is skipped with nothing charged against the daily cap, so a folder that's gone for the afternoon can't march a healthy session into "Recovery failed." You're told once — the red "Project folder is missing" banner over the composer plus a one-line note in the chat (see [project-folder-missing.md](project-folder-missing.md)) — and the moment the folder is back, recovery resumes on its own. Reconnect the drive and walk away; no manual retry needed.

#### Nothing else talks to a cut-off session before the recovery does

A session waiting for recovery is marked with exactly the state the scan looks for, so any other automatic message sent to it first would erase that mark and the cut-off work would never be replayed. That is precisely what Plain Speak's "you forgot your summary card" request used to do: it arrived first, the agent answered it, the turn ended cleanly, and the session then sat idle until someone noticed (four hours in one live case; 58 cut-off turns across 45 sessions in the five days before the fix). Since 2026-09-21 a cut-off turn gets no Plain Speak card and no card request at all, so the recovery scan always gets to resume it.

#### A worker that died before it started is cleaned up, not retried forever

Some sessions are **spawned by another session** — an orchestrator that fans out a batch of workers, like the "Fix all failing tests" run. Once in a while one of those workers dies in the first few seconds, before it ever really starts: no first response, not even its opening task message saved. There is nothing to replay or resume, so the background retry can only fail over and over — about a dozen doomed attempts across ~3½ hours — while leaving a stuck red **"Recovery failed"** card you'd have to clear by hand. Omniscio now recognizes this exact shape — a spawned worker that never established a session and has nothing to replay — and quietly **archives** it instead. The orchestrator that launched it already spins up a fresh replacement, so no work is lost; the dead stub just tidies itself away rather than churning in red. This only ever touches those disposable spawned workers: a session **you** started, or one that got far enough to have real content, is never auto-archived — it stays put for you exactly as before. (For diagnostics, `AMC_DISABLE_PREINIT_ORPHAN_RETIRE=1` turns it off.)

#### A turn that loses its "done" signal is sorted by _why_ — not parked in amber

Once in a while the agent finishes its whole answer, but the CLI never sends the small "turn complete" signal that normally flips a session out of the green **running** state — the signal gets dropped or badly delayed, most often when the machine is so busy it stalls the CLI at the exact moment it wraps up. The finished answer just sits on screen, green, with nothing happening; and if you then message the session to nudge it, that can start a fresh empty turn that times out into a misleading **"error."**

Omniscio treats the agent's **own end-of-message marker** — the deterministic line it prints right before its final message — as a backup "this turn is done" signal. When that marker is present and the session then goes quiet for about 20 seconds, Omniscio **works out _why_ the signal dropped** instead of blindly parking the session in amber "Needs You": if the turn hit a **rate limit** it parks as orange **"waiting"** and recovers on its own; if the message was **genuinely cut off mid-delivery while the machine was slammed** (the marker printed but nothing after it) it becomes a self-healing **"retrying"** row the 60-second auto-retry re-drives; and a **delivered** answer — any real text after the marker, a question for you, or a clean answer on a calm machine — settles straight to amber **"your turn."** The delivered-answer rule is the 2026-08-11 fix: a turn whose final message fully streamed is DONE, so it goes to Needs You immediately even on a slammed machine — hiding it and silently re-running it both wasted money and camouflaged a finished session as green (one sat hidden for 53 minutes). That stops an _interrupted_ turn from masquerading as "the agent is waiting on you," while still fixing the original stuck-green / false-error problem. Any further output from the agent resets the timer, and quiet alone is never enough: while the agent's latest step is a tool call — a command still running, or its result not yet answered — or its own session log shows it is still writing its next step, the turn is not treated as done however long it is silent. (Before 2026-09-27 that check read a field the agent's stream never actually sends, so an agent that printed its marker early — DeepSeek often does — could be cut off 20 seconds into a long command; see the [mid-tool-call false settle postmortem](../../.claude/memory/postmortems/archived/stall-watchdog-mid-tool-call-false-settle-postmortem.md).) If you ever need the old behavior for debugging, `AMC_DISABLE_FINAL_MARKER_SETTLE=1` turns it off.

#### The Retry button

The **Retry** pill appears on red **"Recovery failed"** sessions — the durable give-up state — and (2026-08-11) on a still-recovering session **once its 5-minute invisibility cap expires**: at that point the session is visible in Needs You anyway, so the pill gives you the immediate manual lever alongside the background scan that keeps retrying. Within the first 5 minutes of recovering, no pill shows and the session looks like a normal running session: the 60-second scan is retrying it silently for you, so there is nothing to click. This is deliberate (as of 2026-08-04): the recovering state surfaces nothing until recovery gives up — or the cap expires — mirroring how an auto-recovering rate limit stays hidden while it self-heals.

A session parked in the durable `recovery_failed` state shows the one-click pill in **red** — a failure keeps failure colors. Clicking it resets the retry budget and the backoff ladder outright, then fires one restart immediately — and that works even on a session whose background auto-retry hit its hard cap and stopped for good. Previously the terminal state had no button at all; you had to know that typing a message is what un-parks a given-up session.

**The label is honest about what's happening behind it.** While a "Recovery failed" session is still being retried in the background, the pill's tooltip says so — "Auto-recovery keeps retrying in the background — Retry tries again immediately." Once auto-retry has fully **stopped** — repeated failures reached the hard cap, or the error is known-terminal (out of credit, an invalid API key) — it reads "Auto-retry stopped after repeated failures — Retry resets the limit and tries again now." That flip happens live: a panel you already have open updates the moment auto-retry gives up, instead of claiming background retries that are no longer happening.

**On an errored session, Retry is the ONE recovery button (2026-09-16).** It used to appear alongside a generic **Continue** pill, so a session that stopped with an error offered two ways to say "try again" — plus the separate **Switch model** handoff banner above the transcript. Continue and Retry are the same intent, and Retry is the stronger of the two (it resets the re-arm ladder and drives recovery immediately rather than just messaging the session), so the Continue pill now hides whenever the Retry pill is showing. The `C` keyboard shortcut follows the visible button: on those sessions it fires the same Retry instead of a nudge with no button behind it. Nothing else changes — Continue still appears for the recovery states that have no Retry pill (paused, ended, stalled, a stopped or app-restart-suspended session), and **Switch model** stays, because starting over in a different model is a genuinely different move rather than a second way to retry.

### A reply that arrives unfinished

There is one failure the scan loop above cannot see, because nothing looks wrong: the CLI **reports the turn as finished** while the reply itself was cut off mid-sentence. The model stops mid-word, the provider calls it a clean finish, and there is no error and no `response_aborted` marker to act on — so before 2026-09-29 a half-written answer simply reached you as though it were complete.

The app now recognises the one shape of this it can **prove**. A **dev-pipeline gate report** is defined by the closing marker on its last line, so a reply that OPENS with a gate report's heading and never reaches that closing marker is unfinished by definition — no guesswork about prose. When that happens the app asks the agent to finish the report it was writing and **holds the turn out of your inbox** until it does, so you are never handed a half-written report — and never a half-written question list you could have answered as if it were the real thing. It fires **at most once per message you send**: if the second attempt is cut off as well, the turn concludes normally and you see exactly what you would have seen before.

Only gate reports can be caught this way. A truncated ordinary reply is **not** detected — the obvious test ("does it end mid-word?") flagged thousands of perfectly normal sentences when it was measured, so it is deliberately not used. This is a turn-end rule, not part of the 60-second scan, and it needs no marker on the session.

### Upgrade & retro-fill

When you upgrade to the build that introduced this feature, a one-time database migration (schema **v256**) does two things:

1. Adds the per-session daily-attempt counter column.
2. **Retro-fills** the `response_aborted` marker onto every existing row that matches the original bug's fingerprint — sessions that are `needs_you` or `error` with no pending action and not deleted. This means the silent failures you accumulated _before_ upgrading get picked up by the very first scan after the upgrade, instead of sitting in your inbox forever.

The migration is idempotent — running it again on an already-migrated database is a no-op, and the retro-fill is narrowed to the exact bug fingerprint so it never relabels a real question, a rate-limited row, or a plan-approval row as an aborted response.

#### Manual cleanup script (escape hatch)

If you want to apply the retro-fill _without_ upgrading — or you upgraded, downgraded, then upgraded again and some rows got stamped back to NULL — there's a standalone script: [scripts/triage-aborted-sessions.mjs](../../scripts/triage-aborted-sessions.mjs). It runs against your live database with no Electron boot:

```
node scripts/triage-aborted-sessions.mjs              # dry run (default) — counts + lists the affected rows
node scripts/triage-aborted-sessions.mjs --execute    # stamps response_aborted on the matching rows
node scripts/triage-aborted-sessions.mjs --db-path <path>   # point at a non-default DB
```

Dry-run is read-only and safe even while Omniscio is running (the DB is in WAL mode). The `--execute` write is also WAL-safe but best done with Omniscio closed to avoid churn. The UPDATE is narrowed to the same fingerprint as the migration, so re-runs are idempotent.

### Hand-off with rate-limit recovery

Omniscio already had a recovery service for rate-limited sessions. Aborted-response recovery is its sibling — same restart machinery, different scan signal. The two are wired so they never both claim the same session: a `rate_limited` row belongs to rate-limit recovery (aborted-response recovery skips it), and a `response_aborted` row belongs to aborted-response recovery (rate-limit recovery's own eligibility check skips it). Each service's eligibility predicate explicitly returns "not mine" for the other's signal.

### Behaviour summary

| Situation                                            | What happens                                                                                                                                                                                                                                                                            |
| ---------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| CLI aborts mid-response                              | Session stamped `response_aborted` — stays **invisible** (looks like a normal running session) while the scan retries it, **capped at 5 minutes**: past the cap it surfaces in the sidebar's **Interrupted** section (red dot + Retry pill) while the scan keeps retrying                                    |
| Recovery enabled, 1st/2nd attempt of the day         | Scan replays last message within ~60s, silently, no notification, nothing shown                                                                                                                                                                                                          |
| Recovery enabled, 3rd attempt of the day             | Promoted to durable red **"Recovery failed"** (now the FIRST time it becomes visible); scan never retries it again                                                                                                                                                                       |
| You send the session a real message                  | Daily attempt counter resets to zero                                                                                                                                                                                                                                                    |
| You click **Retry** on a red "Recovery failed" row   | Retry budget + backoff ladder reset, one immediate restart — works even after auto-retry stopped at its hard cap (the pill only appears on this red state)                                                                                                                               |
| Recovery disabled (settings OR env)                  | No auto-retry; row still labeled, **Retry** button still works                                                                                                                                                                                                                          |
| Upgrade to v256                                      | Existing silent-failure rows retro-filled + picked up by the first scan                                                                                                                                                                                                                 |
| Real question / plan / permission races in mid-scan  | Scan defers to the user, leaves the session alone                                                                                                                                                                                                                                       |
| Machine under heavy CPU load                         | First-output kill is deferred (alive process, ≤6-min ceiling) and the scan skips advancement — no false "Recovery failed" from a busy spell                                                                                                                                             |
| Project folder missing (moved / renamed / unplugged) | Recovery pauses without counting an attempt; folder banner + one chat note; resumes by itself when the folder returns                                                                                                                                                                   |
| The session still owns a running gate (2026-09-23) | Recovery **waits** instead of restarting: an auto-restart replaces the turn and would kill that running gate with it. Nothing is charged against the retry cap, and recovery proceeds by itself on the first tick after the gate closes. Clicking **Retry** yourself still restarts immediately. The same hold protects the other automatic restarts (rate-limit resume, account switch, idle release) |
| A spawned worker dies at startup (never ran, nothing to replay) | Auto-**archived** instead of looping in red "Recovery failed" — its orchestrator re-spawns the work; only these disposable spawned workers, never a session you started (`AMC_DISABLE_PREINIT_ORPHAN_RETIRE=1` to disable)                                                     |
| A turn's "done" signal drops after the answer marker | Sorted by _why_ once the stream goes quiet (~20s): rate-limited → orange "waiting" (self-recovers); message genuinely cut off mid-delivery under heavy load → **invisible** "retrying" (auto-retried silently, 5-min cap); a **delivered** answer, a real question, or a clean answer on a calm box → amber "your turn" — a finished session can never hide |
| Model quits mid-reasoning with a "clean" finish (thinking-only final message — seen on third-party reasoning models, e.g. GLM/DeepSeek shims) | Classified as an aborted response and auto-resumed by the 60s scan — never a needs-you with an empty "finished" message; a "resuming it automatically" note explains the pause (2026-08-17, contract I19) |

### Out of scope (carry forward)

- **Per-session retry-cap override.** The 2-attempts-per-day cap is a single global constant, not a per-session or per-project setting. If a session needs more retries, send it a message (which resets the counter) or click Retry.
- **Configurable scan cadence from the UI.** The 60-second interval is a code constant exposed only as the env/settings kill switch, not a per-user tuning knob.
- **Notifications on recovery.** Successful auto-recovery is deliberately silent — no toast, no inbox row. The whole point is that the user never had to notice.
- **Resuming the exact half-finished response.** Recovery replays the last operator message; it does not reconstruct the partial output the CLI abandoned.

## For agents

### Where it lives in code

- **The service** — [src/main/services/aborted-response-recovery-service.ts](../../src/main/services/aborted-response-recovery-service.ts) is the sole owner of the periodic scan, the per-session attempt counter, the operator-message reset, and the `forceRestart()` entry point the Retry button calls. Three tunable constants live at the top: 60-second scan interval, 2-attempts-per-day cap, 90-second replay debounce. The `isResponseAbortedRecoveryEligible()` predicate is the exhaustive switch over every pending-action value that decides whether the scan owns a session.
- **Service lifecycle** — started/stopped from [src/main/index.ts](../../src/main/index.ts) alongside the other periodic services (after the rate-limit rehydrate step).
- **The restart hand-off** — the scan calls `replayLastOperatorMessage()` on the process manager in [src/main/process/process-manager.ts](../../src/main/process/process-manager.ts) with `source: 'response-aborted'`; that same file's `writeToStdin()` fires the operator-message reset on every real user message.
- **Schema migration v256** — [src/main/db/database.ts](../../src/main/db/database.ts) adds the `aborted_response_retry_attempts` column and the idempotent retro-fill UPDATE. The atomic counter queries (`incrementResponseAbortedAttempts`, `resetResponseAbortedAttempts`, `clearResponseAbortedPendingAction`, `getResponseAbortedSessionIds`) live in [src/main/db/queries-sessions/status.ts](../../src/main/db/queries-sessions/status.ts).
- **The `response_aborted` + `recovery_failed` enum values** — declared on the `PendingAction` union in [src/shared/types/session-status.ts](../../src/shared/types/session-status.ts); the inbox dot/label/tooltip mapping (`recovery_failed` → `red-400` "Recovery failed" — a failure, recolored from orange 2026-06-24) is in [src/renderer/src/lib/utils.ts](../../src/renderer/src/lib/utils.ts). As of 2026-08-04 a still-recovering `response_aborted` session renders **invisible** — `pickStatusRow` ([src/renderer/src/lib/status-display.ts](../../src/renderer/src/lib/status-display.ts)) routes it to the normal green "Running" row (no ember "Paused" dot), and the Retry pill in [SessionPanel.tsx](../../src/renderer/src/features/sessions/SessionPanel.tsx) is gated to `recovery_failed` only. The suppression across every attention surface (inbox, badge, count, chime, plugin poll) is owned by the one predicate `isAutoRecoveringAbortedResponse` in [src/shared/session-attention.ts](../../src/shared/session-attention.ts) (repo-only `needs-you-visibility-contract.md` I19). The `setRecoveryFailedPendingAction()` promote-on-exhaustion query lives in [src/main/db/queries-sessions/status.ts](../../src/main/db/queries-sessions/status.ts). `recovery_failed` is **also** set directly (bypassing this scan) by the placeholder-stuck terminal sink — `firePlaceholderStuckExhausted()`'s `cappedOut` branch in [src/main/process/process-manager.ts](../../src/main/process/process-manager.ts) — once a session exhausts its per-account placeholder-stuck retries + cross-account fallback, so a genuinely-wedged session shows durable red at once instead of looping (see the `context-window-1m-overfeed` postmortem, repo-only).
- **Pre-init orphan retire (a `recovery_failed` row that is ARCHIVED, not re-armed)** — a spawned orchestrator child that died before establishing a CLI session (`cliSessionId` null) with no operator message to replay has nothing the universal re-arm can act on, so `retirePreInitOrphan()` in [src/main/services/stuck-session-recovery-service.ts](../../src/main/services/stuck-session-recovery-service.ts) archives it (`archiveSession(id, 'orphan-preinit')`, falling back to `recovery_terminal` on a `stalled` row) instead of doomed-re-arming it, via the pure `isPreInitDeadChild` predicate in [src/main/services/recovery/preinit-orphan.ts](../../src/main/services/recovery/preinit-orphan.ts). Kill switch `AMC_DISABLE_PREINIT_ORPHAN_RETIRE=1`; the D4 invariant in `transient-spawn-retry-contract.md` (repo-only) is the authority.
- **Retry button + IPC** — the pill is rendered in [ComposerStatusBanners.tsx](../../src/renderer/src/features/sessions/SessionPanel/ComposerStatusBanners.tsx) and its click and the `C` hotkey both call the one shared action in [retry-recovery-action.ts](../../src/renderer/src/features/sessions/retry-recovery-action.ts); that invokes the `SESSION_RETRY_ABORTED_RESPONSE` channel, handled in [src/main/ipc/session-handlers.ts](../../src/main/ipc/session-handlers.ts), which calls `forceRestart()`.
- **Settings field** — `abortedResponseRecoveryEnabled` (default `true`) on `AppSettings` in [src/shared/types.ts](../../src/shared/types.ts), with the matching Zod field in [src/shared/ipc-schemas/update-settings.ts](../../src/shared/ipc-schemas/update-settings.ts).
- **Sibling hand-off** — the reciprocating eligibility check (`hasActionablePending` returning true on `response_aborted`) lives in [src/main/services/rate-limit-recovery-service.ts](../../src/main/services/rate-limit-recovery-service.ts).
- **Cleanup script** — [scripts/triage-aborted-sessions.mjs](../../scripts/triage-aborted-sessions.mjs).
- **Final-marker settle** — a turn whose CLI never sent its "done" signal is CLASSIFIED by `handleFinalMarkerSettle` in [src/main/process/session-stall-watchdog.ts](../../src/main/process/session-stall-watchdog.ts) (2026-07-22, Option B; delivered-answer gate 2026-08-11): `turnHadRateLimitEvent` → `waiting`+`rate_limited`; `!turnHadQuestion && isBoxUnderHeavyLoad() && !finalMarkerAnswerDelivered(session)` → `needs_you`+`response_aborted` (the self-healing shape — invisible while recovering, 5-min cap — paired with a kinded notable row that also carries `interruption: true` so the "The turn was interrupted before completing — retrying." bubble is hidden by default in the conversation; the row is preserved for audit); else bare `needs_you` (a real question, a clean answer on a calm box, or a DELIVERED final answer under load — any non-empty text after the marker means the answer reached the user, so it surfaces instead of being hidden + re-run). Keyed on the per-turn `turnHadFinalMarker` flag latched at stream time via `containsFinalMessageMarker` in [src/shared/agent-content-markers.ts](../../src/shared/agent-content-markers.ts). Grace `FINAL_MARKER_SETTLE_MS` + kill switch `AMC_DISABLE_FINAL_MARKER_SETTLE` in [pty-output-constants.ts](../../src/main/process/pty-output-constants.ts); invariant I13 in the contract. A turn's own "done" signal that arrives while the settle is still saving makes the settle stand down, and neither watchdog ever concludes a turn whose normal wrap-up is still saving — both wait for it, up to 3 minutes (`TURN_DISPOSITION_PENDING_CAP_MS`, [turn-disposition-pending.ts](../../src/main/process/turn-disposition-pending.ts)), so a slow save under load no longer drops an "archive me" sign-off, a question, a crew report or a pipeline auto-advance.
- **The 5-minute invisibility cap** — `ABORTED_RECOVERY_HIDE_MAX_MS` + `abortedRecoveryHideExpired` beside the one suppression predicate `isAutoRecoveringAbortedResponse` in [src/shared/session-attention.ts](../../src/shared/session-attention.ts), clocked on the row's `statusChangedAt` (stable while it sits in the shape). Past the cap every surface surfaces the session — inbox/badge (the shared predicate), the dot ([status-display.ts](../../src/renderer/src/lib/status-display.ts)), the OS dock count (the SQL mirror in [cost.ts](../../src/main/db/queries-sessions/cost.ts)), the plugin poll, and the Retry pill. Two once-a-minute pulses catch expiry on a quiet machine: the renderer's attention context ([useAttentionContext.ts](../../src/renderer/src/hooks/useAttentionContext.ts)) and a coalesced badge recompute in [notification-service.ts](../../src/main/services/notification-service.ts). Needs-you-visibility contract I19 (repo-only) is the authority.
- **Thinking-only / empty-final classification** — a turn that concludes with a thinking-only final assistant message (clean `end_turn`, no text — a reasoning model that died mid-thought) or with no visible content at all is routed to the self-healing shape by the `thinking-only-endturn-recovery` gate (`tryThinkingOnlyFinalTurn`) in [src/main/process/ndjson-auto-disposition.ts](../../src/main/process/ndjson-auto-disposition.ts), keyed on the per-turn `lastAssistantBlockType` tracker set in [src/main/process/ndjson-assistant-content.ts](../../src/main/process/ndjson-assistant-content.ts). Guards keep delivered answers, questions, and intentional background waits untouched; kill switch `AMC_DISABLE_THINKING_ONLY_ENDTURN_RECOVERY=1`; invariant I19 in the contract (live incident 2026-08-17, crofai/deepseek "GLM 5.2").
- **Unfinished-gate-report classification (a reply reported FINISHED that was cut off)** — a turn whose final text OPENS WITH a dev-pipeline gate header but carries none of the byte-exact closing marker that ends such a report is re-driven and held by the `incomplete-gate-report-nudge` gate (`tryIncompleteGateReportNudge`) in [src/main/process/turn-disposition-gates/incomplete-gate-report-nudge.ts](../../src/main/process/turn-disposition-gates/incomplete-gate-report-nudge.ts), registered in [src/main/process/ndjson-turn-disposition-gates.ts](../../src/main/process/ndjson-turn-disposition-gates.ts) immediately after the aborted-response recoveries above. Detection reuses the EXISTING fence-aware marker helpers (`isDevPipelineGateHeader` + `hasDevPipelineGateMarker`), so incompleteness is proved from the app's own definition of a gate report rather than guessed from prose. The nudge is sent through the shared `sendTurnEndNudge` in [turn-disposition-gates/shared-helpers.ts](../../src/main/process/turn-disposition-gates/shared-helpers.ts); bounded once per operator-message episode via `incompleteGateReportNudgeCount`, reset with the sibling counters in [src/main/process/user-message-router.ts](../../src/main/process/user-message-router.ts); kill switch `AMC_DISABLE_INCOMPLETE_GATE_REPORT_NUDGE=1`. It deliberately does NOT yield on `pendingAction !== null` — the live incident's truncated report carried a lettered question list, so that guard would make it blind on the only turn it exists for (live incident 294f0e4b, 2026-09-29; contract rule `F10`).
- **Load-aware patience** — the stall watchdog's first-output defer ([src/main/process/session-stall-watchdog.ts](../../src/main/process/session-stall-watchdog.ts)) and the scan's heavy-load skip both read the whole-box CPU signal from [src/main/process/cpu-load-sampler.ts](../../src/main/process/cpu-load-sampler.ts) (`rollingBoxBusyPercent` / `isBoxUnderHeavyLoad`). Tunables: `FIRST_OUTPUT_LOAD_CEILING_MS` (6-min ceiling, [pty-output-constants.ts](../../src/main/process/pty-output-constants.ts)) and `HEAVY_BOX_LOAD_BUSY_PCT` (85%). The watchdog defer (Part A) is bounded by the ceiling; the scan defer (Part B) never escalates to the terminal under load (a busy spell can't manufacture "Recovery failed" — see the `aborted-response-false-recovery-under-load` postmortem, repo-only), but since 2026-08-10 it is no longer a forever-defer: after ~10 consecutive load-deferred ticks a sustained-load **escape valve** forces one paced replay per minute (no cap burn, no terminal promotion), so a permanently-pinned machine can't strand a session either.

The feature's test-locked invariants and safe-change checklist live in the feature contract at `.claude/memory/contracts/aborted-response-recovery-contract.md` (repo-only).

## Related

- [session-stuck-in-needs-you.md](session-stuck-in-needs-you.md) — the other pending-action causes (question, plan, permission, rate-limit, auth, API error, stopped) and how to clear each; aborted-response is the self-healing one that doesn't need you.
- [crash-recovery.md](crash-recovery.md) — auto-resume of sessions that were running when Omniscio itself exited; aborted-response recovery is the in-session analogue for when the CLI (not Omniscio) drops the turn.
- [waiting-detector.md](waiting-detector.md) — suppresses spurious `needs_you` flips when the agent says it's waiting; aborted-response recovery handles the opposite case where the turn genuinely died.
