Omniscio documentation
Browse all documentation
  1. Getting Started13
  2. Sessions & Agents115
  3. Inbox & Notifications59
  4. Projects & Tasks95
  5. Automation & Scheduling75
  6. Knowledge & Memory26
  7. AI Features60
  8. Integrations100
  9. Plugins & Marketplace33
  10. Cloud & Teams56
  11. Settings & Customization58
  12. Account & Billing28
  13. Troubleshooting84
  14. CLI & API Reference22
  15. Legal & Policies4
  16. Uncategorised22

Out-of-context stall recovery

When a heavy turn fills the context window mid-generation, an agent can simply stop with no closing message. This page covers the self-healing that follows: Omniscio nudges the agent to continue — or forces a compaction, on engines that cannot compact themselves — up to three times, then hands the session back to you if it still cannot finish.

What it is

A long agent session keeps filling its context window. The Claude CLI compacts itself between turns, but a single heavy turn can balloon all the way to 100% full mid-generation and end on a bare tool result — the model had no room left to write its closing reply. The turn just stops. The chat shows the agent's last finished line and then nothing; to you it looks like the agent quit mid-thought with no final message.

Out-of-context stall recovery makes that self-healing. When a turn ends with the context window near-full and the agent did not ask you a question, isn't deliberately parked waiting on something, and did not already finish what it was doing, Omniscio automatically sends a "Please continue" nudge. That opens a fresh turn, which lets the CLI compact and reclaim room, and the agent finishes the thought it ran out of space for. You just see a completed reply, with a small "Auto-recovering…" note so you know what happened.

It is bounded: if a few nudges in a row don't free enough room (a genuinely un-compactable wedge), Omniscio stops nudging, posts a "couldn't finish on its own — may need a fresh session" note, and drops the session into your inbox so you can take over. It never loops forever.

Trade-off you accept: recovery spends a little of your token budget on each nudge (it's a real turn). The cap (3 consecutive nudges) stops a wedged session from burning the budget; after the cap it waits quietly for you.

Where to find it

There is nothing on screen to click while a recovery is happening beyond the small "Auto-recovering…" note in the session; the only surface you manage is the switch that turns the behaviour on or off, which lives in Settings → Sessions, in the group about waiting and keeping sessions going. It ships on, so unless you go looking for that switch you will never need to visit it.

How it behaves

How to turn it on / off

Recovery is on by default. Open Settings → Sessions → the Waiting & keeping sessions going group and look for Auto-Recover When Out of Context (contextStallAutoRecoverEnabled, default on). Turn it off and a turn that stops at the context ceiling simply lands in your inbox as it did before — no automatic nudge.

This is a separate toggle from Auto-Continue Mid-Task. Auto-Continue nudges the agent on every mid-task stop (off by default, because it's aggressive). Out-of-context recovery fires only in the narrow "ran out of room" case, which is why it's safe to have on by default.

How recovery works

At the end of every turn Omniscio reads the live /context percentage — plus whether the turn just compacted — and decides:

  1. Is it the out-of-room case? The agent did not ask a question (no pending action), the session is not parked in a waiting hold, the turn did not already deliver its final answer, AND the turn is out of room — either context is at/above 90% full, OR the turn just compacted but produced no answer. That second signal matters: a compaction clears the live context reading on that same turn, so the percent comes back "unknown" exactly when recovery is most needed — and a compaction that produced no answer is, by definition, out of room (the CLI only compacts because it ran out). Without it, such a session loops silently on "This turn ended without a response" while you manually type "continue" (the 2026-07-13 bug this signal fixed). A plain unknown reading with no compaction still never guesses → do nothing. And "full" is not the same as "stuck": if the turn ran to 100% but still wrote its closing reply, it is finished, not out of room — nudging it would only spend another turn re-confirming it and drop a done session into your inbox with fresh questions. (That is exactly what happened on 2026-08-25, seventeen milliseconds after a completed run reported "nothing here needs you"; recovery now leaves a finished turn alone.)
  2. Recover. Send a hidden auto-message and count the attempt. For Claude it's "Please continue." — the fresh turn triggers the CLI's own compaction. For every model that runs on a custom endpoint — Kimi / DeepSeek / GLM (anthropic-compat) and GPT / xAI (proxy-backed) — which never self-compact on a continue turn, Omniscio instead forces /compact directly (verified to shrink a real Kimi session ~54%, 92k → 42k); you see the resulting "Conversation compacted" note. Either way, room is freed for the agent to finish. This is what keeps a big GPT session compacting before it slams into its context wall — those sessions run with the CLI's own auto-compaction turned off (to unlock the model's full window), so this forced /compact is the only thing that trims them, and without it they used to run to the wall and hard-fail with "Prompt is too long."
  3. Give up after the cap. After 3 consecutive nudges with the context still pinned, stop and post a single visible "couldn't finish on its own — may need a fresh session" note — and only that note (the generic "turn ended without a response" line is suppressed on give-up so the two don't contradict each other) — then let the session fall to Needs You.

The attempt counter resets the moment you send the session a real message, or as soon as the context drops back below the threshold (a compaction worked) — so a healthy session that briefly touched the ceiling starts fresh next time.

Behaviour summary

Situation What happens
Turn ends at ≥90% context, no question, not waiting Omniscio nudges "Please continue" (Claude) or forces /compact (Kimi/DeepSeek/GLM/GPT/xAI); the agent compacts + finishes (1–3 tries)
Turn compacted but produced no answer (percent reads unknown) Treated as out of room — Omniscio nudges (or forces /compact) and finishes; same 1–3 cap + give-up
Recovery succeeds You see a small "Auto-recovering…" note then the finished reply
3 nudges in a row, context still pinned Visible "couldn't finish — may need a fresh session" note; → Needs You
Agent ended the turn by asking you a question No nudge — the question goes to your inbox as normal
Turn hit ≥90% but still delivered its final answer No nudge — the turn is finished, not stuck; recovery leaves it alone
Agent is parked in a waiting hold (ScheduleWakeup / wait) No nudge — the deliberate wait is respected
Context drops back under 90% (a compaction worked) Attempt counter resets
You send the session a real message Attempt counter resets
Setting off No nudge; the stalled turn lands in the inbox as before

Out of scope (carry forward)

  • Forcing /compact for Claude. For Claude, recovery relies on a fresh "continue" turn triggering the CLI's own compaction (verified sufficient) — it does not inject /compact. (For every custom-endpoint engine — anthropic-compat Kimi / DeepSeek / GLM and proxy-backed GPT / xAI — which never self-compact, Omniscio does force /compact; see step 2 above.)
  • Tunable threshold / cap from the UI. 90% and 3 are code constants, exposed only as the on/off setting.
  • Silent recovery. A small "Auto-recovering…" note is shown on purpose (transparency). There's no separate "fully silent" mode.
  • Resuming the exact half-written reply. Recovery continues the turn forward; it doesn't reconstruct whatever the model was mid-token on when it ran out of room.

For agents

Where it lives in code

  • The decision — decideContextStallRecovery() in src/main/process/ndjson-context-decisions.ts is a pure function returning recover | give-up | none, from the live percent, the question / waiting-hold / turn-already-finished signals, a knownOutOfRoom flag (the compaction-with-no-answer case), and the two tunable constants CONTEXT_STALL_RECOVER_PERCENT (90) and CONTEXT_STALL_RECOVER_MAX (3).
  • The wiring — applyContextStallRecovery in src/main/process/ndjson-auto-disposition.ts (delegated from handleTurnComplete in ndjson-event-handlers.ts) feeds it the live percent (resolveContextWarningUsage), the question signal (pendingAction), the waiting-hold signal (session.deferredWork), the finished-turn signal (turnFinishedReason() in turn-disposition-gates/shared-helpers.ts — the SAME predicate the Auto-Continue gate reads, so the two nudge paths can't disagree about whether a turn is done), and knownOutOfRoom = turnHadCompactEvent && !turnHadVisibleAnswer; on recover it nudges via writeToStdin, on give-up it posts the banner and finalizeTurnDisposition suppresses the generic "turn ended" line before falling through to Needs You. The counter (session.contextRecoverCount) resets there and in writeToStdin (process-manager.ts).
  • Setting — contextStallAutoRecoverEnabled (default true) on AppSettings (session-behavior-settings.ts), Zod-validated (ipc-schemas), accessor getContextStallAutoRecoverEnabled (accessors-settings.ts), UI toggle in WorkflowSettings-definitions.ts.

The feature's test-locked invariants and safe-change checklist live in the contract at .claude/memory/contracts/context-stall-recovery-contract.md (repo-only).

Related

This is one of two self-healers for a turn that did not finish. Its sibling, for when the CLI drops a turn outright rather than completing it out of room, is on the aborted response recovery page; the broader, off-by-default nudge for any mid-task stop is on the stalled session page; and the detector that tells the two apart from a deliberate wait is on the waiting detector page. What the session shows you when its context is merely getting full, short of a stall, is on the context warnings page.

  • aborted-response-recovery.md — the sibling self-healer for when the CLI drops a turn (process exit / stream end / stall), versus this one for when a turn completes but the agent was out of room.
  • stalled-session.md — the orange "Stalled?" state + Auto-Continue Mid-Task (the broader, off-by-default nudge).
  • waiting-detector.md — suppresses needs_you flips when the agent says it's waiting; out-of-context recovery explicitly defers to it.
  • account-switch-contract.md — rate-limit / account-switch recovery, the other turn-interruption path.
  • context-stall-recovery-blinded-on-compaction-postmortem.md — the 2026-07-13 "turn ended without a response" loop where recovery went blind on compaction turns, and the knownOutOfRoom fix.

Last verified 2026-09-23