---
title: Relentless session relaunch (keep retrying a launch hiccup until the session is back)
---

# Relentless session relaunch (keep retrying a launch hiccup until the session is back)

## What it is

When Omniscio starts a session — a brand-new one, a restart, or an auto-resume after a crash — it spawns a Claude CLI child process. On a heavily-loaded computer that spawn can **fail transiently**: Windows refuses to create the process for a moment (the machine is out of process/handle headroom, antivirus is mid-scan, the box is swap-thrashing). This is not a broken setup — the binary path is fine, the account is fine — the machine was just momentarily too busy to start one more process.

**Relentless relaunch** makes a session that was running **fight its way back** from that kind of hiccup instead of giving up. With the setting on (the default), Omniscio keeps re-attempting the launch — with a gap between tries that grows the longer the machine stays jammed — until the session comes back. You don't have to notice it, click Retry, or restart anything.

This matters most when **many sessions come back at once** (a mass restart, or auto-resume after a crash with a dozen sessions). That burst is exactly what overloads the OS and makes the individual launches hiccup — and it's exactly when "give up after 3 quick tries" used to strand a pile of sessions as red errors. (That was the real incident: a morning restart hit `errno -4094` across ~25 sessions, each retried 3 times in under a second, then went red.)

## Where to find it

### How to turn it off

**Settings → Workflow → "Keep Relaunching a Session Until It's Back"** — a toggle, default **on**. Helper text: _"When a session that was running fails to launch because the machine is momentarily overloaded (a transient OS hiccup), keep retrying — with a growing gap between tries — until it comes back, instead of giving up after a few attempts. Genuine errors still stop and turn red."_

Turn it off if you'd rather a session that can't launch after a few quick tries show up as a red error for manual triage. With it off, Omniscio uses the original behavior exactly: at most 3 short retries, then a red error.

The setting key is `relentlessSessionRelaunch` (boolean). There's also a launch-time kill switch for diagnostics: `AMC_DISABLE_RELENTLESS_RELAUNCH=1` forces the original bounded give-up regardless of the toggle.

## How it behaves

### When it kicks in

Relentless relaunch only applies to a **transient OS launch failure** — the machine-too-busy kind. Under the hood these are the `errno -4094` (Windows `CreateProcess` "UNKNOWN"), `EMFILE`, `ENFILE`, `EAGAIN`, and `EBUSY` failures. A launch that fails for a **real** reason — a missing/ moved binary, an authentication error, a genuine crash — is **not** transient: it still surfaces immediately as a red error, exactly as before. Omniscio never hides a real problem behind an endless retry.

### The growing gap (why it doesn't make things worse)

A relentless retry is **paced**, not a hot loop. The first few tries are fast (a quarter-second, then doubling), but the gap keeps growing toward about **30 seconds** the longer the machine stays saturated. So a permanently-overloaded box is retried **patiently** — roughly once every half-minute — rather than being hammered, which would only deepen the very overload that caused the failure. There's also a small ceiling on how many sessions retry at the exact same instant (a throttle): if that's momentarily full, a waiting session just holds its place and re-attempts after a backoff instead of failing. The session is never abandoned; it's only ever slowed down.

### What you see

The session stays in its "coming back" state (it does **not** flash red) while it retries, and Omniscio keeps you informed without flooding the chat:

- **Once, at the first hiccup**: a system line — _"Launch hiccup — the machine looks busy. Keeping at it until this session is back…"_
- **Then only every 10th attempt** (if it's really taking a while): _"Still relaunching this session… (attempt N)"_

So a long, patient retry looks **alive**, not frozen — but a slightly slow recovery doesn't spam your transcript with a line per try.

### What does NOT get relentlessly retried

- **Genuine errors** — auth failures, a real crash, a missing binary. These are not transient launch hiccups, so they surface as a red error right away for you to look at. Relentless relaunch is specifically about riding out a _busy machine_, not papering over a real fault.
- The behavior is also **scoped to the transient launch path** — it does not change crash-recovery's "which sessions auto-resume" rules, the stuck-turn/aborted-response recovery, or any other give-up. See [crash recovery](crash-recovery.md) and [aborted-response recovery](aborted-response-recovery.md) for those.

## For agents

### Where it lives in code

For agents working in the repo:

- **Classifier + backoff:** [src/main/process/transient-spawn-retry.ts](../../src/main/process/transient-spawn-retry.ts) — `isTransientSpawnError()` decides what counts as transient; `relentlessRetryDelayMs()` is the escalating backoff capped at `RELENTLESS_RETRY_MAX_DELAY_MS` (30s); `tryAcquireSpawnRetrySlot()` / `MAX_CONCURRENT_SPAWN_RETRIES` is the concurrency throttle.
- **The gate:** `maybeRetryTransientSpawn()` in [src/main/process/process-manager.ts](../../src/main/process/process-manager.ts) reads `relentlessSessionRelaunch` (default `true`) + the env kill switch. When on, the per-session `MAX_SPAWN_RETRIES` cap is bypassed and a full retry budget DEFERS via `scheduleRelentlessRetryWait()` instead of going terminal; `emitSpawnRetryStatus()` is the announce-once-then-every-10th status. When off, the path is byte-identical to the original bounded retry.
- **Watchdog interaction:** a session waiting on a relentless retry has `spawnRetryPending = true`, which the liveness watchdog skips ([src/main/process/liveness-watchdog.ts](../../src/main/process/liveness-watchdog.ts)) — so a long patient retry is never falsely flipped to error mid-wait.
- **Setting wiring:** type + default in [src/shared/types/settings/session-behavior-settings.ts](../../src/shared/types/settings/session-behavior-settings.ts), Zod in [src/shared/ipc-schemas/settings/session-behavior-settings.ts](../../src/shared/ipc-schemas/settings/session-behavior-settings.ts), UI toggle in [src/renderer/src/features/settings/WorkflowSettings-definitions.ts](../../src/renderer/src/features/settings/sections/workflow/WorkflowSettings-definitions.ts).
- **Invariants:** the [transient-spawn-retry contract](../../.claude/memory/contracts/transient-spawn-retry-contract.md) (A1-I3 conditional + A1-I8) names what a future change must not break and the tests that catch it.

## Related

This covers only the transient launch hiccup. If what you are looking at is a session that came back wrong after the app restarted, read [crash-recovery.md](crash-recovery.md); if it is a turn that died mid-answer, read [aborted-response-recovery.md](aborted-response-recovery.md). Those two own the decisions about which sessions come back at all, which this feature deliberately does not touch.
