CLI Session Recovery — Get Stuck Sessions Going Again
Three targeted CLI calls let an AI agent diagnose and fix a wedged Claude session with no approval round-trip: moving a session off an account that hit a rate limit or auth error, restarting a process that died, and nudging a session that is alive but idle. Covers finding a healthy account and what each call returns when it goes wrong.
What it is
An AI agent holding the Omniscio bearer token can diagnose and fix a wedged Claude session directly over HTTP — no user approval round-trip. Three targeted endpoints cover the three failure modes: account bad, process dead, session idle.
The problem
A Claude session in Omniscio can get stuck in several ways:
- The Anthropic account it was on hit a rate limit or got frozen (auth error, quota exhausted).
- An account switch failed mid-recovery and the session is pointing at a bad account.
- The Claude Code process died (crash, OOM, signal) and needs to be restarted with full history.
- The session is technically alive but has gone idle and just needs a push to continue.
All three are fixable over the CLI control server at 127.0.0.1:19519 with the Omniscio bearer token (auto-delivered to ~/.amc/cli-token for the omniscio-control skill; see omniscio-control.md and cli-control.md for connection details). No approval queue — the three mutating routes (plus the read-only GET /accounts discovery helper) apply immediately.
Where to find it
There is no screen for this. CLI Session Recovery exists only for an AI agent, which reaches it over the same local control server the rest of Omniscio's automation is driven from; nothing appears in the app for a person to click. The fixing calls apply the moment they return rather than waiting in an approval queue, so a wedged session can be put right without stopping to ask. What follows is therefore what the agent needs: the token it authenticates with, the read-only call that finds a healthy account to target, and the three calls that fix a session.
Fetch the bearer token
TOKEN=$(powershell.exe -NoProfile -File "$HOME/.claude/secrets/get-secret.ps1" amc-cli | tr -d '\r\n')
How it behaves
Recovery is immediate by design: there is no approval card, so a fix takes effect as soon as the call comes back, and every call is deliberate about which single problem it addresses. The read-only account list is there to help you choose a target, and the three fixing calls each cover exactly one failure mode — a bad account, a dead process, or a session that is merely idle. One call can outlast its own answer: restarting or moving a session has to spawn a Claude process, and when the machine is busy that work is accepted and continues in the background, which the response says plainly rather than reporting a failure.
For agents
Step 1 — List accounts to find a healthy target
Before moving a session, find an account that is healthy and not overloaded.
GET /accounts — returns every account in the pool with its health, breaker state, live-session count, and 5-hour utilization. No credentials are ever returned.
curl -s -H "Authorization: Bearer $TOKEN" http://127.0.0.1:19519/accounts
Response shape:
{
"ok": true,
"data": {
"accounts": [
{
"id": "acct_abc123",
"label": "work@example.com",
"type": "login",
"isActive": true,
"isSpawnFallback": false,
"health": "ok",
"breakerOpen": false,
"usage": { "utilizationPct": 34, "resetAt": "2026-06-09T18:00:00.000Z" },
"liveSessionCount": 3
},
{
"id": "acct_def456",
"label": "backup-key",
"type": "apikey",
"isActive": false,
"isSpawnFallback": true,
"health": "rate_limit",
"breakerOpen": true,
"usage": null,
"liveSessionCount": 0
}
]
}
}
Field reference:
| Field | Values | Meaning |
|---|---|---|
health |
'ok' | 'auth' | 'rate_limit' | 'degraded' |
Why the breaker is open (or 'ok' if it is not) |
breakerOpen |
boolean | true means Omniscio is already avoiding this account for new spawns |
usage.utilizationPct |
0–100 or null |
Fraction of the 5-hour token window consumed; null if no usage data yet |
usage.resetAt |
ISO-8601 or null |
When the 5-hour window resets |
liveSessionCount |
integer | Number of sessions currently running on this account |
isSpawnFallback |
boolean | Account is the designated last-resort fallback when all others are busy |
A good move target is one where health === 'ok' and breakerOpen === false. Prefer lower liveSessionCount and utilizationPct when multiple healthy accounts exist.
Step 2 — Fix the session
Which endpoint to use
| Situation | Use |
|---|---|
| Account hit rate limit / auth error / frozen | POST /session/:id/move-account to a healthy account |
| Process crashed or died, account is fine | POST /session/:id/restart |
| Session is alive-but-idle, needs a continue | POST /session/:id/nudge |
POST /session/:id/restart
Kills the current Claude Code process and relaunches it with --resume so the full conversation history is preserved. Use this when the process wedged (crash, OOM, lost contact) but the account the session was on is still healthy.
No request body.
curl -s -X POST \
-H "Authorization: Bearer $TOKEN" \
http://127.0.0.1:19519/session/<session-uuid>/restart
Success (200):
{ "ok": true, "data": { "sessionId": "<uuid>", "status": "starting" } }
Errors:
| Status | When |
|---|---|
| 401 | Missing or wrong bearer token |
| 404 | Session not found |
| 409 | A restart or terminate is already in flight for this session |
| 504 | The restart is still starting — see below. Not a failure. |
| 500 | Relaunch failed (detail in error) |
A 504 here does NOT mean the restart failed
Restarting a session means spawning a Claude Code process, and Omniscio deliberately paces spawns so a burst of them can't freeze the computer. When the machine is busy that queue can run long, so the route now answers within about 10 seconds either way:
{
"ok": false,
"code": "REQUEST_TIMEOUT",
"error": "The restart was accepted and is still starting in the background — ..."
}
The restart was accepted and is still running. What to do:
- Do not re-issue it — a second restart would kill and respawn the session again. (Retry
anyway and you'll get a
409while the first one is still going, which is also not a failure.) - Poll
GET /session/:id/statusto see when it comes back. - Send an
X-Client-Request-Idheader on the original call if you want a retry to replay the first result instead of starting a second restart.
Before this, a restart on a busy box simply never answered and the connection died, which read
like the whole server was down. An operator who needs a longer wait can set
AMC_CLI_SPAWN_ROUTE_DEADLINE_MS.
POST /session/:id/nudge
Wakes the session from any status (ended, paused, archived, error) and injects a "keep going" turn. Best when the session is alive but idle, or you want to resume it with a specific instruction.
Body (all optional — omit entirely for the default "Please continue." nudge):
{ "text": "The migration script finished. Please run the smoke tests now." }
# Default nudge
curl -s -X POST \
-H "Authorization: Bearer $TOKEN" \
-H "Content-Type: application/json" \
http://127.0.0.1:19519/session/<session-uuid>/nudge
# Nudge with custom text
curl -s -X POST \
-H "Authorization: Bearer $TOKEN" \
-H "Content-Type: application/json" \
-d '{"text": "Tests are passing now. Please continue with the refactor."}' \
http://127.0.0.1:19519/session/<session-uuid>/nudge
Success (200):
{ "ok": true, "data": { "sessionId": "<uuid>", "delivered": true, "wokenFrom": "paused" } }
wokenFrom is the prior status the session was woken from ("archived", "ended", "error") or
null if it was already running or if the target was PAUSED and the message was HELD instead
(2026-09-28): a nudge is machinery, and a pause is a person's own command, so the nudge goes into
that session's own queue and arrives when a person unpauses it. A held nudge answers
{ "delivered": false, "queued": true, "wokenFrom": null } — nothing is lost, and re-sending would
deliver it twice.
Errors:
| Status | When |
|---|---|
| 400 | Body parse failure (e.g. text exceeds 32 000 chars) |
| 401 | Missing or wrong bearer token |
| 404 | Session not found |
| 409 | A turn is already in flight — wait and retry |
| 429 | Rate limit hit (120 calls/hour, shared with peer-message) — retryAfterMs in response |
| 500 | Wake or send failed (detail in error) |
Nudge shares the cross-session-messaging rate-limit bucket (120 calls per bearer token per hour). See cross-session-messaging.md.
POST /session/:id/move-account
Moves a stuck Claude session onto a specific account and resumes it there. This is the fix for "account frozen / switch failed" — pick a healthy account from GET /accounts and route the session to it.
The named account is honored even if its breaker is open (operator intent overrides health-avoidance), so consult GET /accounts first and pick wisely. The response's targetHealth shows the health of the account you chose.
Non-Claude sessions (Codex, external engines) return 400 — those providers don't use Anthropic accounts.
Body (required):
{ "accountId": "acct_abc123" }
curl -s -X POST \
-H "Authorization: Bearer $TOKEN" \
-H "Content-Type: application/json" \
-d '{"accountId": "acct_abc123"}' \
http://127.0.0.1:19519/session/<session-uuid>/move-account
Success (200):
{
"ok": true,
"data": {
"sessionId": "<uuid>",
"movedTo": "acct_abc123",
"targetHealth": "ok",
"wokenFrom": "error"
}
}
wokenFrom mirrors the nudge semantics — the prior status the session was in before it was relaunched on the new account.
Errors:
| Status | When |
|---|---|
| 400 | Missing/invalid accountId, or session is not a Claude session |
| 401 | Missing or wrong bearer token |
| 404 | Session not found |
| 409 | A turn is already in flight |
| 429 | Mutation rate limit hit |
| 504 | The move is still resuming in the background — same meaning as the restart 504 above. Not a failure; the session is already pinned to the account. Do not re-issue; poll GET /session/:id/status. |
| 500 | Resume failed (detail in error) |
Full recovery example
# 1. Find a healthy account
ACCOUNTS=$(curl -s -H "Authorization: Bearer $TOKEN" http://127.0.0.1:19519/accounts)
echo "$ACCOUNTS" | jq '.data.accounts[] | select(.health == "ok" and .breakerOpen == false)'
# 2. Move the stuck session to the healthy account
curl -s -X POST \
-H "Authorization: Bearer $TOKEN" \
-H "Content-Type: application/json" \
-d '{"accountId": "acct_abc123"}' \
http://127.0.0.1:19519/session/550e8400-e29b-41d4-a716-446655440000/move-account
Notes
- The three mutating routes (plus the read-only
GET /accountsdiscovery helper) apply immediately — no approval inbox row is created. - All require the Omniscio bearer token (
Authorization: Bearer <token>). GET /accountsis read-only (high read-budget rate limit). The three POST routes share the standard mutation rate limit.- The account breaker (
AMC_DISABLE_ACCOUNT_BREAKER=1to disable) normally prevents spawns onto degraded accounts.move-accountintentionally bypasses it — the operator chose the account explicitly.
Related
The approval queue these three calls deliberately skip is described on the CLI pending actions page. How Omniscio spreads sessions across accounts by itself — so that a stuck session is rarer in the first place — is on Distribute sessions across accounts, and messaging another session rather than reviving it uses the endpoint covered by Cross-session messaging.
See also
- cli-server-gating.md — full gating rules for every CLI endpoint
- cross-session-messaging.md — the
peer-messageendpoint that shares the 120/hr nudge bucket - distribute-sessions-across-accounts.md — how Omniscio normally spreads sessions across accounts automatically
- omniscio-control.md — the bundled skill that wraps the CLI control server for agents
Last verified 2026-09-28