---
title: CLI Session Recovery — Get Stuck Sessions Going Again
---

# CLI Session Recovery — Get Stuck Sessions Going Again

## What it is

> An AI agent holding the Omniscio bearer token can diagnose and fix a wedged Claude session directly over HTTP — no user approval round-trip. Three targeted endpoints cover the three failure modes: account bad, process dead, session idle.

### The problem

A Claude session in Omniscio can get stuck in several ways:

- The Anthropic account it was on hit a rate limit or got frozen (auth error, quota exhausted).
- An account switch failed mid-recovery and the session is pointing at a bad account.
- The Claude Code process died (crash, OOM, signal) and needs to be restarted with full history.
- The session is technically alive but has gone idle and just needs a push to continue.

All three are fixable over the CLI control server at `127.0.0.1:19519` with the Omniscio bearer token (auto-delivered to `~/.amc/cli-token` for the `omniscio-control` skill; see [omniscio-control.md](omniscio-control.md) and [cli-control.md](cli-control.md) for connection details). No approval queue — the three mutating routes (plus the read-only `GET /accounts` discovery helper) apply immediately.

## Where to find it

There is no screen for this. CLI Session Recovery exists only for an AI agent, which reaches it over the same local control server the rest of Omniscio's automation is driven from; nothing appears in the app for a person to click. The fixing calls apply the moment they return rather than waiting in an approval queue, so a wedged session can be put right without stopping to ask. What follows is therefore what the agent needs: the token it authenticates with, the read-only call that finds a healthy account to target, and the three calls that fix a session.

### Fetch the bearer token

```bash
TOKEN=$(powershell.exe -NoProfile -File "$HOME/.claude/secrets/get-secret.ps1" amc-cli | tr -d '\r\n')
```

## How it behaves

Recovery is immediate by design: there is no approval card, so a fix takes effect as soon as the call comes back, and every call is deliberate about which single problem it addresses. The read-only account list is there to help you choose a target, and the three fixing calls each cover exactly one failure mode — a bad account, a dead process, or a session that is merely idle. One call can outlast its own answer: restarting or moving a session has to spawn a Claude process, and when the machine is busy that work is accepted and continues in the background, which the response says plainly rather than reporting a failure.

## For agents

### Step 1 — List accounts to find a healthy target

Before moving a session, find an account that is healthy and not overloaded.

**`GET /accounts`** — returns every account in the pool with its health, breaker state, live-session count, and 5-hour utilization. No credentials are ever returned.

```bash
curl -s -H "Authorization: Bearer $TOKEN" http://127.0.0.1:19519/accounts
```

Response shape:

```json
{
  "ok": true,
  "data": {
    "accounts": [
      {
        "id": "acct_abc123",
        "label": "work@example.com",
        "type": "login",
        "isActive": true,
        "isSpawnFallback": false,
        "health": "ok",
        "breakerOpen": false,
        "usage": { "utilizationPct": 34, "resetAt": "2026-06-09T18:00:00.000Z" },
        "liveSessionCount": 3
      },
      {
        "id": "acct_def456",
        "label": "backup-key",
        "type": "apikey",
        "isActive": false,
        "isSpawnFallback": true,
        "health": "rate_limit",
        "breakerOpen": true,
        "usage": null,
        "liveSessionCount": 0
      }
    ]
  }
}
```

**Field reference:**

| Field                  | Values                                               | Meaning                                                                   |
| ---------------------- | ---------------------------------------------------- | ------------------------------------------------------------------------- |
| `health`               | `'ok'` \| `'auth'` \| `'rate_limit'` \| `'degraded'` | Why the breaker is open (or `'ok'` if it is not)                          |
| `breakerOpen`          | boolean                                              | `true` means Omniscio is already avoiding this account for new spawns     |
| `usage.utilizationPct` | 0–100 or `null`                                      | Fraction of the 5-hour token window consumed; `null` if no usage data yet |
| `usage.resetAt`        | ISO-8601 or `null`                                   | When the 5-hour window resets                                             |
| `liveSessionCount`     | integer                                              | Number of sessions currently running on this account                      |
| `isSpawnFallback`      | boolean                                              | Account is the designated last-resort fallback when all others are busy   |

A **good move target** is one where `health === 'ok'` and `breakerOpen === false`. Prefer lower `liveSessionCount` and `utilizationPct` when multiple healthy accounts exist.

### Step 2 — Fix the session

#### Which endpoint to use

| Situation                                    | Use                                                   |
| -------------------------------------------- | ----------------------------------------------------- |
| Account hit rate limit / auth error / frozen | `POST /session/:id/move-account` to a healthy account |
| Process crashed or died, account is fine     | `POST /session/:id/restart`                           |
| Session is alive-but-idle, needs a continue  | `POST /session/:id/nudge`                             |

---

#### `POST /session/:id/restart`

Kills the current Claude Code process and relaunches it with `--resume` so the full conversation history is preserved. Use this when the process wedged (crash, OOM, lost contact) but the account the session was on is still healthy.

No request body.

```bash
curl -s -X POST \
  -H "Authorization: Bearer $TOKEN" \
  http://127.0.0.1:19519/session/<session-uuid>/restart
```

**Success (200):**

```json
{ "ok": true, "data": { "sessionId": "<uuid>", "status": "starting" } }
```

**Errors:**
| Status | When |
| ------ | ---- |
| 401 | Missing or wrong bearer token |
| 404 | Session not found |
| 409 | A restart or terminate is already in flight for this session |
| 504 | The restart is still starting — see below. **Not a failure.** |
| 500 | Relaunch failed (detail in `error`) |

##### A 504 here does NOT mean the restart failed

Restarting a session means spawning a Claude Code process, and Omniscio deliberately paces
spawns so a burst of them can't freeze the computer. When the machine is busy that queue can
run long, so the route now answers within about 10 seconds either way:

```json
{
  "ok": false,
  "code": "REQUEST_TIMEOUT",
  "error": "The restart was accepted and is still starting in the background — ..."
}
```

The restart **was accepted and is still running**. What to do:

- **Do not re-issue it** — a second restart would kill and respawn the session again. (Retry
  anyway and you'll get a `409` while the first one is still going, which is also not a failure.)
- **Poll `GET /session/:id/status`** to see when it comes back.
- Send an `X-Client-Request-Id` header on the original call if you want a retry to replay the
  first result instead of starting a second restart.

Before this, a restart on a busy box simply never answered and the connection died, which read
like the whole server was down. An operator who needs a longer wait can set
`AMC_CLI_SPAWN_ROUTE_DEADLINE_MS`.

---

#### `POST /session/:id/nudge`

Wakes the session from any status (ended, paused, archived, error) and injects a "keep going" turn. Best when the session is alive but idle, or you want to resume it with a specific instruction.

**Body** (all optional — omit entirely for the default "Please continue." nudge):

```json
{ "text": "The migration script finished. Please run the smoke tests now." }
```

```bash
# Default nudge
curl -s -X POST \
  -H "Authorization: Bearer $TOKEN" \
  -H "Content-Type: application/json" \
  http://127.0.0.1:19519/session/<session-uuid>/nudge

# Nudge with custom text
curl -s -X POST \
  -H "Authorization: Bearer $TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"text": "Tests are passing now. Please continue with the refactor."}' \
  http://127.0.0.1:19519/session/<session-uuid>/nudge
```

**Success (200):**

```json
{ "ok": true, "data": { "sessionId": "<uuid>", "delivered": true, "wokenFrom": "paused" } }
```

`wokenFrom` is the prior status the session was woken from (`"archived"`, `"ended"`, `"error"`) or
`null` if it was already running **or if the target was PAUSED and the message was HELD instead**
(2026-09-28): a nudge is machinery, and a pause is a person's own command, so the nudge goes into
that session's own queue and arrives when a person unpauses it. A held nudge answers
`{ "delivered": false, "queued": true, "wokenFrom": null }` — nothing is lost, and re-sending would
deliver it twice.

**Errors:**
| Status | When |
| ------ | ---- |
| 400 | Body parse failure (e.g. `text` exceeds 32 000 chars) |
| 401 | Missing or wrong bearer token |
| 404 | Session not found |
| 409 | A turn is already in flight — wait and retry |
| 429 | Rate limit hit (120 calls/hour, shared with `peer-message`) — `retryAfterMs` in response |
| 500 | Wake or send failed (detail in `error`) |

Nudge shares the cross-session-messaging rate-limit bucket (120 calls per bearer token per hour). See [cross-session-messaging.md](cross-session-messaging.md).

---

#### `POST /session/:id/move-account`

Moves a stuck **Claude** session onto a specific account and resumes it there. This is the fix for "account frozen / switch failed" — pick a healthy account from `GET /accounts` and route the session to it.

The named account is honored even if its breaker is open (operator intent overrides health-avoidance), so consult `GET /accounts` first and pick wisely. The response's `targetHealth` shows the health of the account you chose.

Non-Claude sessions (Codex, external engines) return `400` — those providers don't use Anthropic accounts.

**Body (required):**

```json
{ "accountId": "acct_abc123" }
```

```bash
curl -s -X POST \
  -H "Authorization: Bearer $TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"accountId": "acct_abc123"}' \
  http://127.0.0.1:19519/session/<session-uuid>/move-account
```

**Success (200):**

```json
{
  "ok": true,
  "data": {
    "sessionId": "<uuid>",
    "movedTo": "acct_abc123",
    "targetHealth": "ok",
    "wokenFrom": "error"
  }
}
```

`wokenFrom` mirrors the nudge semantics — the prior status the session was in before it was relaunched on the new account.

**Errors:**
| Status | When |
| ------ | ---- |
| 400 | Missing/invalid `accountId`, or session is not a Claude session |
| 401 | Missing or wrong bearer token |
| 404 | Session not found |
| 409 | A turn is already in flight |
| 429 | Mutation rate limit hit |
| 504 | The move is still resuming in the background — same meaning as the restart 504 above. **Not a failure**; the session is already pinned to the account. Do not re-issue; poll `GET /session/:id/status`. |
| 500 | Resume failed (detail in `error`) |

### Full recovery example

```bash
# 1. Find a healthy account
ACCOUNTS=$(curl -s -H "Authorization: Bearer $TOKEN" http://127.0.0.1:19519/accounts)
echo "$ACCOUNTS" | jq '.data.accounts[] | select(.health == "ok" and .breakerOpen == false)'

# 2. Move the stuck session to the healthy account
curl -s -X POST \
  -H "Authorization: Bearer $TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"accountId": "acct_abc123"}' \
  http://127.0.0.1:19519/session/550e8400-e29b-41d4-a716-446655440000/move-account
```

### Notes

- The three mutating routes (plus the read-only `GET /accounts` discovery helper) apply **immediately** — no approval inbox row is created.
- All require the Omniscio bearer token (`Authorization: Bearer <token>`).
- `GET /accounts` is read-only (high read-budget rate limit). The three POST routes share the standard mutation rate limit.
- The account breaker (`AMC_DISABLE_ACCOUNT_BREAKER=1` to disable) normally prevents spawns onto degraded accounts. `move-account` intentionally bypasses it — the operator chose the account explicitly.

## Related

The approval queue these three calls deliberately skip is described on the [CLI pending actions](cli-pending-actions.md) page. How Omniscio spreads sessions across accounts by itself — so that a stuck session is rarer in the first place — is on [Distribute sessions across accounts](distribute-sessions-across-accounts.md), and messaging another session rather than reviving it uses the endpoint covered by [Cross-session messaging](cross-session-messaging.md).

### See also

- [cli-server-gating.md](../../.claude/memory/cli-server-gating.md) — full gating rules for every CLI endpoint
- [cross-session-messaging.md](cross-session-messaging.md) — the `peer-message` endpoint that shares the 120/hr nudge bucket
- [distribute-sessions-across-accounts.md](distribute-sessions-across-accounts.md) — how Omniscio normally spreads sessions across accounts automatically
- [omniscio-control.md](omniscio-control.md) — the bundled skill that wraps the CLI control server for agents
