---
title: Usage Cascade (what to run when your Claude subscription is out)
---

# Usage Cascade (what to run when your Claude subscription is out)

## What it is

When **every** Claude account in your pool has hit its usage limit, Omniscio's normal behaviour is to
stop: affected sessions park in a **waiting** state, a banner appears in the sessions sidebar, and
everything resumes on its own when your subscription window resets. That is the right default —
it never spends money you did not ask it to — but it means work stops, sometimes for hours.

The **usage cascade** lets you say, in advance and in your own order, what should happen instead.
You build a short ladder. Each rung names one other AI vendor and how many sessions may use it at
the same time:

| #   | Fall back to                                                       | How many at once     |
| --- | ------------------------------------------------------------------ | -------------------- |
| 1   | GLM                                                                | No limit _(default)_ |
| 2   | DeepSeek                                                           | At most 5            |
| —   | _wait for your subscription to reset_ (always last, not removable) | —                    |

When the subscription runs dry, each affected session walks that ladder from the top and takes the
first rung that is **switched on**, **set up** (its credential is present), and **not already
full**. If no rung qualifies, the session waits exactly as it does today.

"Runs dry" means **every** one of your live Claude accounts is genuinely at its limit. If one
account is merely busy and another still has room, Omniscio just moves the session to the account
with room, exactly as it always has — the cascade does not fire, and nothing is billed.

### It also catches a vendor running out

The same ladder is used when a session that is **already running on one of these vendors** runs
out on it — GLM hits its usage limit, say, or your DeepSeek balance hits zero. Instead of waiting
for that vendor to reset, the session walks your ladder from the top, skipping the vendor it is
leaving (and any vendor it already ran out on), and carries on at the first rung that qualifies.

For example, with just **DeepSeek** on your ladder, a GLM session that hits GLM's usage limit keeps
going on DeepSeek (Flash, DeepSeek's default model) with its conversation intact, and re-runs the
step that was interrupted.

A few details:

- **The vendor's own backups come first.** A second GLM key and every other row of the model
  family's supply list (Settings → Accounts → Who pays & who serves — your own account, a reseller,
  Omniscio credits) are used before a session leaves its vendor.
- **No bouncing.** A session never walks back onto a vendor it already ran out on in the same
  stretch, so it moves at most once per rung and then waits. After its next successful turn it
  forgets that list, so a later run-out can use a vendor that has since recovered.
- **A session started on a vendor stays where it moved.** Nothing can tell when a vendor's limit has
  reset, so it is not sent back automatically; new sessions still start on the vendor you chose.
- **Existing ladders pick this up** — and an empty ladder still changes nothing.

**It ships empty.** Until you add a rung, nothing changes at all.

## Where to find it

**Settings → Accounts → Usage cascade.** The card shows your ladder in order:

- **Add a fallback…** — a dropdown of the vendors you can use. Only vendors that run through
  Omniscio's own Claude process are offered (DeepSeek, Kimi, GLM, MiniMax and the like), because
  only those can hand the conversation back to Claude afterwards. A reseller such as DeepInfra is
  not offered as a rung: it is a row inside a model family's supply list, and an older rung that
  named one now runs on that model's family, whose list decides who serves it. Engines with their own
  separate session manager — Codex, Gemini, Devin, Cursor and friends — are deliberately not
  offerable: a session placed on one could never come home.
- **No limit / Limit how many at once** — a small switch per rung, **off by default**. Leave it off
  and the fallback runs as many sessions as it needs. Switch it on and a number box appears, capping
  how many sessions may sit on that one fallback simultaneously.
- **An on/off switch** per rung — turn a fallback off without losing its place in your order.
- **Up / down arrows** — reorder. Order is the whole point: the first rung that works wins.
- **A trash icon** — remove the rung entirely.
- **A "Bills per token" label** on any vendor that charges per use, so a paid rung is never a
  surprise.

The final rung — _wait for your subscription to reset_ — is shown as a non-editable row so the full
behaviour is visible rather than implied.

## How it behaves

### What happens to a session that falls back

1. The session starts on the fallback vendor and **keeps its conversation** — the transcript and
   context carry over, because these vendors run through the same underlying CLI.
2. It says so once in its own transcript, naming the vendor it moved to and why.
3. The next time that session starts, if your Claude subscription has room again, it **moves back
   to Claude automatically**, restored to the exact model it was using before.

You do not have to do anything to bring sessions home, and nothing is left stranded on a paid
vendor after your subscription recovers. (A session you started on a vendor — GLM, say — is the one
exception: it stays on the rung it moved to, as described above.)

### Why there is no "cheaper Claude model" rung

An earlier design offered "drop to a cheaper Claude model" as the first rung. It was removed
deliberately, because it cannot work at the moment it would be consulted.

Omniscio decides the pool is exhausted by looking at your **overall 5-hour and 7-day usage
windows** — not at per-model limits. So by the time the cascade runs, a cheaper Claude model would
be asking the same account that is already at its cap, and would simply fail again. Worse, Claude
always counts as "set up", so the rung could never be skipped: it would be taken every time,
announce itself, and deliver nothing.

If you want Omniscio to use a cheaper Claude model when usage gets tight, that is a **separate,
already-shipped feature** that fires _earlier_ — when you approach a threshold, rather than after
you hit the wall. See the rate-limit model downgrade settings.

### The promises this feature makes

- **Nothing happens until you ask for it.** An empty ladder behaves exactly like today.
- **It only fires when a lane is genuinely out** — every Claude account at its limit, or the vendor
  a session runs on refusing it for its usage limit or an empty balance. One Claude account running
  out still just switches accounts, as it always has.
- **A vendor's own backups come first**, and **a session never bounces** between two vendors that
  are both out.
- **A cap is opt-in and means "at most this many at once" — never "run fewer sessions".** A rung
  starts with no limit; you switch a cap on if you want one. It never limits your Claude sessions.
- **Money is never spent by surprise.** A rung that bills per token is labelled where you configure
  it, and can only ever be reached because you put it there.
- **Every session that left Claude comes home** once your subscription has capacity again.
- **A full or broken rung is skipped silently**, never an error. Nothing fails because a fallback
  was unavailable.
- **You can always see it happened** — a move is announced in that session's transcript.

### Turning it off

Two ways:

- **Clear the ladder** (or switch every rung off) in Settings. The feature becomes inert.
- **Set `AMC_DISABLE_USAGE_CASCADE=1`** in the environment. This restores today's behaviour exactly
  while leaving your ladder saved, so you can switch it back on without reconfiguring.

There is also a second, inherited switch worth knowing about: the cascade lives inside the
capacity-park route, so `AMC_DISABLE_CAPACITY_PARK=1` turns it off too, one layer up.

## For agents

### How it works under the hood

Everything happens at **one decision point**: the moment a session tries to start and cannot get a
Claude credential because every account is capped.

- **Going out** — `decideEmptyCredsRoute`
  ([src/shared/account-pool-exhaustion.ts](../../src/shared/account-pool-exhaustion.ts)) gains a
  `cascade` route. The spawn guard
  ([src/main/process/spawn-guard-router.ts](../../src/main/process/spawn-guard-router.ts)) stamps
  the session onto the chosen vendor and relaunches it there. If nothing resolves, it parks exactly
  as before.
- **Coming home** — just before a session respawns
  ([session-relaunch-service.ts](../../src/main/services/session/session-relaunch-service.ts) and
  [session-auto-resume.ts](../../src/main/services/session/session-auto-resume.ts)), if it is on a
  fallback and Claude has capacity, its original vendor and model are restored first.

There is deliberately **no background sweep** in either direction. Each session decides for itself,
at the only moment it can act, which is why a crash or restart can never leave one stranded.

The **vendor trigger** has its own decision point: the end of the turn a vendor refused. The two
turn-end vendor gates
([rate-limit-turn-end.ts](../../src/main/process/turn-end-error-gates/rate-limit-turn-end.ts),
[vendor-insufficient-balance.ts](../../src/main/process/turn-end-error-gates/vendor-insufficient-balance.ts))
ask `usageCascadeWalkFor`
([recovery-helpers.ts](../../src/main/process/turn-end-error-gates/recovery-helpers.ts)) after the
vendor's own fallbacks. With no candidate rung they park exactly as before, in the same tick;
otherwise the walk resolves a rung (excluding the lanes this session ran out on), re-checks that no
new turn started meanwhile, moves the session in one statement (`stepSessionOntoCascadeLane`), and
re-runs the interrupted turn there.

**One rung's re-run never blocks the next rung's.** Every one of these re-runs goes through the same
door (`replayLastOperatorMessage`), which holds a short cross-path window so that two recoveries
reacting to the *same* failure cannot both re-run it. That window is scoped to the failure, not to
the clock: once the turn a re-run started has **come back with a result** — the usual shape when a
new lane refuses within seconds — the next fallback is free to re-run it. Before 2026-09-25 the
window was time-only, so a ladder that moved a session onto a lane which then refused immediately
stopped there and parked in **Needs You** with its next source sitting ready. It now carries on down
the ladder. The shared per-turn replay budget still caps the whole chain, so a lane that keeps
failing cannot spend without limit.

- The **ordering rules** are a pure function, `planUsageCascade`
  ([src/shared/providers/usage-cascade.ts](../../src/shared/providers/usage-cascade.ts)) — it does
  no I/O, so it is fully testable.
- The **rung it picks** is resolved by
  [usage-cascade-resolver.ts](../../src/main/services/providers/usage-cascade-resolver.ts), which
  checks each vendor's readiness and how full each lane is.
- **Which lane a session is on** is recorded on the session row itself, together with the way home.
  That record is also what counts toward each rung's cap, so a session that is waiting still holds
  its slot.

Full invariants and rationale: [usage-cascade-contract.md](../../.claude/memory/contracts/usage-cascade-contract.md).

## Related

- [account-pool.md](account-pool.md) — how Omniscio picks between multiple Claude logins, and the
  exhaustion banner this feature takes over from.
- [distribute-sessions-across-accounts.md](distribute-sessions-across-accounts.md) — spreading
  sessions across the Claude accounts you already have, which happens _before_ any of this.
