Usage Cascade (what to run when your Claude subscription is out)
The usage cascade is a short ordered ladder of other AI vendors that Omniscio falls back to when every Claude account in your pool has genuinely run out, or when the vendor a session is running on (GLM, DeepSeek, Kimi, MiniMax) hits its own usage limit or runs out of balance, instead of parking sessions in a waiting state for hours. You build it in Settings, it ships empty, and a session that left Claude comes home once capacity returns.
What it is
When every Claude account in your pool has hit its usage limit, Omniscio's normal behaviour is to stop: affected sessions park in a waiting state, a banner appears in the sessions sidebar, and everything resumes on its own when your subscription window resets. That is the right default — it never spends money you did not ask it to — but it means work stops, sometimes for hours.
The usage cascade lets you say, in advance and in your own order, what should happen instead. You build a short ladder. Each rung names one other AI vendor and how many sessions may use it at the same time:
| # | Fall back to | How many at once |
|---|---|---|
| 1 | GLM | No limit (default) |
| 2 | DeepSeek | At most 5 |
| — | wait for your subscription to reset (always last, not removable) | — |
When the subscription runs dry, each affected session walks that ladder from the top and takes the first rung that is switched on, set up (its credential is present), and not already full. If no rung qualifies, the session waits exactly as it does today.
"Runs dry" means every one of your live Claude accounts is genuinely at its limit. If one account is merely busy and another still has room, Omniscio just moves the session to the account with room, exactly as it always has — the cascade does not fire, and nothing is billed.
It also catches a vendor running out
The same ladder is used when a session that is already running on one of these vendors runs out on it — GLM hits its usage limit, say, or your DeepSeek balance hits zero. Instead of waiting for that vendor to reset, the session walks your ladder from the top, skipping the vendor it is leaving (and any vendor it already ran out on), and carries on at the first rung that qualifies.
For example, with just DeepSeek on your ladder, a GLM session that hits GLM's usage limit keeps going on DeepSeek (Flash, DeepSeek's default model) with its conversation intact, and re-runs the step that was interrupted.
A few details:
- The vendor's own backups come first. A second GLM key and every other row of the model family's supply list (Settings → Accounts → Who pays & who serves — your own account, a reseller, Omniscio credits) are used before a session leaves its vendor.
- No bouncing. A session never walks back onto a vendor it already ran out on in the same stretch, so it moves at most once per rung and then waits. After its next successful turn it forgets that list, so a later run-out can use a vendor that has since recovered.
- A session started on a vendor stays where it moved. Nothing can tell when a vendor's limit has reset, so it is not sent back automatically; new sessions still start on the vendor you chose.
- Existing ladders pick this up — and an empty ladder still changes nothing.
It ships empty. Until you add a rung, nothing changes at all.
Where to find it
Settings → Accounts → Usage cascade. The card shows your ladder in order:
- Add a fallback… — a dropdown of the vendors you can use. Only vendors that run through Omniscio's own Claude process are offered (DeepSeek, Kimi, GLM, MiniMax and the like), because only those can hand the conversation back to Claude afterwards. A reseller such as DeepInfra is not offered as a rung: it is a row inside a model family's supply list, and an older rung that named one now runs on that model's family, whose list decides who serves it. Engines with their own separate session manager — Codex, Gemini, Devin, Cursor and friends — are deliberately not offerable: a session placed on one could never come home.
- No limit / Limit how many at once — a small switch per rung, off by default. Leave it off and the fallback runs as many sessions as it needs. Switch it on and a number box appears, capping how many sessions may sit on that one fallback simultaneously.
- An on/off switch per rung — turn a fallback off without losing its place in your order.
- Up / down arrows — reorder. Order is the whole point: the first rung that works wins.
- A trash icon — remove the rung entirely.
- A "Bills per token" label on any vendor that charges per use, so a paid rung is never a surprise.
The final rung — wait for your subscription to reset — is shown as a non-editable row so the full behaviour is visible rather than implied.
How it behaves
What happens to a session that falls back
- The session starts on the fallback vendor and keeps its conversation — the transcript and context carry over, because these vendors run through the same underlying CLI.
- It says so once in its own transcript, naming the vendor it moved to and why.
- The next time that session starts, if your Claude subscription has room again, it moves back to Claude automatically, restored to the exact model it was using before.
You do not have to do anything to bring sessions home, and nothing is left stranded on a paid vendor after your subscription recovers. (A session you started on a vendor — GLM, say — is the one exception: it stays on the rung it moved to, as described above.)
Why there is no "cheaper Claude model" rung
An earlier design offered "drop to a cheaper Claude model" as the first rung. It was removed deliberately, because it cannot work at the moment it would be consulted.
Omniscio decides the pool is exhausted by looking at your overall 5-hour and 7-day usage windows — not at per-model limits. So by the time the cascade runs, a cheaper Claude model would be asking the same account that is already at its cap, and would simply fail again. Worse, Claude always counts as "set up", so the rung could never be skipped: it would be taken every time, announce itself, and deliver nothing.
If you want Omniscio to use a cheaper Claude model when usage gets tight, that is a separate, already-shipped feature that fires earlier — when you approach a threshold, rather than after you hit the wall. See the rate-limit model downgrade settings.
The promises this feature makes
- Nothing happens until you ask for it. An empty ladder behaves exactly like today.
- It only fires when a lane is genuinely out — every Claude account at its limit, or the vendor a session runs on refusing it for its usage limit or an empty balance. One Claude account running out still just switches accounts, as it always has.
- A vendor's own backups come first, and a session never bounces between two vendors that are both out.
- A cap is opt-in and means "at most this many at once" — never "run fewer sessions". A rung starts with no limit; you switch a cap on if you want one. It never limits your Claude sessions.
- Money is never spent by surprise. A rung that bills per token is labelled where you configure it, and can only ever be reached because you put it there.
- Every session that left Claude comes home once your subscription has capacity again.
- A full or broken rung is skipped silently, never an error. Nothing fails because a fallback was unavailable.
- You can always see it happened — a move is announced in that session's transcript.
Turning it off
Two ways:
- Clear the ladder (or switch every rung off) in Settings. The feature becomes inert.
- Set
AMC_DISABLE_USAGE_CASCADE=1in the environment. This restores today's behaviour exactly while leaving your ladder saved, so you can switch it back on without reconfiguring.
There is also a second, inherited switch worth knowing about: the cascade lives inside the
capacity-park route, so AMC_DISABLE_CAPACITY_PARK=1 turns it off too, one layer up.
For agents
How it works under the hood
Everything happens at one decision point: the moment a session tries to start and cannot get a Claude credential because every account is capped.
- Going out —
decideEmptyCredsRoute(src/shared/account-pool-exhaustion.ts) gains acascaderoute. The spawn guard (src/main/process/spawn-guard-router.ts) stamps the session onto the chosen vendor and relaunches it there. If nothing resolves, it parks exactly as before. - Coming home — just before a session respawns (session-relaunch-service.ts and session-auto-resume.ts), if it is on a fallback and Claude has capacity, its original vendor and model are restored first.
There is deliberately no background sweep in either direction. Each session decides for itself, at the only moment it can act, which is why a crash or restart can never leave one stranded.
The vendor trigger has its own decision point: the end of the turn a vendor refused. The two
turn-end vendor gates
(rate-limit-turn-end.ts,
vendor-insufficient-balance.ts)
ask usageCascadeWalkFor
(recovery-helpers.ts) after the
vendor's own fallbacks. With no candidate rung they park exactly as before, in the same tick;
otherwise the walk resolves a rung (excluding the lanes this session ran out on), re-checks that no
new turn started meanwhile, moves the session in one statement (stepSessionOntoCascadeLane), and
re-runs the interrupted turn there.
One rung's re-run never blocks the next rung's. Every one of these re-runs goes through the same
door (replayLastOperatorMessage), which holds a short cross-path window so that two recoveries
reacting to the same failure cannot both re-run it. That window is scoped to the failure, not to
the clock: once the turn a re-run started has come back with a result — the usual shape when a
new lane refuses within seconds — the next fallback is free to re-run it. Before 2026-09-25 the
window was time-only, so a ladder that moved a session onto a lane which then refused immediately
stopped there and parked in Needs You with its next source sitting ready. It now carries on down
the ladder. The shared per-turn replay budget still caps the whole chain, so a lane that keeps
failing cannot spend without limit.
- The ordering rules are a pure function,
planUsageCascade(src/shared/providers/usage-cascade.ts) — it does no I/O, so it is fully testable. - The rung it picks is resolved by usage-cascade-resolver.ts, which checks each vendor's readiness and how full each lane is.
- Which lane a session is on is recorded on the session row itself, together with the way home. That record is also what counts toward each rung's cap, so a session that is waiting still holds its slot.
Full invariants and rationale: usage-cascade-contract.md.
Related
- account-pool.md — how Omniscio picks between multiple Claude logins, and the exhaustion banner this feature takes over from.
- distribute-sessions-across-accounts.md — spreading sessions across the Claude accounts you already have, which happens before any of this.
Last verified 2026-09-28