---
title: Usage Forecast (predicted rate-limit exhaustion)
---

# Usage Forecast (predicted rate-limit exhaustion)

## What it is

Every Claude.ai subscription enforces a **5-hour rolling rate-limit window** — once you hit 100% of your usage budget inside that window, new turns refuse until the window resets. Omniscio already shows the raw percentage as a usage bar in **Settings → Accounts** for each OAuth account, but a percentage is hard to plan against. The Usage Forecast feature translates the current burn rate into plain-English ETAs and surfaces them in two places:

- A **global rollup banner** at the very top of the Accounts popover, answering the headline question "when am I going to run out across all my accounts?"
- A **per-account projection row** under each account's 5-hour bar — for example "~2.3h until exhausted, resets in 4h" or "Almost out — ≈18 min until exhausted, resets in 1h" or "won't reach limit before reset".

Read either at a glance and you know whether you can keep working or whether you should pace yourself.

The forecast also drives **opt-in early warnings**. If you set the "5-hour exhaustion warning" threshold above zero, Omniscio fires a one-shot warning event the first time the projected exhaustion drops below your threshold inside any given 5-hour window — so you can switch accounts, slow down, or wrap up before you actually hit the wall. The warning is deduplicated per window (you'll never get two alerts for the same 5-hour bucket) and re-arms automatically when that window resets.

Both buckets are forecasted, and as of 2026-06-01 **the headline verdict and the coverage line come from a single shared walk per bucket** — so the top line ("~Nh until you run out" / "Won't run out before reset") can never disagree with the "5h/7d coverage" line beneath it. Each bucket walks a **fixed planning horizon** (5 hours for the 5-hour bucket, 7 days for the weekly bucket) hour-by-hour, drawing down the pool's usable headroom at the projected burn rate. The weekly bucket **refills each account's 100pp window at the instant its reset actually fires** — time-correct relief — rather than pre-crediting all of it up front. So the headline answers from the headroom you have RIGHT NOW against near-term burn: a pool that is near-full on the 7-day limit shows a short honest runway, not a false multi-day one.

The total weekly capacity _supplied over the 7-day horizon_ still equals "current headroom + 100pp per in-horizon reset" (so the weekly **coverage number** is unchanged), but the headline's exhaust ETA reflects when the pool actually runs dry given that those refills arrive gradually. This fixed two related bugs: an earlier one where the weekly buffer used the soonest-reset horizon (~16-24h) and over-stated coverage (a 14-account pool looked like 380-530% when reality was ~150-220%), and a later one where the headline pre-credited the future-reset capacity as an up-front lump and projected multi-day runway (~136h) for a pool whose accounts were all sitting at 89-100% weekly.

## Where to find it

Click the utilization percentage chip in the top-right of the main window (for example `91% v`). The floating **Accounts** popover opens with the global rollup banner at the very top, and below it each OAuth account row shows its `5h` / `7d` / `Extra` mini-bars with the forecast row directly underneath them.

The same per-account projection row also appears in **Settings → Accounts**, under each OAuth account card's coloured 5-hour bar. The global banner is popover-only — Settings → Accounts does not show it — and API-key accounts show no forecast row in either place, because an API key has no 5-hour rate-limit window to project against.

Its controls live in **Settings → Notifications**, at the bottom of the page: the **Usage forecast** master toggle, and nested under it the **5-hour exhaustion warning** threshold.

For the full picture there is also a **Usage Stats** sub-page under the **Statistics** virtual project in the Omniscio sidebar group (open it and click the **Usage** tab), and an optional **Usage widget** chip you can add to the top header row from the header's **⋯ overflow** menu or **Settings → Widgets**.

## How it behaves

### Global prediction (top of popover)

When you click the utilization chip in the top-right of the main window (e.g. `91% v`), the floating Accounts popover opens. The first thing you see — above the per-account list — is a single one-line banner that aggregates every OAuth account into one verdict. It uses the same color tiers as the per-account row but operates on the rollup of every login account.

**Visual states the global banner can be in:**

- **Hidden** (renders nothing) — when there's no signal worth showing: zero login accounts, any login account is missing usage data, all login accounts at 0% utilization (no usage yet), or the master `Usage forecast` toggle is off.
- **`Calculating…`** (muted grey) — at least one account has fewer than ~10 minutes of signal in the current 5-hour window. Slope estimates from short samples are noise, so the global verdict holds off until every account has enough data.
- **`All accounts exhausted — resets in <Xh>`** (red) — every login account is at 100% utilization. The number shown is the time until the SOONEST account refills — that's when you can start working again on at least one account.
- **`Won't run out before reset`** (muted green) — at the current burn rate, cascading resets save you. Some accounts may individually run out, but the slowest account doesn't exhaust until AFTER an earlier-exhausted account has already reset, so there's never a moment in the current windows when every account is simultaneously at 100%. (See "How it works" for the algorithm.) As of 2026-06-07 the `· data N min stale` suffix can append to this green verdict too (not just the countdown states below), so even a "won't run out" line discloses when it's drawn from stale data.
- **`<duration> until your 5-hour limit`** / **`<duration> until your weekly limit`** (3-tier color) — the standard projection, labelled with WHICH limit binds (from `bindingBucket`) so a transient 5-hour dip never reads like the multi-day weekly wall. "Duration" uses the same human-friendly projection formatter as the per-account row (`~2.3h`, `≈45 min`, and now `~1.8d` for multi-day weekly spans — ≥24h collapses to one-decimal days). When the binding bucket is unknown (older payloads), it falls back to the generic `<duration> until you run out`. Color and prefix shift on urgency:
  - **Less than 30 minutes** — red, prefixed `Almost out — `.
  - **30 minutes to 2 hours** — amber, prefixed `Running low — `.
  - **2 hours or more** — green, no prefix (calm).

  The `Almost out — ` and `Running low — ` prefixes are intentional so urgency isn't communicated by color alone — that matters for color-blind users and screen-readers (WCAG 1.4.1).

  Optional suffixes (only when the banner is in this projected state):
  - **`· data N min stale`** — only appears when the most recent successful usage poll for at least one account is more than 30 minutes old. (Zero minutes is fresh, not stale, so the suffix is suppressed.)
  - **`· unusual rate`** — at least one account is burning faster than 80%/hour. Could be a runaway agent, another machine logged into the same account, or a real intense session. Worth a glance.

**Plain-English algorithm.** Omniscio asks each OAuth account when it expects to hit 100% at the current burn rate. It then finds the LATEST exhaustion time among them — the moment when even the slowest account hits 100%. That time `T` is your candidate global ETA. Then Omniscio checks whether the earlier-exhausted accounts have already reset by `T`. If yes — cascading resets save you, and the banner shows the muted-green "Won't run out before reset". If no — there is a moment in the current 5-hour windows when every account is simultaneously at 100%, and `T` is reported as `<duration> until your <binding> limit` (the binding bucket — 5-hour or weekly — names the limit).

**Known limitations.**

- **OAuth-only.** API-key accounts have no 5-hour rate-limit window (they bill per-token), so they're filtered out of the global rollup entirely. If you have ONLY API-key accounts, the banner stays hidden.
- **All login accounts must have usage data.** If even one OAuth account is missing its usage payload, the banner stays hidden rather than aggregating partial capacity. (We don't extrapolate from "some accounts" — you might be expecting that missing account's capacity to be available.)
- **Reset-aware within a fixed horizon.** The headline walk steps forward to a fixed horizon (5h / 7d) and DOES account for resets that land inside that window — each weekly reset refills its account's headroom at the moment it fires. What it does NOT model is windows that open _after_ the horizon ends. So "Won't run out before reset" means "across the next 5h / 7d, accounting for the resets that land in that span, the pool never hits zero". For most users this matches their mental model, but it's worth disclosing.
- **Headline walks both buckets, and matches the coverage line.** The headline picks the earlier-exhausting bucket and reports it as the binding constraint. Because it shares a walk with the coverage estimate, `coverage < 100%` always lines up with "will run out" and `coverage ≥ 100%` with "won't run out". A `bindingBucket: 'fiveHour' | 'weekly' | null` field on the `GlobalBucketForecast` payload identifies which one is binding, and the headline uses it to label the limit ("until your 5-hour limit" vs "until your weekly limit") — so a 5-hour dip that refills in hours never reads like the weekly wall that strands you for days. `null` is returned when neither bucket runs dry inside its planning horizon (the headline then shows the muted-green "Won't run out before reset").

### Hour-of-day forecast model

Below the global headline (the "<duration> until you run out" / "Won't run out before reset" / "All accounts exhausted" line described above), the same banner now renders a second muted line that answers a different question: **at the rate you _typically_ burn at this hour and weekday, how much headroom do you have?** Instead of extrapolating the last 10–60 minutes of activity (which the headline above does), this line predicts forward using your last 14 days of usage, weighted by the cell of the hour-of-day × day-of-week grid you're currently in.

**Visual states the coverage line can be in:**

- **Hidden** (renders nothing) — when the model has no usable numbers for this bucket (`point`, `low`, and `high` are all non-finite). Defensive only; the production model shouldn't produce this.
- **`Learning your patterns — N more days needed`** (muted grey) — true cold start. Omniscio has **zero** distinct calendar days of usage samples in the rolling 14-day window — i.e. the very first session, before the sampler has logged anything. `N` is `7 − distinctDays`, so you'll see `7 more days needed` here. This is the only visible state where the coverage line replaces both the 5h and the weekly numbers — neither is shown until at least one sample lands.
- **`5h <P>% · 7d <P>% coverage · still learning — N of 7 days`** (muted grey, limited data) — between 1 and 6 distinct calendar days of samples. The model still renders predictions (the fallback chain — cell → weekday-mean → global-mean — fills the empty cells), but a `· still learning — N of 7 days` suffix discloses that you haven't hit the full-confidence threshold yet. The suffix only appears when **both buckets render as point estimates** — when one or both render as a range, the range itself already conveys uncertainty visually and adding the text would be redundant noise.
- **`5h <P>% · 7d <P>% coverage`** (muted grey, full confidence) — 7+ distinct days of samples. Both buckets have enough samples in their HoD × weekday cell that the within-cell standard deviation is small enough to commit to a point number. `<P>` is `capacity ÷ projected_burn × 100`, rounded to the nearest integer — where `capacity = 100 − current_utilization` (room left in the bucket) and `projected_burn = projected_total − current_utilization` (consumption during the remaining window). `>100%` means you're projected to finish the bucket with headroom; `<100%` means projected to overshoot. `100%` is the break-even line. No disclosure suffix.
- **`5h <Lo>–<Hi>% · 7d <P>% coverage`** (muted grey, low confidence on one bucket) — when the within-cell std exceeds 25 percentage points for a bucket, Omniscio widens that bucket's number to a range (`mean ± std` projected through the bucket horizon). Either bucket can be a range independently — the 5h line might show a range while the weekly line shows a point, or vice versa. The en-dash `–` (U+2013) is intentional. No `still learning` suffix here even at 1-6 distinct days, because the range itself signals uncertainty.
- **One bucket only** — if one bucket has no usable numbers but the other is fine, only the fine one renders (e.g. just `5h 142% coverage`).

The coverage estimate is intentionally muted (`text-surface-500`) and — since the two-line redesign — shares ONE detail line with the sessions-remaining estimate directly under the headline, joined by `·` (e.g. `5h 74% · 7d 157% coverage · ≈59 sessions left`). The headline's existing `py-2` provides the gap. This is visual subordination on purpose: the headline answers "am I in trouble right now?", the detail line answers "given my usual patterns, how much room do I have — and how much work does that buy?". The whole banner is at most two lines, headline first.

**Plain-English algorithm.** Omniscio samples each OAuth account's utilization every 15 minutes and stores `(timestamp, accountId, bucket, utilization%)` rows for 14 days, then drops anything older. To predict, it does five things:

1. **Convert samples into hourly burn deltas.** Pair adjacent samples (within 90 minutes of each other; wider gaps are treated as silence and skipped); the delta is `utilization_now − utilization_prev` in percentage points. Anything negative (a window reset wrapped around) is clamped to 0 — those don't represent burn, they represent the bucket refilling.
2. **Bucket those deltas into a 168-cell grid.** Each cell is one combination of `hour-of-day (0–23) × day-of-week (Sun–Sat) × bucket (5h | weekly)`. Per cell, Omniscio computes the mean and population standard deviation of the deltas. Cells with no observations stay empty and fall back at lookup time (see step 4).
3. **Walk forward from "now" to bucket reset.** Omniscio looks up the current cell, projects that mean × the remaining minutes inside this hour, then advances to the next hour, looks up THAT cell, projects forward, and repeats until it hits the bucket's reset time (5 hours away for the 5h bucket; up to 7 days for the weekly bucket). The projected cumulative burn is the predicted total usage between now and reset.
4. **Fall back when a cell is sparse.** If a cell has no observations, Omniscio falls back to the same hour-of-day across all weekdays. If THAT mean is also empty, it falls back to the bucket's overall mean. (Sample size of 1 in a cell widens the std defensively.)
5. **Compute the coverage.** `capacity ÷ projected_burn × 100`, where `capacity = 100 − utilization_now` (remaining percentage points before the bucket hits 100%) and `projected_burn = projected_total − utilization_now` (the cumulative burn from step 3, expressed as the consumption during the remaining window). `100%` is the break-even line (capacity equals projected burn, zero headroom); `>100%` means projected to finish with headroom; `<100%` means projected to overshoot. Edge cases: when `capacity ≤ 0` (already at the wall) coverage is `0`; when `projected_burn ≤ 0` (resets-at already in the past, or no burn projected) coverage is `Infinity` and the renderer suppresses the line. The same walk is run twice more — once with each cell's `mean − std` (low burn variant) and once with `mean + std` (high burn variant) — to produce the low/high range. The low-burn walk produces the HIGH coverage, and the high-burn walk produces the LOW coverage (more burn → less coverage). When the per-cell std is low (≤25 percentage points across the cells touched in the walk), the range collapses tightly enough that Omniscio commits to a single point estimate; otherwise it renders the range.

**Cold start vs limited data.** Two thresholds apply, and they're decoupled on purpose:

- **`distinctDays === 0` — true cold start.** No samples at all (very first session, before the 15-minute sampler has logged anything). Omniscio suppresses both predictions and shows the `Learning your patterns — 7 more days needed` message instead. This is the ONLY case where the coverage line replaces the predictions.
- **`1 ≤ distinctDays < 7` — limited data.** The model has at least one sample, so the fallback chain (cell → weekday-mean → global-mean) is enough to project numbers — they're imperfect but still the best signal Omniscio has. Omniscio renders the full coverage line and appends a `· still learning — N of 7 days` disclosure suffix, but only when both buckets render as point estimates. When one or both render as a range, the range itself signals uncertainty and the text would be redundant noise.
- **`distinctDays ≥ 7` — full confidence.** Predictions render with no disclosure suffix.

This was a deliberate UX shift in 2026-04: the original design suppressed all numbers for 7 days and showed only the cold-start hint, on the theory that "imperfect numbers are worse than no numbers". User feedback rejected that — they explicitly want continually-updating predictions with disclosure of low confidence, not nothing. Distinct calendar days are counted in the user's local timezone, so an overnight Tue→Wed session counts as two days. The threshold (7) and the window (14 days) are NOT user-tunable on purpose — per design, the coverage indicator has one knob and it's the existing `usageForecastEnabled` flag (off → no coverage line, no headline, no warnings).

**What this does NOT do.** It does not change the 3-tier color or the "Almost out / Running low" prefix logic on the headline above — those still come from the flat-extrapolation `forecastGlobalExhaustion` math. It does not fire warnings; only the headline math drives the warning dispatcher. It is computed in the main process per cache write (same cadence as the headline) and pushed to the renderer in the same payload — no extra IPC call.

**Where the code lives.**

- The model itself is a pure function in [/src/shared/usage-forecast-hod.ts](/src/shared/usage-forecast-hod.ts) — `samplesToHourlyDeltas` (step 1), `buildHodCells` (step 2), `forecastWithHodPool` (steps 3–5). All three are pure-functional with no Electron / SQLite / `Date.now()` dependencies, so they're fully unit-testable. Constants are inline: `COLD_START_DAYS = 7` (full-confidence threshold; cold-start gate is now `distinctDays === 0` regardless), `RANGE_THRESHOLD = 25` (percentage-point std cutoff for collapsing to a point), `MAX_GAP_MS = 90 × 60 × 1000` (max gap between paired samples). The `HodForecast` return type carries both `coldStartDaysRemaining` (kept for back-compat with the existing payload field) and `distinctDays` (added for the new "still learning" disclosure). `forecastWithHodPool` takes a `PoolAccountInput[]` (one entry per login account, with that account's own `currentFiveHourPct` / `currentWeeklyPct` / `fiveHourResetsAt` / `weeklyResetsAt`) and walks each account independently, summing per-account capacity and burn into a pool-wide coverage line — exhausted accounts (≥100%) are skipped, since they can't contribute work until reset.
- Sample storage is the `usage_samples` table introduced in schema migration v116. Query helpers — `insertUsageSample`, `getUsageSamplesSince`, `deleteUsageSamplesOlderThan` — live in [/src/main/db/queries-usage-samples.ts](/src/main/db/queries-usage-samples.ts).
- The periodic sampler service is [/src/main/services/usage/usage-sampler.ts](/src/main/services/usage/usage-sampler.ts). It writes a row every 15 minutes per OAuth account and runs a TTL cleanup pass to drop rows older than 14 days. It's wired into the app lifecycle as the `Usage sampler` startup task in [/src/main/startup/registry.ts](/src/main/startup/registry.ts) (run via `runStartupTasks`). At ingest time the model collapses multiple samples within the same `(accountId, local-hour)` bucket to a single observation (last sample wins) before computing deltas, so future cadence changes (faster OR slower) don't skew the per-cell mean.
- The IPC handler that combines live utilization + samples + the HOD model into the `GlobalBucketForecast.buffer` field is `computeGlobalForecastWithHod(now)` in [/src/main/ipc/usage-forecast-handlers.ts](/src/main/ipc/usage-forecast-handlers.ts). It reads `getSettings()`, `authService.getCachedUsage()`, and `getUsageSamplesSince(now − 14d)`, builds a `PoolAccountInput[]` with one entry per login account (each carrying that account's own utilization + reset times), runs the three-step pipeline via `forecastWithHodPool`, and returns the result on the existing `GlobalBucketForecast` payload alongside the headline state. The pool walk preserves headroom from less-utilized accounts — collapsing to a single MAX-utilization account before walking (the previous shape) silently dropped that headroom, so the coverage line could read `0% coverage` while the headline above said "won't run out before reset".
- The renderer-side coverage text is built by `formatCoverageCluster()` in [/src/renderer/src/features/settings/GlobalUsageForecastRow.tsx](/src/renderer/src/features/settings/GlobalUsageForecastRow.tsx) (near the bottom of the file), which returns: (1) `distinctDays === 0` → the cold-start hint (`Learning your patterns — N more days needed`); (2) `1 ≤ distinctDays < 7` AND both buckets are point estimates → the coverage numbers plus a `· still learning — N of 7 days` suffix; (3) otherwise → coverage numbers only. Range vs point is decided per-bucket by `formatBucket()` from each bucket's `confidence` field, and the single trailing `coverage` label is appended once by the cluster. The cluster's string is then merged with the sessions estimate inside `ForecastDetailLine` (see the next section), which returns `null` when neither segment has anything to show — so a missing buffer just shrinks the banner back to the single headline line.
- The cross-process types — `BufferEstimate`, `GlobalBufferLine` — are re-exported through [/src/shared/types.ts](/src/shared/types.ts).

### Sessions remaining (part of the detail line)

Whenever the pool is in a **projected** state, the muted detail line appends a sessions-remaining estimate after the coverage cluster — for example `5h 74% · 7d 157% coverage · ≈59 sessions left`. It answers the most concrete planning question of all: **"how much more work does that headroom actually buy me?"**

Where the countdown answers "how long" and the coverage cluster answers "do I have room given my usual patterns", this segment answers "how many more things can I start". As of 2026-06-07 it shows in **both** projected sub-states: when you're **projected to run out** it counts down the binding limit, AND when you're **healthy** ("Won't run out before reset") it shows your current runway too. In the healthy case there's no single binding limit, so Omniscio measures your **current** headroom against **whichever of your two limits (5-hour or weekly) you'd hit first** at your recent burn and shows that — your honest right-now runway (no borrowing against future resets). Because it counts only what you have _now_, in a healthy pool it can read modest (e.g. `≈6 sessions left`) even under the green headline — a reset simply lands before you'd burn through it.

**Why it's an estimate (and why it can't be exact).** Anthropic reports your limit as an opaque **utilization percentage**, not a token bucket — there is no published "tokens = 100%" number anywhere in the API. So Omniscio **calibrates empirically**: over the last 7 days it measures how many percentage-points of the binding limit the pool burned against how many real sessions (cost > $0) Omniscio logged in that same window, giving a "percent-per-session" rate it multiplies by your **current** headroom in the binding bucket. The result is deliberately fuzzy — rendered with `≈` — and only as accurate as the usage Omniscio actually sees. If you also burn these same Claude accounts outside Omniscio, the calibration skews conservative (it under-counts how much you have left).

> **Tokens are computed but not shown.** Omniscio still calibrates a `tokensRemaining` figure on the same data and carries it on the wire (`GlobalBucketForecast.tokensRemaining`, for other consumers), but the banner **no longer displays the raw token count** — it duplicated the sessions estimate and added a fourth number to a line we were deliberately condensing to two. The user-facing runway is sessions only.

**When the segment stays hidden.** It shows nothing — rather than a misleading number — whenever it can't calibrate honestly: fewer than 3 real (cost-bearing) sessions in the last 7 days, or essentially zero recent burn (an idle pool). With too little signal the rate is unstable, so Omniscio declines rather than extrapolate a wild number (the same class of guard that keeps the coverage line from exploding near a reset boundary). In the healthy state, where Omniscio scores both buckets, it hides only when _neither_ bucket has enough signal.

**Rounding.** Sessions round to a whole number: `≈59 sessions left`, singular `≈1 session left`, and `≈ less than 1 session left` when the estimate is positive but rounds below one.

**Scope.** Like the rest of the banner, it's **pool-wide** (all your login accounts together) and keyed to **whichever limit binds first** (5-hour or weekly) — the binding limit when you're running out, or the tighter of the two when you're healthy. It never contradicts the headline above it.

**Where the code lives.**

- The pure bridge is `estimateRemainingWork()` in [/src/shared/usage-forecast-hod.ts](/src/shared/usage-forecast-hod.ts) — a walk-free ratio (it never touches the forecast walk, so it can't regress the coverage/headline invariants) that returns `null` unless there's enough signal (`MIN_PP_FOR_ESTIMATE`, `MIN_SESSIONS_FOR_ESTIMATE = 3`, positive finite headroom). It still computes `tokensRemaining` alongside `sessionsRemaining`.
- The handler glue is `computeRemainingWork()` in [/src/main/ipc/usage-forecast-handlers.ts](/src/main/ipc/usage-forecast-handlers.ts): per bucket it sums the pool's current `usableCapacity`, sums that bucket's percentage-point burn over the last 7 days (reusing the same `samplesToHourlyDeltas`→`aggregateToPoolDeltas` deltas the model already derives), and reads the recent real-session workload via `getRecentSessionWorkload()` in [/src/main/db/queries-sessions/cost.ts](/src/main/db/queries-sessions/cost.ts) — queried once. When a bucket binds it scores exactly that bucket; when nothing binds (the healthy state) it scores both buckets and surfaces the tighter via the `tighterEstimate` helper (the whole estimate with fewer `sessionsRemaining`, never field-mixing). The result rides on the existing `GlobalBucketForecast` payload as optional `sessionsRemaining` / `tokensRemaining` fields.
- The renderer is `formatSessionsLeft()` inside the merged `ForecastDetailLine` in [/src/renderer/src/features/settings/GlobalUsageForecastRow.tsx](/src/renderer/src/features/settings/GlobalUsageForecastRow.tsx); the `tokensRemaining` field is intentionally not rendered, so `formatTokensShort()` is no longer used by this component.

## Related

This page continues in [Usage Forecast (part 2)](usage-forecast-part-2.md), which carries how to read each per-account forecast state, the early-warning threshold, the data-freshness line, the four-card Usage Stats page, the optional header widget and the architecture behind it all.

- [account-pool.md](account-pool.md) — the multi-account pool that this forecast helps you manage; an exhausted projection here often precedes the pool falling back to a different account
- [notifications-and-silence.md](notifications-and-silence.md) — the master notification controls that gate other OS toasts (the forecast warning isn't gated by global silence mute, but the controls live in the same Notifications settings page)
- [session-stuck-in-needs-you.md](session-stuck-in-needs-you.md) — when you DO hit a rate-limit, this is the recovery path that explains how Omniscio resumes sessions on a different account
