Omniscio documentation
Browse all documentation
  1. Getting Started13
  2. Sessions & Agents115
  3. Inbox & Notifications59
  4. Projects & Tasks95
  5. Automation & Scheduling75
  6. Knowledge & Memory26
  7. AI Features60
  8. Integrations100
  9. Plugins & Marketplace33
  10. Cloud & Teams56
  11. Settings & Customization58
  12. Account & Billing28
  13. Troubleshooting84
  14. CLI & API Reference22
  15. Legal & Policies4
  16. Uncategorised22

MCP Overhead alert (when your MCP servers cost too much)

An amber inbox card that fires when the MCP servers configured for your sessions add up to too much per-request overhead — too many tool definitions in context, too much memory while sessions run, or servers you never actually use. Covers the three independent trip conditions and the one global rule behind them.

What it is

Amber inbox card that fires when the MCP servers configured across your sessions add up to too much per-request overhead. It is the MCP analogue of the Doc Token Alert: same inbox-card + snooze + Settings-limit shape, but global (MCP config is not per-project in Omniscio, so there is exactly one rule) and with three independent trip legs instead of one.

What it watches

Every MCP server you have enabled injects its tool definitions into every request's context (tokens) and, while a session is running, uses memory. The checker watches the global set of servers — your ~/.claude.json servers plus Omniscio's custom-server registry — and fires when any of these crosses its limit:

  1. Estimated tokens — a heuristic estimate of the tool-definition token overhead. Always labeled an estimate: Omniscio never starts a server to count its real tool-definition tokens (that would stall sessions for a minute+, the same reason spawns never pass --mcp-config).
  2. Active-session RAM — measured memory (MB) of live MCP child processes, summed from the Resources sampler. Best-effort: when no session is active (or no fresh sample exists) ramMb is null and this leg is simply skipped — it never blocks the other two.
  3. Unused servers — count of configured servers with zero tool calls in the usage window (from the tool_invocations scanner). The most reliable "wasted weight" signal.

A limit of 0 disables that leg. The alert fires when any enabled leg trips (token OR RAM OR unused).

Where to find it

The alert arrives as an amber card in your inbox, titled "MCP overhead over budget", alongside your other alerts — open it for the numbers and your options. (It is not grouped under the MCP Servers sidebar entry; that panel is for managing your servers, not for this alert.) Its limits live in Settings → Notifications → MCP Overhead.

How it behaves

Honesty contract

  • Estimate: the token figure (clearly marked est.).
  • Measured: server count, active-session RAM, and per-server usage (call counts, "never used").

A non-technical user never sees a guess dressed up as a fact.

The card

The card is titled "MCP overhead over budget". Opening it shows the headline estimate (red when over the token limit), active-session RAM (or "Not measured — no active session"), the unused/total server count, the configured limits, and a per-server breakdown (each server's estimated tokens, source, call count, and a "never used" flag). Three actions:

  • Archive — the card's own Archive button (top of the card; also the keyboard E shortcut, middle-click, bulk archive, or an inbox rule) acknowledges the overhead at its current size (sticky-to-size, tiered re-alert: won't re-fire until a tripping metric grows past the dismissed value by the re-alert growth %, or the re-alert cadence elapses).
  • Manage MCP servers — opens the MCP Servers panel so you can turn off the dead weight.
  • Start session — the universal button every alert card carries, for asking an agent to help instead.

The card also supports the universal inbox snooze, and you can mute MCP Overhead alerts entirely — from the card, or from Settings → Notifications → Alert types.

Settings

Settings → Notifications → MCP Overhead holds the single global rule: a master enable toggle plus tunable limits (estimated-token limit, active-RAM MB limit, unused-server limit) and the two re-alert knobs (growth %, cadence days). The master switch is the enabled flag on the seeded 'global' rule — there is no AppSettings field.

For agents

Implementation notes

  • Decision logic is one pure module: src/shared/alert-features/mcp-bloat-alert.ts — decideMcpBloatCard(scan, rule, now) (the firing / not-firing / unknown decision the inbox card producer reads) plus shouldFireMcpBloatAlert and the per-leg tokenLegTrips / ramLegTrips / unusedLegTrips helpers. The clock is injected, so it is testable without faking time.
  • Scan engine (src/main/services/mcp/mcp-bloat-scan.ts) resolves the global server set, estimates tokens, joins the 30-day usage aggregate for unused detection, and reads the active-session MCP RAM from the resource sampler. It reuses existing data — no server is ever started.
  • Storage is a single global rule + a single global scan row (no per-project rows), forward-only migration on the frozen baseline.
  • The inbox card is a central alert, not its own inbox source: src/main/services/mcp/mcp-bloat-alert-card.ts turns the scan/rule into one global card (dedup key mcp-bloat-alert:global) via the shared reconcileAlertFamily / registerAlertArchiveListener building blocks every alert family uses. Its custom body is src/renderer/src/features/mcp-bloat-alert/McpBloatAlertCardBody.tsx, mounted inside the shared AlertInboxViewer. Triggers and IPC handlers otherwise mirror the Doc Token Alert feature.

See the contract for the full invariant list: .claude/memory/contracts/mcp-bloat-alert-contract.md.

Related

  • mcp-servers.md — the servers this alert is measuring, and how to add or trim them.
  • doc-token-alert.md — the same kind of alert for documentation overhead rather than MCP tool definitions.

Last verified 2026-09-28