---
title: MCP Overhead alert (when your MCP servers cost too much)
---

# MCP Overhead Alert (MCP Bloat checker)

## What it is

Amber inbox card that fires when the **MCP servers configured across your sessions** add up to too much per-request overhead. It is the MCP analogue of the [Doc Token Alert](doc-token-alert.md): same inbox-card + snooze + Settings-limit shape, but **global** (MCP config is not per-project in Omniscio, so there is exactly one rule) and with **three independent trip legs** instead of one.

### What it watches

Every MCP server you have enabled injects its **tool definitions into every request's context** (tokens) and, while a session is running, uses **memory**. The checker watches the global set of servers — your `~/.claude.json` servers plus Omniscio's custom-server registry — and fires when any of these crosses its limit:

1. **Estimated tokens** — a heuristic estimate of the tool-definition token overhead. **Always labeled an estimate**: Omniscio never starts a server to count its real tool-definition tokens (that would stall sessions for a minute+, the same reason spawns never pass `--mcp-config`).
2. **Active-session RAM** — measured memory (MB) of live MCP child processes, summed from the Resources sampler. **Best-effort**: when no session is active (or no fresh sample exists) `ramMb` is `null` and this leg is simply skipped — it never blocks the other two.
3. **Unused servers** — count of configured servers with **zero tool calls** in the usage window (from the `tool_invocations` scanner). The most reliable "wasted weight" signal.

A limit of `0` disables that leg. The alert fires when **any** enabled leg trips (token OR RAM OR unused).

## Where to find it

The alert arrives as an **amber card in your inbox**, titled "MCP overhead over budget", alongside your other alerts — open it for the numbers and your options. (It is not grouped under the **MCP Servers** sidebar entry; that panel is for managing your servers, not for this alert.) Its limits live in **Settings → Notifications → MCP Overhead**.

## How it behaves

### Honesty contract

- **Estimate:** the token figure (clearly marked `est.`).
- **Measured:** server count, active-session RAM, and per-server usage (call counts, "never used").

A non-technical user never sees a guess dressed up as a fact.

### The card

The card is titled "MCP overhead over budget". Opening it shows the headline estimate (red when over the token limit), active-session RAM (or "Not measured — no active session"), the unused/total server count, the configured limits, and a **per-server breakdown** (each server's estimated tokens, source, call count, and a "never used" flag). Three actions:

- **Archive** — the card's own Archive button (top of the card; also the keyboard `E` shortcut, middle-click, bulk archive, or an inbox rule) acknowledges the overhead at its current size (sticky-to-size, tiered re-alert: won't re-fire until a tripping metric grows past the dismissed value by the re-alert growth %, or the re-alert cadence elapses).
- **Manage MCP servers** — opens the MCP Servers panel so you can turn off the dead weight.
- **Start session** — the universal button every alert card carries, for asking an agent to help instead.

The card also supports the universal inbox **snooze**, and you can mute MCP Overhead alerts entirely — from the card, or from Settings → Notifications → Alert types.

### Settings

**Settings → Notifications → MCP Overhead** holds the single global rule: a master enable toggle plus tunable limits (estimated-token limit, active-RAM MB limit, unused-server limit) and the two re-alert knobs (growth %, cadence days). The master switch is the `enabled` flag on the seeded `'global'` rule — there is no `AppSettings` field.

## For agents

### Implementation notes

- **Decision logic** is one pure module: [`src/shared/alert-features/mcp-bloat-alert.ts`](../../src/shared/alert-features/mcp-bloat-alert.ts) — `decideMcpBloatCard(scan, rule, now)` (the firing / not-firing / unknown decision the inbox card producer reads) plus `shouldFireMcpBloatAlert` and the per-leg `tokenLegTrips` / `ramLegTrips` / `unusedLegTrips` helpers. The clock is injected, so it is testable without faking time.
- **Scan engine** ([`src/main/services/mcp/mcp-bloat-scan.ts`](../../src/main/services/mcp/mcp-bloat-scan.ts)) resolves the global server set, estimates tokens, joins the 30-day usage aggregate for unused detection, and reads the active-session MCP RAM from the resource sampler. It reuses existing data — no server is ever started.
- **Storage** is a single global rule + a single global scan row (no per-project rows), forward-only migration on the frozen baseline.
- **The inbox card** is a central alert, not its own inbox source: [`src/main/services/mcp/mcp-bloat-alert-card.ts`](../../src/main/services/mcp/mcp-bloat-alert-card.ts) turns the scan/rule into one global card (dedup key `mcp-bloat-alert:global`) via the shared `reconcileAlertFamily` / `registerAlertArchiveListener` building blocks every alert family uses. Its custom body is [`src/renderer/src/features/mcp-bloat-alert/McpBloatAlertCardBody.tsx`](../../src/renderer/src/features/mcp-bloat-alert/McpBloatAlertCardBody.tsx), mounted inside the shared `AlertInboxViewer`. Triggers and IPC handlers otherwise mirror the Doc Token Alert feature.

See the contract for the full invariant list: `.claude/memory/contracts/mcp-bloat-alert-contract.md`.

## Related

- [mcp-servers.md](mcp-servers.md) — the servers this alert is measuring, and how to add or trim them.
- [doc-token-alert.md](doc-token-alert.md) — the same kind of alert for documentation overhead rather than MCP tool definitions.

