---
title: Inbox Pilot (an AI that triages your inbox for you)
---

# Inbox Pilot

## What it is

**Inbox Pilot** is a cheap classifier LLM call (Qwen3-32B via OpenRouter by default, with a one-shot Haiku fallback) that runs at the end of every Claude session turn and decides what should happen to the session in your inbox. It looks at the latest agent end-of-turn and classifies it into one of six outcomes — keep it as **Needs You**, **hide** it, **archive** it, **snooze** it for an hour, **respond** to the agent on your behalf, or **schedule a response** to be sent later (up to 24h ahead) — and writes a one-sentence reason for every decision.

Inbox Pilot is **off by default** at three levels and **opt-in per session**:

1. **Top-level flag `inboxPilotEnabled` (hidden / config-only):** the master kill switch. Shipped 2026-05-23 as a **Settings → Features → Enable Inbox Pilot** toggle, but **as of 2026-06-30 that toggle and its Settings-search entry were removed — the flag is now a completely hidden, config-only setting** (default off; settable only by editing `config.json` or via the CLI control server's `PATCH /settings/inboxPilotEnabled`, never from the app UI). Its gating behaviour is unchanged: when this is off, the Inbox Pilot sidebar row is hidden, the per-session three-dot menu drops its **Inbox Pilot** / **Custom rule for this session** / **Allow Inbox Pilot to reply** / **Test on latest message…** rows entirely (2026-05-24 — previously they kept rendering and toggling them silently failed against the gated IPC), the `aiManagerService.evaluate()` end-of-turn pipeline returns immediately as a no-op (no Haiku calls, no audit rows, no overlays), and every Inbox Pilot IPC handler (rules, decisions, preview, per-session toggles, bulk apply, daily cost, pending overlays) rejects with `{success: false, error: "Inbox Pilot is disabled..."}`. As of **2026-07-20** — when Plain Speak split into its own Settings section — Inbox Pilot was retired from the UI entirely: the opt-in auto-enable was **removed** from `src/main/services/opt-in-toggle-migration.ts`, and a one-shot `src/main/services/migrations/inbox-pilot-retire-migration.ts` (own cookie `inboxPilotRetireMigrationDone`) **resets the flag OFF** for any install that still had it on, so `inboxPilotEnabled` is now uniformly false. The code + flag are kept dormant, re-enable-able only via `config.json` / the CLI.
2. **Master enable inside the sidebar tile** — at the top of the Inbox Pilot page.
3. **Per-session checkbox** in each session's three-dot menu.

It runs only on sessions where all three are on, and only while the daily cap and pause flag also allow it. Every decision Inbox Pilot makes is audited in a decisions log so you can see exactly what it did, why, what it cost, and which model it used. Inbox Pilot is no longer surfaced in Settings search — the feature toggle and its search-index entry were removed on 2026-06-30, and as of **2026-07-20** it is retired from the UI everywhere: the sidebar tile stays hidden, the per-session menus drop their Inbox-Pilot rows, and the Quick-Reply **"Allow Inbox Pilot to auto-send"** card is hidden too — all gated on `inboxPilotEnabled`, now uniformly false (the flag lives on config-only).

### What the classifier sees

Each Inbox Pilot call sends the classifier a **narrow** transcript — only the **latest agent message** and the **operator (user) message immediately preceding it**, formatted as `[operator] …\n[agent] …`. Earlier turns in the session are not included. The system prompt already directed the model to classify the latest agent turn and treat earlier turns as background context, so sending the rest was wasted tokens and an extra prompt-injection surface for older agent output the model wasn't supposed to weigh heavily.

Three follow-on details worth knowing:

- **`system` rows are excluded** as candidates for the operator slot. These are CLI/Omniscio bookkeeping (compact-divider, recovery banners), not operator input — including them would let the classifier reason about meta-events as if they were user instructions.
- **Recipe-injected operator rows ARE included.** When a recipe or automation injects a prompt and the agent replies, the classifier needs to see what the agent was responding to even though no human typed it. Filtering them would make the agent's reply read as a non-sequitur.
- **An operator typed AFTER the latest agent reply is not paired with that agent.** The pairing always picks the operator strictly BEFORE the agent's timestamp — an unanswered operator is a separate new turn awaiting its own agent reply.

If the latest agent message alone exceeds the classifier's 128k-char input budget, it's truncated from the end with an explicit `… [truncated, full N chars]` marker so the model knows the message was clipped. If `operator + agent` combined exceeds the budget, the operator is dropped and only the agent line is sent.

Plain Speak (the sibling overlay feature) uses a wider transcript — it's paraphrasing a longer span and has its own context builder. The narrowing described here applies to Inbox Pilot only.

## Where to find it

In your **inbox**, on the sessions it acts on — each one is marked so you can see the Pilot was there. The controls that govern it live in **Settings → Inbox Pilot**, and the per-session switch is on the session itself.

## How it behaves

### What you see

You see Inbox Pilot in three places:

- **The Inbox Pilot tile in the left sidebar.** Click the airplane (plane-takeoff) icon in the projects bar and a full-page Inbox Pilot view opens. At the very top is a single **Enable Inbox Pilot** master toggle — flip it off and the entire rest of the tile collapses away (only the toggle remains visible). Flipping it off prompts a confirm dialog and, on confirm, resets every sub-setting to its safe default while preserving your rule prompt text — see the **Master on/off toggle** bullet under "Cost and pause controls" below for the full reset behavior. With the master toggle on, the tile reveals: a **rollout-controls block** — a **Shadow mode** toggle, a **Default on for new sessions** toggle, a **Bulk apply** card with the **Enable on N active sessions** and **Disable on all sessions** buttons, the **Auto-respond** toggle, and the **Schedule reply for later** toggle. When today's spend reaches the daily cap, an amber **Daily cap reached** notice appears at the top of this block. Below the rollout controls is a **Cost & access** card with the **daily cost cap** input (range 0–100, displayed with a visual `$` prefix so it reads "$1" not "1"), today's spend bar shown inline beneath the cap input, and the **Allow on API-key accounts** toggle. The rule editor below has its own enable switch, a multi-line prompt textbox where you tune the rule in your own words, and a **Preview** button that runs the rule against a real session in dry-run mode without applying any side effects. Recent decisions are listed at the bottom (most recent first).
- **Each session's three-dot overflow menu.** Open the kebab menu on any session row and you'll see an **Inbox Pilot** entry as a top-level item with a checkbox you can flip on or off for that session. It defaults to off for every session — unless you've turned **Default on for new sessions** on in the Inbox Pilot tile, in which case new sessions arrive with the checkbox already on. The helper text under the checkbox tells you the per-turn cost (~$0.0001 per turn) when the toggle is on. This is a per-session opt-in; flipping a master switch in Settings doesn't enable a session by itself. When Inbox Pilot is **on** for a session, a **"Custom rule for this session"** disclosure appears directly below the checkbox. Expanding it reveals a textarea where you can add a session-specific instruction that layers on top of your global rule — for example, _"For this session, only flag me if a deploy fails. Otherwise hide."_ Edits are debounced 500ms and saved automatically; max 8000 chars; leaving it empty disables the per-session layer and falls back to the global rule alone. Below the custom-rule disclosure is a **"Test on latest message…"** menu item with a flask icon — clicking it opens a dry-run modal that runs Inbox Pilot against this specific session's transcript right now, layering in the saved global rule **and** this session's saved per-session rule (mirroring what the live router would do). The result shows the action Pilot would take and the one-sentence reason, so you can sanity-check that your rule does what you expect before letting it act for real. The test runs purely free — no action fires, the daily cap is not charged, the result is written to the decisions log with `was_dry_run = 1` (and prunes after 7 days like all preview rows). When Inbox Pilot is **off** for the session, the test row is greyed out with a tooltip telling you to turn the session toggle on first.
- **The "Pilot took a look" pill in the chat.** When Inbox Pilot fires on a session — whether in shadow mode or real mode — the agent turn it evaluated gets a small expandable pill rendered inline in the chat (one pill per turn, anchored after the agent's reply). Closed, the pill shows a status dot whose colour carries the outcome (green = applied, amber = cap hit, red = failed, grey = skipped or in-flight) and a label whose phrasing carries the action ("Pilot archived this thread" in real mode, "Pilot would have archived this thread" in shadow mode). Click to expand and the pill reveals the LLM's one-sentence reason, the model used, the call's cost, and the first 8 characters of the decision id. See "The 'Pilot took a look' pill" section below for the full anatomy.

### The six outcomes

Inbox Pilot classifies each agent end-of-turn into exactly one of these six labels (the model emits the upper-case label; the value persists in the decisions log as the lowercase enum value shown in parentheses):

- **`KEEP_NEEDS_YOU`** (`keep_needs_you`) — the agent has asked the user a question or hit a blocker that needs human input. The session stays in your **Needs You** queue. This is the default when Inbox Pilot is unsure.
- **`HIDE`** (`hide`) — the message is informational, mid-task narration, or "thinking out loud" that the user doesn't need to act on. The session's status flips to `ended` so it drops out of the attention list. Sending a new message to the session reopens it as normal.
- **`ARCHIVE`** (`archive`) — the agent has explicitly finished its task with no outstanding issues. The session is soft-deleted with the standard archive flow (you can undo from the toast). Auto-unarchives on the next send.
- **`SNOOZE`** (`snooze`) — the agent says it's waiting for an external trigger (a build, deploy, test result, scheduled event, or some other thing it can't control). The session is hidden for a fixed **1 hour** and reappears in your inbox when the snooze expires.
- **`RESPOND`** (`respond`) — the agent asked a question Inbox Pilot can answer from your global + per-session rules (typically a routine yes/no, "go ahead", or short factual reply). Inbox Pilot types the response into the session on your behalf. The message renders with a **"Sent by Inbox Pilot"** attribution badge (the badge tooltip carries the decision id so you can correlate back to the audit row). RESPOND is **off by default** at three layers — the global `Auto-respond` setting in the Inbox Pilot sidebar tile must be on, AND the **Allow Inbox Pilot to reply** checkbox must be checked in that session's three-dot menu, AND the session must be in a non-busy status (not `running` / `starting` / `stalled` / `terminating` / `archived`) at the moment the classifier returns RESPOND. Two anti-runaway gates fire on top of that: a **loop guard** (max 3 applied RESPONDs per session in any rolling 5-minute window) and a **60-second cooldown** between any two applied RESPONDs in the same session. When the classifier returns RESPOND but a gate blocks the actual reply, the decision row is logged with outcome `skipped` and an `error` string identifying the gate (`global-disabled`, `session-disabled`, `session-busy`, `loop-guard`, `cooldown`) so the audit log still tells you what the model wanted to do.
- **`SCHEDULE_RESPOND`** (`scheduled_respond`) — the agent asked a question Inbox Pilot can answer, but sending **now** would be inappropriate (e.g. it's the middle of the night and the agent is fine waiting until morning, or the user has stated a "send tomorrow at 9am" preference). The classifier emits both the reply text and an ISO-8601 `scheduledFor` timestamp; the dispatcher queues a one-shot Send-Later delivery for that time. The decision row records `action_taken = 'scheduled_respond'` and a `scheduled_for` column so the audit log preserves _when_ the queued reply was meant to fire. **Hard cap: 24 hours into the future** — anything further out is rejected at the dispatcher with outcome `failed` and a "scheduledFor is out of range (>24h ahead)" reason. Past times and unparseable timestamps are also rejected on the spot. SCHEDULE_RESPOND is **off by default** behind its own toggle in the Inbox Pilot sidebar tile → **Schedule reply for later** (`aiManager.routerScheduleRespondEnabled`) which is _disabled_ unless Auto-respond is also on — it sits below the Auto-respond switch as a child option. Once enabled, the classifier sees a "Schedule" outcome it can emit; otherwise it falls back to plain RESPOND or one of the other four outcomes. SCHEDULE_RESPOND shares the same per-session opt-in, session-busy gate, loop guard, and cooldown as RESPOND — the throttle counts both action types against the same 3-per-5-min and 60s budgets so a session can't dodge the rate limit by alternating between RESPOND and SCHEDULE_RESPOND. If a session already has a Send-Later reply queued (one slot per session), a second SCHEDULE_RESPOND is logged with outcome `skipped` and error `scheduled-exists` — the existing schedule is _not_ overwritten.

Every decision carries a one-sentence `reason` (max 30 words) which is shown inline on the affected message and in the decisions log. `RESPOND` decisions also carry a `responseText` field (capped at 2,000 characters) — the exact text Inbox Pilot typed into the session. When `RESPOND` is fired via a pre-approved snippet (see next section), the decision row instead carries `snippetId` (the resolved snippet's row id) plus `frozenSnippetText` (the snippet's text captured at dispatch time so the audit trail survives later edits to the saved snippet).

### Snippet replies (pre-approved auto-sends)

When you've curated **Quick Reply snippets** (sidebar **Quick Replies** virtual project → Quick replies tab), Inbox Pilot can pick one by **code** instead of generating freeform reply text. The classifier sees a list of available snippet codes appended to its system prompt and, when one matches the situation exactly, emits a `snippetCode` field (e.g. `SNIP_THANKS_FOR_YOUR_HELP`) instead of `responseText`. The dispatcher looks up the snippet's saved text and types it verbatim.

**Why this exists.** A freeform reply is one new piece of LLM-generated text per turn. A snippet reply is _exactly_ a string you already wrote and approved. For routine acknowledgements ("thanks, please continue", "yes go ahead", "looks good") this is both cheaper and safer: the text is fixed, has been read by you at least once, and changes only when you edit the snippet.

**Which snippets are eligible.** A snippet is shown to the Inbox Pilot classifier only when **all** of the following are true:

- It's a regular snippet (not a folder or divider).
- **Auto-submit is on** — the snippet sends immediately when typed, with no edit-then-confirm step.
- **The text contains no `{var}` template placeholders** — Inbox Pilot can't fill in template variables, so snippets that require them are excluded entirely.
- **The "Available to Inbox Pilot" toggle on the snippet is on.** This is on by default for every snippet; flipping it off in the snippet editor removes that snippet from the classifier's view while leaving it available for manual quick-reply use. (As of 2026-07-20 this toggle is **hidden in the snippet editor when Inbox Pilot is retired/off** — `inboxPilotEnabled` false — while the stored value is preserved for a future re-enable.)

**What the classifier sees.** Each eligible snippet appears in the prompt as `- SNIP_<slugified-label>: <description>`. The description is the snippet's optional **"Inbox Pilot description"** field if you've set one (a short prose hint like "close out a completed task"); when blank, Inbox Pilot falls back to the first ~80 characters of the snippet's own text. The prompt is capped at **25 snippets**, ordered by your existing sidebar `displayOrder` — if you have more than 25 eligible snippets, the ones at the bottom of your list are silently dropped from the prompt (still usable manually). Snippets whose label has no usable alphanumeric content (emoji-only, e.g. `👍`) can't be slugified and are dropped from the prompt; they remain visible as manual quick-replies.

**Per-snippet opt-out.** Open the snippet editor (sidebar **Quick Replies** virtual project → Quick replies tab → click a snippet, or right-click any snippet in the sidebar and choose **Edit**) and you'll find an **"Available to Inbox Pilot"** toggle with a description field below it. Turning the toggle off makes the snippet invisible to the classifier; the snippet still works for manual sends. Use this for snippets that are situational, sarcastic, or that you don't want fired without you reading the message first.

**Fail-closed mid-flight.** The eligibility check happens at two points: when the prompt is built (the classifier only sees eligible snippets) AND again at dispatch time, _after_ all the RESPOND gates pass. If you toggle a snippet's "Available to Inbox Pilot" off between the moment the classifier emits its code and the moment the dispatcher tries to fire — or if you flip its auto-submit off, or add a `{var}` placeholder — the dispatcher refuses to fire and the decision row is logged with outcome `failed` and error `snippet code not found or no longer eligible`. The session stays in your inbox waiting for human input. There is no silent fallback to "send anyway".

**Audit trail freezes the text at dispatch time.** When a snippet reply fires, the decision row stores both the snippet's id (so you can trace back to the row) and a copy of the snippet's text at the moment of dispatch (`frozenSnippetText`). Editing the snippet later changes future fires but does not rewrite the audit log — the row still shows what was actually typed into the session.

**No `use_count` bump on auto-fires.** The snippet `use_count` shown in the Quick Replies editor counts **manual** fires only. Inbox Pilot's auto-sends deliberately do not increment it — otherwise classifier-driven fires would bias the user-curated `displayOrder` ranking, and snippets that won a 25-cap slot would self-reinforce that win. `use_count` stays a pure manual-usage signal.

### Per-session rule

In addition to the global Inbox Pilot rule you set in Settings, each session can carry its own **per-session rule** that you author in the session's three-dot menu (see "What you see" above). The session rule is appended to the prompt **after** the global rule and **before** the agent transcript — the classifier reads global → session-specific → transcript in that order, with the transcript marked as untrusted content.

The session rule is framed to the classifier as a _modifier_, not a replacement. The global rule still applies; the session rule narrows or specializes it for that session only. A session rule like _"only flag failed deploys"_ won't silence every message — it tells the classifier that, when in doubt for this particular session, lean toward `HIDE` unless the message clearly meets the session-specific criterion.

The audit decisions log captures a `promptHash` that combines both rules. A session with the same global rule, same transcript, but a different session rule produces a different hash — so audit rows distinguish "decision was made under global-only" from "decision was made under global + session". Empty or whitespace-only session prompts normalize to a "no per-session layer" hash so toggling between "cleared" and "never set" stays stable.

### Cost and pause controls

Inbox Pilot has its **own** independent set of cost and pause knobs — they are not shared with Plain Speak. There are four layers of control:

- **Daily cost cap (default $1.00).** Inbox Pilot has its own running daily total in USD, persisted in `aiManager.routerDailyCapUsd`. When the cap is crossed, Inbox Pilot stops firing for the rest of the day; Plain Speak (which has its own separate cap) is unaffected. The Inbox Pilot tile's status banner shows today's spend so you can see the running total at a glance. Edit the cap in the **Cost & access** card on the same tile (range 0–100, displayed with a visual `$` prefix so it reads "$1" not "1"). When a cap is hit, Omniscio shows a one-time-per-rule-per-day toast whose action button jumps straight to the Inbox Pilot tile. Capped fires are recorded in the decisions log with outcome `capped` so you can audit what got skipped.
- **Master on/off toggle.** A single **Enable Inbox Pilot** toggle sits at the very top of the Inbox Pilot tile (it flips `aiManager.routerPaused` under the hood). When the toggle is **off**, Inbox Pilot does not fire for any session AND the rest of the tile — shadow mode, default-on-for-new, bulk apply, auto-respond, schedule-respond, daily cap, allow-on-API-key, rule editor, recent activity — is hidden entirely so the tile collapses to just the toggle. The toggle's description text reads **"Turn on to enable everything below."** This pause is independent of Plain Speak's pause flag.
  - **Turning the master OFF is a deliberate two-step (mirror of the shadow-mode-OFF flow).** Because every sub-setting that's currently on would silently survive an OFF cycle and re-fire the moment the user flipped the master back on, Omniscio prompts a **"Turn Inbox Pilot off?"** confirm dialog with a warning. On confirm, Omniscio first calls `bulkDisableAll()` to flip `ai_router_enabled = 0` on every session in the database (the same one-click rollback that the "Disable on all sessions" button performs), THEN persists the master setting alongside a full sub-setting reset: shadow mode is forced back ON (the safe observe-only default), default-on-for-new-sessions OFF, auto-respond OFF, schedule-respond OFF, and the routing rule's own enable switch OFF. **Your rule prompt text is preserved** — only the boolean gates reset, so you don't lose the rule you spent time tuning. Cancelling the confirm leaves the master ON and changes nothing. If the `bulkDisableAll()` step fails, the master setting is **not** persisted and an error toast surfaces so you can retry — the master is never flipped to OFF without the bulk-disable having succeeded first. After a reset, turning the master back ON is a simple one-click flip (no confirm) but every sub-setting you want will need a deliberate re-toggle.
- **Master enable.** The rule editor in the tile has its own enable toggle. Turning Inbox Pilot off here disables it for **every** session at once.
- **Per-session enable.** Even when the pause is off and the master enable is on, Inbox Pilot still won't fire on a given session unless that session's three-dot menu has the **Inbox Pilot** checkbox checked. This is the gate that stops Inbox Pilot from acting on sessions you didn't opt in.

By default, Inbox Pilot only runs on **OAuth** accounts. To allow it on API-key accounts (where every call is billed directly to your API key), enable the **Allow on API-key accounts** toggle in the tile's Cost & access card (`aiManager.routerAllowApiKey`) — this permission is also Inbox-Pilot-specific and does not affect Plain Speak.

## Related

How the Pilot behaves once it is running — observing first, the pill that tells you it looked, the decisions log, privacy and its known gaps — continues on [Inbox Pilot — running it, privacy and limits](inbox-pilot-part-2.md). The inbox it triages is described on the [Inbox overview](inbox-overview.md) page.

- [plain-speak.md](plain-speak.md) — the sibling feature that rewrites the latest agent message in plain English; same backend, separate cap / pause / api-key permission. Plain Speak's settings live in Settings → Plain Speak (its own section)
- [catch-up-card.md](catch-up-card.md) — the pinned 4-line summary at the top of a Needs You session (**currently disabled** as of 2026-04-30 — code preserved, Settings UI unwired)
- [coaching-engine.md](coaching-engine.md) — the other AI-driven nudge system in Omniscio, but for teaching you the app rather than triaging your inbox
- [automations-and-auto-replies.md](automations-and-auto-replies.md) — for autonomous actions on incoming messages (which Inbox Pilot intentionally doesn't do)
- [inbox-overview.md](inbox-overview.md) — the unified inbox that Inbox Pilot's decisions feed into
- [ai-providers.md](ai-providers.md) — how Omniscio routes AI features between Anthropic, Groq, Together, and OpenRouter. Inbox Pilot deliberately pins its classifier to OpenRouter (Qwen primary + Haiku fallback) for JSON-mode reliability — see "Evaluator routing" above
