Inbox Pilot (an AI that triages your inbox for you)
An AI that reads the sessions waiting for you in your inbox and handles the routine ones — replying, archiving, escalating — so you only see the items that genuinely need a person. It can be run in observe-only mode first, is bounded by cost and pause controls, and is per-session opt-in.
What it is
Inbox Pilot is a cheap classifier LLM call (Qwen3-32B via OpenRouter by default, with a one-shot Haiku fallback) that runs at the end of every Claude session turn and decides what should happen to the session in your inbox. It looks at the latest agent end-of-turn and classifies it into one of six outcomes — keep it as Needs You, hide it, archive it, snooze it for an hour, respond to the agent on your behalf, or schedule a response to be sent later (up to 24h ahead) — and writes a one-sentence reason for every decision.
Inbox Pilot is off by default at three levels and opt-in per session:
- Top-level flag
inboxPilotEnabled(hidden / config-only): the master kill switch. Shipped 2026-05-23 as a Settings → Features → Enable Inbox Pilot toggle, but as of 2026-06-30 that toggle and its Settings-search entry were removed — the flag is now a completely hidden, config-only setting (default off; settable only by editingconfig.jsonor via the CLI control server'sPATCH /settings/inboxPilotEnabled, never from the app UI). Its gating behaviour is unchanged: when this is off, the Inbox Pilot sidebar row is hidden, the per-session three-dot menu drops its Inbox Pilot / Custom rule for this session / Allow Inbox Pilot to reply / Test on latest message… rows entirely (2026-05-24 — previously they kept rendering and toggling them silently failed against the gated IPC), theaiManagerService.evaluate()end-of-turn pipeline returns immediately as a no-op (no Haiku calls, no audit rows, no overlays), and every Inbox Pilot IPC handler (rules, decisions, preview, per-session toggles, bulk apply, daily cost, pending overlays) rejects with{success: false, error: "Inbox Pilot is disabled..."}. As of 2026-07-20 — when Plain Speak split into its own Settings section — Inbox Pilot was retired from the UI entirely: the opt-in auto-enable was removed fromsrc/main/services/opt-in-toggle-migration.ts, and a one-shotsrc/main/services/migrations/inbox-pilot-retire-migration.ts(own cookieinboxPilotRetireMigrationDone) resets the flag OFF for any install that still had it on, soinboxPilotEnabledis now uniformly false. The code + flag are kept dormant, re-enable-able only viaconfig.json/ the CLI. - Master enable inside the sidebar tile — at the top of the Inbox Pilot page.
- Per-session checkbox in each session's three-dot menu.
It runs only on sessions where all three are on, and only while the daily cap and pause flag also allow it. Every decision Inbox Pilot makes is audited in a decisions log so you can see exactly what it did, why, what it cost, and which model it used. Inbox Pilot is no longer surfaced in Settings search — the feature toggle and its search-index entry were removed on 2026-06-30, and as of 2026-07-20 it is retired from the UI everywhere: the sidebar tile stays hidden, the per-session menus drop their Inbox-Pilot rows, and the Quick-Reply "Allow Inbox Pilot to auto-send" card is hidden too — all gated on inboxPilotEnabled, now uniformly false (the flag lives on config-only).
What the classifier sees
Each Inbox Pilot call sends the classifier a narrow transcript — only the latest agent message and the operator (user) message immediately preceding it, formatted as [operator] …\n[agent] …. Earlier turns in the session are not included. The system prompt already directed the model to classify the latest agent turn and treat earlier turns as background context, so sending the rest was wasted tokens and an extra prompt-injection surface for older agent output the model wasn't supposed to weigh heavily.
Three follow-on details worth knowing:
systemrows are excluded as candidates for the operator slot. These are CLI/Omniscio bookkeeping (compact-divider, recovery banners), not operator input — including them would let the classifier reason about meta-events as if they were user instructions.- Recipe-injected operator rows ARE included. When a recipe or automation injects a prompt and the agent replies, the classifier needs to see what the agent was responding to even though no human typed it. Filtering them would make the agent's reply read as a non-sequitur.
- An operator typed AFTER the latest agent reply is not paired with that agent. The pairing always picks the operator strictly BEFORE the agent's timestamp — an unanswered operator is a separate new turn awaiting its own agent reply.
If the latest agent message alone exceeds the classifier's 128k-char input budget, it's truncated from the end with an explicit … [truncated, full N chars] marker so the model knows the message was clipped. If operator + agent combined exceeds the budget, the operator is dropped and only the agent line is sent.
Plain Speak (the sibling overlay feature) uses a wider transcript — it's paraphrasing a longer span and has its own context builder. The narrowing described here applies to Inbox Pilot only.
Where to find it
In your inbox, on the sessions it acts on — each one is marked so you can see the Pilot was there. The controls that govern it live in Settings → Inbox Pilot, and the per-session switch is on the session itself.
How it behaves
What you see
You see Inbox Pilot in three places:
- The Inbox Pilot tile in the left sidebar. Click the airplane (plane-takeoff) icon in the projects bar and a full-page Inbox Pilot view opens. At the very top is a single Enable Inbox Pilot master toggle — flip it off and the entire rest of the tile collapses away (only the toggle remains visible). Flipping it off prompts a confirm dialog and, on confirm, resets every sub-setting to its safe default while preserving your rule prompt text — see the Master on/off toggle bullet under "Cost and pause controls" below for the full reset behavior. With the master toggle on, the tile reveals: a rollout-controls block — a Shadow mode toggle, a Default on for new sessions toggle, a Bulk apply card with the Enable on N active sessions and Disable on all sessions buttons, the Auto-respond toggle, and the Schedule reply for later toggle. When today's spend reaches the daily cap, an amber Daily cap reached notice appears at the top of this block. Below the rollout controls is a Cost & access card with the daily cost cap input (range 0–100, displayed with a visual
$prefix so it reads "$1" not "1"), today's spend bar shown inline beneath the cap input, and the Allow on API-key accounts toggle. The rule editor below has its own enable switch, a multi-line prompt textbox where you tune the rule in your own words, and a Preview button that runs the rule against a real session in dry-run mode without applying any side effects. Recent decisions are listed at the bottom (most recent first). - Each session's three-dot overflow menu. Open the kebab menu on any session row and you'll see an Inbox Pilot entry as a top-level item with a checkbox you can flip on or off for that session. It defaults to off for every session — unless you've turned Default on for new sessions on in the Inbox Pilot tile, in which case new sessions arrive with the checkbox already on. The helper text under the checkbox tells you the per-turn cost (~$0.0001 per turn) when the toggle is on. This is a per-session opt-in; flipping a master switch in Settings doesn't enable a session by itself. When Inbox Pilot is on for a session, a "Custom rule for this session" disclosure appears directly below the checkbox. Expanding it reveals a textarea where you can add a session-specific instruction that layers on top of your global rule — for example, "For this session, only flag me if a deploy fails. Otherwise hide." Edits are debounced 500ms and saved automatically; max 8000 chars; leaving it empty disables the per-session layer and falls back to the global rule alone. Below the custom-rule disclosure is a "Test on latest message…" menu item with a flask icon — clicking it opens a dry-run modal that runs Inbox Pilot against this specific session's transcript right now, layering in the saved global rule and this session's saved per-session rule (mirroring what the live router would do). The result shows the action Pilot would take and the one-sentence reason, so you can sanity-check that your rule does what you expect before letting it act for real. The test runs purely free — no action fires, the daily cap is not charged, the result is written to the decisions log with
was_dry_run = 1(and prunes after 7 days like all preview rows). When Inbox Pilot is off for the session, the test row is greyed out with a tooltip telling you to turn the session toggle on first. - The "Pilot took a look" pill in the chat. When Inbox Pilot fires on a session — whether in shadow mode or real mode — the agent turn it evaluated gets a small expandable pill rendered inline in the chat (one pill per turn, anchored after the agent's reply). Closed, the pill shows a status dot whose colour carries the outcome (green = applied, amber = cap hit, red = failed, grey = skipped or in-flight) and a label whose phrasing carries the action ("Pilot archived this thread" in real mode, "Pilot would have archived this thread" in shadow mode). Click to expand and the pill reveals the LLM's one-sentence reason, the model used, the call's cost, and the first 8 characters of the decision id. See "The 'Pilot took a look' pill" section below for the full anatomy.
The six outcomes
Inbox Pilot classifies each agent end-of-turn into exactly one of these six labels (the model emits the upper-case label; the value persists in the decisions log as the lowercase enum value shown in parentheses):
KEEP_NEEDS_YOU(keep_needs_you) — the agent has asked the user a question or hit a blocker that needs human input. The session stays in your Needs You queue. This is the default when Inbox Pilot is unsure.HIDE(hide) — the message is informational, mid-task narration, or "thinking out loud" that the user doesn't need to act on. The session's status flips toendedso it drops out of the attention list. Sending a new message to the session reopens it as normal.ARCHIVE(archive) — the agent has explicitly finished its task with no outstanding issues. The session is soft-deleted with the standard archive flow (you can undo from the toast). Auto-unarchives on the next send.SNOOZE(snooze) — the agent says it's waiting for an external trigger (a build, deploy, test result, scheduled event, or some other thing it can't control). The session is hidden for a fixed 1 hour and reappears in your inbox when the snooze expires.RESPOND(respond) — the agent asked a question Inbox Pilot can answer from your global + per-session rules (typically a routine yes/no, "go ahead", or short factual reply). Inbox Pilot types the response into the session on your behalf. The message renders with a "Sent by Inbox Pilot" attribution badge (the badge tooltip carries the decision id so you can correlate back to the audit row). RESPOND is off by default at three layers — the globalAuto-respondsetting in the Inbox Pilot sidebar tile must be on, AND the Allow Inbox Pilot to reply checkbox must be checked in that session's three-dot menu, AND the session must be in a non-busy status (notrunning/starting/stalled/terminating/archived) at the moment the classifier returns RESPOND. Two anti-runaway gates fire on top of that: a loop guard (max 3 applied RESPONDs per session in any rolling 5-minute window) and a 60-second cooldown between any two applied RESPONDs in the same session. When the classifier returns RESPOND but a gate blocks the actual reply, the decision row is logged with outcomeskippedand anerrorstring identifying the gate (global-disabled,session-disabled,session-busy,loop-guard,cooldown) so the audit log still tells you what the model wanted to do.SCHEDULE_RESPOND(scheduled_respond) — the agent asked a question Inbox Pilot can answer, but sending now would be inappropriate (e.g. it's the middle of the night and the agent is fine waiting until morning, or the user has stated a "send tomorrow at 9am" preference). The classifier emits both the reply text and an ISO-8601scheduledFortimestamp; the dispatcher queues a one-shot Send-Later delivery for that time. The decision row recordsaction_taken = 'scheduled_respond'and ascheduled_forcolumn so the audit log preserves when the queued reply was meant to fire. Hard cap: 24 hours into the future — anything further out is rejected at the dispatcher with outcomefailedand a "scheduledFor is out of range (>24h ahead)" reason. Past times and unparseable timestamps are also rejected on the spot. SCHEDULE_RESPOND is off by default behind its own toggle in the Inbox Pilot sidebar tile → Schedule reply for later (aiManager.routerScheduleRespondEnabled) which is disabled unless Auto-respond is also on — it sits below the Auto-respond switch as a child option. Once enabled, the classifier sees a "Schedule" outcome it can emit; otherwise it falls back to plain RESPOND or one of the other four outcomes. SCHEDULE_RESPOND shares the same per-session opt-in, session-busy gate, loop guard, and cooldown as RESPOND — the throttle counts both action types against the same 3-per-5-min and 60s budgets so a session can't dodge the rate limit by alternating between RESPOND and SCHEDULE_RESPOND. If a session already has a Send-Later reply queued (one slot per session), a second SCHEDULE_RESPOND is logged with outcomeskippedand errorscheduled-exists— the existing schedule is not overwritten.
Every decision carries a one-sentence reason (max 30 words) which is shown inline on the affected message and in the decisions log. RESPOND decisions also carry a responseText field (capped at 2,000 characters) — the exact text Inbox Pilot typed into the session. When RESPOND is fired via a pre-approved snippet (see next section), the decision row instead carries snippetId (the resolved snippet's row id) plus frozenSnippetText (the snippet's text captured at dispatch time so the audit trail survives later edits to the saved snippet).
Snippet replies (pre-approved auto-sends)
When you've curated Quick Reply snippets (sidebar Quick Replies virtual project → Quick replies tab), Inbox Pilot can pick one by code instead of generating freeform reply text. The classifier sees a list of available snippet codes appended to its system prompt and, when one matches the situation exactly, emits a snippetCode field (e.g. SNIP_THANKS_FOR_YOUR_HELP) instead of responseText. The dispatcher looks up the snippet's saved text and types it verbatim.
Why this exists. A freeform reply is one new piece of LLM-generated text per turn. A snippet reply is exactly a string you already wrote and approved. For routine acknowledgements ("thanks, please continue", "yes go ahead", "looks good") this is both cheaper and safer: the text is fixed, has been read by you at least once, and changes only when you edit the snippet.
Which snippets are eligible. A snippet is shown to the Inbox Pilot classifier only when all of the following are true:
- It's a regular snippet (not a folder or divider).
- Auto-submit is on — the snippet sends immediately when typed, with no edit-then-confirm step.
- The text contains no
{var}template placeholders — Inbox Pilot can't fill in template variables, so snippets that require them are excluded entirely. - The "Available to Inbox Pilot" toggle on the snippet is on. This is on by default for every snippet; flipping it off in the snippet editor removes that snippet from the classifier's view while leaving it available for manual quick-reply use. (As of 2026-07-20 this toggle is hidden in the snippet editor when Inbox Pilot is retired/off —
inboxPilotEnabledfalse — while the stored value is preserved for a future re-enable.)
What the classifier sees. Each eligible snippet appears in the prompt as - SNIP_<slugified-label>: <description>. The description is the snippet's optional "Inbox Pilot description" field if you've set one (a short prose hint like "close out a completed task"); when blank, Inbox Pilot falls back to the first ~80 characters of the snippet's own text. The prompt is capped at 25 snippets, ordered by your existing sidebar displayOrder — if you have more than 25 eligible snippets, the ones at the bottom of your list are silently dropped from the prompt (still usable manually). Snippets whose label has no usable alphanumeric content (emoji-only, e.g. 👍) can't be slugified and are dropped from the prompt; they remain visible as manual quick-replies.
Per-snippet opt-out. Open the snippet editor (sidebar Quick Replies virtual project → Quick replies tab → click a snippet, or right-click any snippet in the sidebar and choose Edit) and you'll find an "Available to Inbox Pilot" toggle with a description field below it. Turning the toggle off makes the snippet invisible to the classifier; the snippet still works for manual sends. Use this for snippets that are situational, sarcastic, or that you don't want fired without you reading the message first.
Fail-closed mid-flight. The eligibility check happens at two points: when the prompt is built (the classifier only sees eligible snippets) AND again at dispatch time, after all the RESPOND gates pass. If you toggle a snippet's "Available to Inbox Pilot" off between the moment the classifier emits its code and the moment the dispatcher tries to fire — or if you flip its auto-submit off, or add a {var} placeholder — the dispatcher refuses to fire and the decision row is logged with outcome failed and error snippet code not found or no longer eligible. The session stays in your inbox waiting for human input. There is no silent fallback to "send anyway".
Audit trail freezes the text at dispatch time. When a snippet reply fires, the decision row stores both the snippet's id (so you can trace back to the row) and a copy of the snippet's text at the moment of dispatch (frozenSnippetText). Editing the snippet later changes future fires but does not rewrite the audit log — the row still shows what was actually typed into the session.
No use_count bump on auto-fires. The snippet use_count shown in the Quick Replies editor counts manual fires only. Inbox Pilot's auto-sends deliberately do not increment it — otherwise classifier-driven fires would bias the user-curated displayOrder ranking, and snippets that won a 25-cap slot would self-reinforce that win. use_count stays a pure manual-usage signal.
Per-session rule
In addition to the global Inbox Pilot rule you set in Settings, each session can carry its own per-session rule that you author in the session's three-dot menu (see "What you see" above). The session rule is appended to the prompt after the global rule and before the agent transcript — the classifier reads global → session-specific → transcript in that order, with the transcript marked as untrusted content.
The session rule is framed to the classifier as a modifier, not a replacement. The global rule still applies; the session rule narrows or specializes it for that session only. A session rule like "only flag failed deploys" won't silence every message — it tells the classifier that, when in doubt for this particular session, lean toward HIDE unless the message clearly meets the session-specific criterion.
The audit decisions log captures a promptHash that combines both rules. A session with the same global rule, same transcript, but a different session rule produces a different hash — so audit rows distinguish "decision was made under global-only" from "decision was made under global + session". Empty or whitespace-only session prompts normalize to a "no per-session layer" hash so toggling between "cleared" and "never set" stays stable.
Cost and pause controls
Inbox Pilot has its own independent set of cost and pause knobs — they are not shared with Plain Speak. There are four layers of control:
- Daily cost cap (default $1.00). Inbox Pilot has its own running daily total in USD, persisted in
aiManager.routerDailyCapUsd. When the cap is crossed, Inbox Pilot stops firing for the rest of the day; Plain Speak (which has its own separate cap) is unaffected. The Inbox Pilot tile's status banner shows today's spend so you can see the running total at a glance. Edit the cap in the Cost & access card on the same tile (range 0–100, displayed with a visual$prefix so it reads "$1" not "1"). When a cap is hit, Omniscio shows a one-time-per-rule-per-day toast whose action button jumps straight to the Inbox Pilot tile. Capped fires are recorded in the decisions log with outcomecappedso you can audit what got skipped. - Master on/off toggle. A single Enable Inbox Pilot toggle sits at the very top of the Inbox Pilot tile (it flips
aiManager.routerPausedunder the hood). When the toggle is off, Inbox Pilot does not fire for any session AND the rest of the tile — shadow mode, default-on-for-new, bulk apply, auto-respond, schedule-respond, daily cap, allow-on-API-key, rule editor, recent activity — is hidden entirely so the tile collapses to just the toggle. The toggle's description text reads "Turn on to enable everything below." This pause is independent of Plain Speak's pause flag.- Turning the master OFF is a deliberate two-step (mirror of the shadow-mode-OFF flow). Because every sub-setting that's currently on would silently survive an OFF cycle and re-fire the moment the user flipped the master back on, Omniscio prompts a "Turn Inbox Pilot off?" confirm dialog with a warning. On confirm, Omniscio first calls
bulkDisableAll()to flipai_router_enabled = 0on every session in the database (the same one-click rollback that the "Disable on all sessions" button performs), THEN persists the master setting alongside a full sub-setting reset: shadow mode is forced back ON (the safe observe-only default), default-on-for-new-sessions OFF, auto-respond OFF, schedule-respond OFF, and the routing rule's own enable switch OFF. Your rule prompt text is preserved — only the boolean gates reset, so you don't lose the rule you spent time tuning. Cancelling the confirm leaves the master ON and changes nothing. If thebulkDisableAll()step fails, the master setting is not persisted and an error toast surfaces so you can retry — the master is never flipped to OFF without the bulk-disable having succeeded first. After a reset, turning the master back ON is a simple one-click flip (no confirm) but every sub-setting you want will need a deliberate re-toggle.
- Turning the master OFF is a deliberate two-step (mirror of the shadow-mode-OFF flow). Because every sub-setting that's currently on would silently survive an OFF cycle and re-fire the moment the user flipped the master back on, Omniscio prompts a "Turn Inbox Pilot off?" confirm dialog with a warning. On confirm, Omniscio first calls
- Master enable. The rule editor in the tile has its own enable toggle. Turning Inbox Pilot off here disables it for every session at once.
- Per-session enable. Even when the pause is off and the master enable is on, Inbox Pilot still won't fire on a given session unless that session's three-dot menu has the Inbox Pilot checkbox checked. This is the gate that stops Inbox Pilot from acting on sessions you didn't opt in.
By default, Inbox Pilot only runs on OAuth accounts. To allow it on API-key accounts (where every call is billed directly to your API key), enable the Allow on API-key accounts toggle in the tile's Cost & access card (aiManager.routerAllowApiKey) — this permission is also Inbox-Pilot-specific and does not affect Plain Speak.
Related
How the Pilot behaves once it is running — observing first, the pill that tells you it looked, the decisions log, privacy and its known gaps — continues on Inbox Pilot — running it, privacy and limits. The inbox it triages is described on the Inbox overview page.
- plain-speak.md — the sibling feature that rewrites the latest agent message in plain English; same backend, separate cap / pause / api-key permission. Plain Speak's settings live in Settings → Plain Speak (its own section)
- catch-up-card.md — the pinned 4-line summary at the top of a Needs You session (currently disabled as of 2026-04-30 — code preserved, Settings UI unwired)
- coaching-engine.md — the other AI-driven nudge system in Omniscio, but for teaching you the app rather than triaging your inbox
- automations-and-auto-replies.md — for autonomous actions on incoming messages (which Inbox Pilot intentionally doesn't do)
- inbox-overview.md — the unified inbox that Inbox Pilot's decisions feed into
- ai-providers.md — how Omniscio routes AI features between Anthropic, Groq, Together, and OpenRouter. Inbox Pilot deliberately pins its classifier to OpenRouter (Qwen primary + Haiku fallback) for JSON-mode reliability — see "Evaluator routing" above
Last verified 2026-09-28