---
title: Inbox Pilot (an AI that triages your inbox for you) (part 2)
---

# Inbox Pilot (an AI that triages your inbox for you) (part 2)

## What it is

This is part 2 of the [Inbox Pilot](inbox-pilot.md) page. It covers everything around the Pilot running rather than what it decides: watching it work before you trust it, switching it on for new or existing sessions, the record of what it did, what leaves your machine, and where it falls short.

## Where to find it

The same places as the main page — the **inbox** and **Settings → Inbox Pilot** — plus the per-session controls, which are where the observe-only and bulk switches actually take effect.

## How it behaves

### Shadow mode (observe only)

The Inbox Pilot tile has a master toggle called **Shadow mode** (`aiManager.aiRouterDryRunByDefault`). When shadow mode is on, the Haiku classifier still runs on every agent turn — it still spends money against your daily cap, it still writes a row to the decisions log, it still picks one of the six outcomes — but the dispatcher short-circuits **before** any actual archive, hide, snooze, response, or scheduled-respond side effect. The session stays exactly where it was; only the audit record changes.

Shadow mode is the recommended way to roll Inbox Pilot out across your inbox. Turn shadow on, flip Inbox Pilot on for every session (the bulk-enable button below makes this one click), then watch the **"Pilot took a look"** pills accumulate for a day or two and confirm the Pilot's judgment matches yours before letting it act for real. Each shadow-mode decision renders an inline pill that reads in the conditional — "Pilot would have archived this thread", "Pilot would have drafted a reply" — and carries a small uppercase **`shadow`** badge so you can never mistake a shadow run for a real one. The daily cost cap still applies in shadow mode (the LLM call is the same call whether or not the action fires), and capped fires still log as `outcome = 'capped'` for the audit trail.

**Turning shadow mode OFF is a deliberate two-step**, by design. Because turning shadow off means the Pilot's _next_ decision on every opted-in session will fire for real, Omniscio prompts a **"Turn shadow mode off?"** confirm dialog with a warning. On confirm, Omniscio auto-disables Inbox Pilot on every session in the database (the same one-click rollback that the "Disable on all sessions" button performs) **before** the shadow setting itself flips. This forces you to deliberately re-opt-in per session before any real archive, hide, snooze, respond, or scheduled-respond can fire — preventing the worst-case scenario where flipping shadow off would cause fifty threads to suddenly archive themselves the moment the gate drops. Cancelling the confirm leaves shadow mode on and changes nothing. If the bulk-disable step fails for any reason (e.g. a database write error), shadow mode is left on and an error toast surfaces so you can retry — the shadow setting is never persisted to the OFF state without the bulk-disable having succeeded first.

### Default on for new sessions

Below the shadow toggle is **Default on for new sessions** (`aiManager.aiRouterDefaultOnForNewSessions`). When on, every newly created session gets Inbox Pilot turned on at creation — the per-session opt-in flag is set to 1 in the row at insert time. Existing sessions are unaffected by toggling this flag (use the bulk buttons below for those). When off (the default), new sessions are created with Inbox Pilot off and you have to opt each one in manually from its three-dot menu.

Silent recipe sessions are deliberately excluded from this default — even with the flag on, a silent session (one created by a recipe that the user never sees in the inbox) is created with Inbox Pilot off. Silent sessions never reach the inbox, so evaluating their turns is pure wasted cost.

### Bulk enable / Bulk disable

Below the toggles is a **Bulk apply** card with two buttons:

- **Enable on N active sessions** flips the per-session Inbox Pilot opt-in to ON for every eligible session in one click. The `N` is a live count of eligible sessions visible on the button label; when it's zero the button is disabled with a tooltip explaining why. The success toast reports the real server count (the button's `N` is a UI estimate that can drift by one or two if a session just transitioned status mid-render). When shadow mode is on the toast reads "Inbox Pilot observing N sessions"; when shadow is off it reads "Inbox Pilot enabled on N sessions" — same database operation, different framing chosen to match the user's mental model.
- **Disable on all sessions** is the one-click rollback. It prompts a confirm dialog first, then flips the per-session opt-in to OFF on every session in the database. Existing decision-log rows are kept untouched; only the per-session toggles get cleared. This is exactly the same operation that fires automatically when you turn shadow mode off.

The bulk-enable scope deliberately excludes silent recipe sessions (the silent flag is sticky — attention statuses like `needs_you` / `error` / `stalled` do not clear it). Sessions in terminal statuses (`archived` / `ended` / `error` / `paused`) are also skipped — they aren't going to receive new agent turns, so there's nothing for the Pilot to evaluate. Bulk-enable is idempotent: clicking it again immediately after a successful run updates zero rows and shows "No active sessions to enable Inbox Pilot on" (rather than re-touching already-enabled rows).

### The "Pilot took a look" pill

When Inbox Pilot fires on a session — whether in shadow mode or real mode — the agent turn it evaluated gets a small **"Pilot took a look"** pill rendered inline in the chat. The pill anchors to the turn, not to any individual agent message inside it, so each turn has at most one pill no matter how many tool calls or sub-messages the agent emitted in that turn.

The pill is closed by default. Closed, it shows three elements left-to-right:

- A chevron (▸) that rotates 90° when expanded.
- A small status dot whose colour tells you the **outcome**: green for `applied` (the action — real or would-have-been-real — was successfully resolved), amber for `capped` (Pilot didn't run because the daily cap was hit), red for `failed` (the LLM call errored), grey for `skipped` (a gate refused the action — see RESPOND's session-busy / loop-guard / cooldown gates above), and a pulsing grey for `in_flight` (Pilot is still thinking — the LLM call hasn't returned yet).
- A label whose phrasing tells you the **action**. In real mode you see past-tense: "Pilot kept this in your inbox", "Pilot hid this thread", "Pilot archived this thread", "Pilot snoozed this thread", "Pilot drafted a reply", "Pilot scheduled a reply". In shadow mode you see the conditional: "Pilot would have kept this in your inbox", "Pilot would have hidden this thread", and so on. Non-applied outcomes get their own labels: "Pilot is taking a look…", "Pilot couldn't evaluate this turn", "Pilot skipped this turn", "Pilot hit its daily spend cap".

When the outcome is `applied` and the decision was a shadow run (`was_dry_run = true`), a small uppercase **`shadow`** badge sits to the right of the label. The badge never appears on `in_flight` (you don't yet know whether the action will be a real or shadow result by the time the call finishes) or on the failure / cap / skip outcomes (those didn't result in an applied action either way).

Clicking the pill expands it inline, revealing four read-only pieces of context: the model used (e.g. `claude-haiku-4-5`), the cost of the call (formatted to four decimals — or `<$0.0001` for sub-fractional calls, or `—` for the rare row where cost couldn't be captured), the first 8 characters of the decision id (a stable handle you can correlate with a row in the decisions log at the bottom of the Inbox Pilot tile), and — most importantly — the Pilot's one-sentence `reason` for the action it picked. If the outcome is `failed` and an error string was captured, the error renders in red beneath the reason.

Five of the six possible outcomes render a pill: `applied`, `in_flight`, `failed`, `skipped`, `capped`. The sixth — `cancelled` — does not render at all by design, because cancellations are pre-emption events the user does not need to see (the dispatcher cancels the in-flight row when a newer agent turn arrives in the same session before the LLM call returns; the next turn's pill is the one that matters).

The pill is keyboard-accessible: it's a native `<button>`, so Tab focuses it, Enter or Space toggles it open, and the `aria-expanded` attribute reflects the open/closed state for screen readers. Each pill carries a stable `data-decision-id` attribute matching the decisions-log row id, so tests and external tools can pinpoint a specific Pilot decision by id without having to scrape text.

### Decisions log

Every Inbox Pilot invocation writes a row to the decisions log with `feature = 'router'` (the internal feature identifier — Inbox Pilot is the user-facing label). The row captures: the action taken, the LLM's reasoning, input character count, prompt hash, model used, tokens in/out, cost in USD, outcome (`in_flight` / `applied` / `cancelled` / `failed` / `skipped` / `capped`), error string if any, and whether it was a dry-run preview.

Retention runs on a daily prune:

- **Real (non-dry-run) decisions** are kept for **90 days**.
- **Dry-run preview decisions** are kept for **7 days** (they're noisy and short-lived).
- **Per-session cap** of **1000 most recent rows per session** — older rows beyond that count get pruned even within the 90-day window, so a single very-busy session can't fill the table.
- Orphaned `in_flight` rows older than 5 minutes are reaped at startup and marked `failed` with `error = "orphan reaped at startup"`.

You can see decisions in two places: an inline expandable card on the affected message, and the full list at the bottom of the Inbox Pilot sidebar tile (most recent first). A **View in decisions log** button on each inline card jumps to that exact row.

### Privacy

A few hard rules about what Inbox Pilot does and does not store:

- **Your rule prompt text is NEVER logged in feature events.** The five tracked feature events (`ai_manager_router`, `ai_manager_overlay`, `ai_manager_session_enabled`, `ai_manager_preview`, `ai_manager_pause_all`) carry only a small allow-list of metadata: `outcome`, `action`, `model`, `was_dry_run`, `feature`, `paused`. The prompt itself, the agent transcript, and the LLM's reasoning text are never serialized into analytics events.
- **Prompt content is hashed, not stored, in the decisions log.** Each row carries a `promptHash` (SHA-256 of the system + user message) so you can prove a given decision came from a given prompt without storing the prompt itself in the log.
- **The agent transcript is never duplicated.** Inbox Pilot reads from the existing `conversation_messages` table on each call; it does not maintain a separate copy of your session content. Deleting a session removes the source data Inbox Pilot was reasoning about.
- **Only the latest agent message and the operator message immediately preceding it ever leave your machine.** The classifier call doesn't include the rest of the transcript — earlier turns, sidechain (sub-agent) rows, and system bookkeeping never reach the LLM provider. A long session with hundreds of messages still sends at most a single operator+agent pair (capped at 128k chars combined). This shrinks the prompt-injection surface area and reduces the volume of session content that a model provider sees per classification.
- **Calls run on Qwen3-32B via OpenRouter by default**, with a one-shot fallback to `anthropic/claude-haiku-4.5` (also via OpenRouter) when the primary returns malformed JSON or a schema-invalid object. The primary is ~$0.00007/call (roughly 1/300th the cost of Haiku). The fallback path is the safety net for model drift / new edge cases; both decisions still write to the decisions log with the model that actually fired. The provider+model defaults are overridable via `inboxPilotEvaluatorProvider` / `inboxPilotEvaluatorModel` (and `inboxPilotEvaluatorFallbackModel`) on the Inbox Pilot tile in the sidebar, but **never route the classifier to Anthropic-direct** — see "Evaluator routing" below.

### Security model

`RESPOND` is the only Inbox Pilot outcome that puts text into a running agent session, so it carries a tighter security framing than the four read-only outcomes:

- **The per-session prompt you author is treated as trusted user-authored input — but it's framed to the classifier as a _modifier_, not a replacement.** The global rule still applies on top of it; the session rule narrows or specializes the global rule for one session. The classifier sees `global rule → session rule → transcript` in that order, and the transcript is explicitly marked as untrusted content. A session rule that says _"answer 'yes' to anything"_ doesn't override the global rule's safety framing — it tells the classifier to lean toward `RESPOND` only when the agent's question is one the global rule would already consider safe to auto-answer.
- **The agent transcript is never trusted.** Anything in the session conversation — including text the agent generates that looks like an instruction to Inbox Pilot — is reasoned over but not obeyed. The classifier's only authoritative inputs are the global rule and the per-session rule; the transcript exists to be classified, not to instruct.
- **Confused-deputy mitigations are layered.** The five gates that make a runaway RESPOND loop hard to trigger are: (1) global Auto-respond off by default; (2) per-session opt-in required (checkbox off by default for every session); (3) the loop guard caps applied RESPONDs at 3 per rolling 5-minute window per session; (4) the 60-second cooldown blocks back-to-back RESPONDs in the same session; (5) the session-busy gate refuses to type while the CLI is mid-turn. Each gate fires independently — defense in depth, not a single chokepoint.
- **Every RESPOND is auditable.** The message renders with a visible "Sent by Inbox Pilot" attribution badge so a sent reply can never be mistaken for one you typed; the decisions log row carries the prompt hash, the response text, the model used, the cost, and the decision id linked from the badge tooltip.

### Limits and known gaps

Inbox Pilot is intentionally minimal. The following are explicitly **not** in scope and may show up in later versions:

- **Autonomous send is narrow and rate-limited.** The `RESPOND` and `SCHEDULE_RESPOND` outcomes are deliberately scoped: the classifier returns a single reply text (capped at 2,000 chars) for one agent turn, with a 3-per-5-min loop guard (shared across both action types), 60-second cooldown, global Auto-respond off by default, and per-session opt-in. There is no general "send `continue`" or follow-up-loop capability, and the classifier cannot chain multiple replies across turns from a single decision. `SCHEDULE_RESPOND` adds a hard 24-hour cap on how far ahead a queued reply can land — anything further out is rejected at the dispatcher. For general autonomous-send behavior (scheduled prompts, multi-turn workflows, condition-driven follow-ups), see the **Automations** feature (a separate system).
- **The daily cost cap covers Inbox Pilot's own Haiku spend only.** The default $1/day cap counts the Haiku classifier calls Inbox Pilot itself makes. It does **not** account for the downstream agent turns that `RESPOND` triggers — once Inbox Pilot types into a session, the agent's reply runs on the session's own (Sonnet / Opus) account and is billed there, not against the Inbox Pilot cap. If a runaway loop chained replies on an expensive model, the loop guard and cooldown above are what stop it, not the Haiku cap.
- **No project-level overrides.** The daily cap is global to your Omniscio install. You can't configure a different rule for project A vs. project B at the project level — per-session rules (above) are the only sub-global knob.
- **Snooze duration is fixed at 1 hour.** Inbox Pilot cannot pick a longer or shorter snooze; if it returns `SNOOZE`, the session is hidden for exactly 60 minutes.
- **No A/B preview.** The Preview button in the Inbox Pilot tile (and the kebab menu's per-session "Test on latest message…" item) runs your current rule against one real session in dry-run mode (writes a row with `wasDryRun = true`, no side effects) — there's no side-by-side comparison of two versions of a prompt. The tile-level Preview tests your **draft global rule in isolation** (no per-session layer); the kebab "Test on latest message…" tests **what the live router would actually do** for that session (global rule + that session's saved per-session rule, layered).

### Evaluator routing

The Inbox Pilot classifier call is the single hottest LLM call in Omniscio (one per agent end-of-turn on every opted-in session, all day, every day). Its routing defaults are tuned for **cheap + reliable JSON-mode output**, not for "use whichever provider the rest of Omniscio's AI features use".

**Defaults** (locked by [router-evaluator-routing-contract.test.ts](../../tests/unit/ai-manager/router-evaluator-routing-contract.test.ts)):

- **Primary**: `inboxPilotEvaluatorProvider = 'openrouter'`, `inboxPilotEvaluatorModel = 'qwen/qwen3-32b'`.
- **Fallback**: `inboxPilotEvaluatorFallbackModel = 'anthropic/claude-haiku-4.5'` (gated by `inboxPilotEvaluatorFallbackEnabled = true`), **hard-pinned to OpenRouter** (the user-facing field controls the model id only — the provider is never read from settings for the fallback path).
- **Token budget**: `max_tokens: 600`, `temperature: 0`.
- **Prompt**: every classifier system prompt ends with a fixed `CRITICAL_OUTPUT_RULES` block ("Do NOT think out loud", "begin your response with the opening brace `{`", "no markdown, no code fences"). The block is appended by `router-evaluator.ts` to both the global-rule and per-session-rule prompt branches; do not regress to only injecting it on one side.

**Why Anthropic-direct is forbidden for the classifier.** Anthropic's Messages API does not accept a top-level `response_format: { type: 'json_object' }` parameter. Omniscio's internal `llmProviderService.callAnthropic()` accepts the `responseFormat` field on its option bag but **silently drops it** when assembling the request — the call goes out without any JSON-mode enforcement, and the model is free to wrap its reply in prose, markdown fences, or a chain-of-thought preface. The classifier's downstream parser then trips on the first non-JSON byte and the inline pill reads **"Pilot couldn't evaluate this turn — response was not JSON."** OpenRouter, by contrast, honours `response_format` for both Qwen and `anthropic/claude-haiku-4.5` (it injects the appropriate guardrails on the upstream call), which is why both the primary and the fallback are pinned to it.

**`extractJson()` is mandatory before every `JSON.parse`.** Even with JSON mode and the CRITICAL OUTPUT RULES, occasional models still wrap output in ` ```json ` fences or prefix it with a single line of preamble. The [`extractJson`](../../src/main/services/email/email-prescreen.ts) helper strips fences, finds the outer `{…}` span, and tolerates a trailing comma; running it before `JSON.parse` turns most "almost-JSON" responses into a clean object. The primary call uses `extractJson` once; the fallback uses it again on the fallback response.

**Forensics on every parse / schema failure.** When either the primary or fallback call returns a string `extractJson` can't massage into a schema-valid object, the evaluator emits a `log.warn` with a 500-char truncation of the raw response. This is the single best signal for diagnosing a future Qwen drift or a new edge case — without it, all you see is the user-facing "couldn't evaluate this turn" pill with no model output to inspect. Do not silence these warnings; if you need to lower their volume, gate the log emission behind a debug flag, not delete the call.

**When the fallback fires**, the evaluator returns an `EvaluateRouterResult` with `fallbackTriggered: true` and the model on the decision row reflects the **fallback** model (so the inline pill shows `claude-haiku-4.5`, not `qwen3-32b`). The audit log thus tells you which model actually decided each row, regardless of which model was nominally configured as primary.

**Validation methodology** (when Qwen drifts or a new corpus edge case surfaces): the [ghost-eval postmortem](../../.claude/memory/postmortems/inbox-pilot-ghost-eval-json-parse-postmortem.md) captures the V0→V1 prompt iteration — start there, then run an in-app **shadow-mode soak** before flipping the default model. Note there is **no runnable smoke harness in the repo**: the original `tools/qwen-smoke/` harness + `results/SUMMARY.md` were never committed (code comments in `ai-manager-defaults.ts` and `router-evaluator.ts` still cite that path, but nothing exists behind it). The intended 18-fixture corpus was small and biased toward edge cases (questions with no clear answer, partial completions, ambiguous "let me know when you want to push" tails), but you cannot re-run it as-is — the shadow-mode soak is the practical signal.

## Related

What the Pilot is and what it actually decides — the outcomes, the snippet replies, and the cost controls — is the [Inbox Pilot](inbox-pilot.md) page. The alerts it may raise are described on [Inbox alerts](inbox-alerts.md).
