---
title: Plain Speak — gates, cost and privacy (part 3)
---

# Plain Speak — gates, cost and privacy (part 3)

## What it is

This is part 3 of the [Plain Speak](plain-speak.md) page — the controls and the boundaries: what it may look at, what it costs, what it records and what it sends.

## Where to find it

The switches are in **Settings → Plain Speak**; the decisions log is read from the settings surface, and the gate reports arrive on a build pipeline's own output.

## How it behaves

### Dev-pipeline gate reports

The one message class with its own rules — and the answer, in one place, to "does a gate report get a Plain Speak card?"

**Yes — the agent's own card, free. Never a paid or synthesized one.** A dev-pipeline gate report (a message leading with one of the six byte-exact gate headers — `## 🔍 🟢 Plan Ready`, `## ⚔️ 🟢 Red Team Complete`, `## 🔨 🟢 Build Complete`, `## 💎 🟢 Elegance Pass Complete`, `## 📝 🟢 Docs Complete`, `## 🚀 🟢 Ready to Merge`, including the `🟡` needs-you / `🔴` blocked variants) is detected in `evaluate()` at GATE 4.6 via `isGateReport`. That branch does two things:

- it **skips overlay-TEXT generation entirely** — no paid cascade, ever (a gate report is never billable), and the old client-side per-phase synthesis is RETIRED and must never be rebuilt;
- it then calls `applyGateReportInlineCard`, which takes the **agent's own** `[[AMC_PLAIN_SPEAK]]` card (written below the `[DEV-PIPELINE | …]` marker) and applies it through the shared `applyInlineCard`, persisting `overlayText`.

Persisting `overlayText` is what makes the pill render, so the result is: **the Plain Speak ⇄ Original pill appears**, Plain Speak shows the agent's card, and Original keeps the native gate report (its header + the pipeline stepper the renderer already draws). If the agent wrote no valid card, a `skipped` `'gate report: no inline card'` row is recorded and the bubble renders natively with no pill.

**Two deterministic touches on the way through.** The TLDR is guaranteed to lead with the gate's **phase emoji + status circle** (e.g. `🔨 🟢`), mirrored from the report's own header when the agent left them out — a two-glyph prepend, idempotent, never a rewrite (`gateHeaderGlyphs` + `ensureGatePhaseGlyphsInTldr`, invariant `a-dev-pipeline-gate-report`). And on a **green** gate a routine "approve to proceed?" widget is stripped from the card's `## Questions`, because the gate auto-advances and the toggle would be a dead control; a 🟡/🔴 gate keeps its real question.

**Timing.** Unlike the short-deferral and content skips, this branch runs **before** the needs-you alert is held back — a gate card IS a needs-you message, so its chime / badge / inbox row fire immediately and are never deferred. `OVERLAY_PENDING` never engages. Inbox Pilot still classifies the turn if enabled.

**History, so the dates stop confusing people.** A 2026-07-19 experiment machine-synthesized a card; the 2026-07-22 owner directive reversed it to a blanket "no overlay at all"; **2026-07-23 superseded that** with the current design — no paid/synthesized card, but the agent's own card applied free. Prose written against the 2026-07-22 state is stale.

Mechanics: `ai-manager/index.ts` (the GATE 4.6 branch) → `run-overlay.ts` (`applyGateReportInlineCard`); grammar in [gate-header-grammar.ts](../../src/shared/gate-header-grammar.ts). Contract: [plain-speak-overlay-gating-contract.md](../../.claude/memory/contracts/plain-speak-overlay-gating-contract.md) `evaluate-gate-a-dev-pipeline` / `dev-pipeline-gate-report-shows` / `a-dev-pipeline-gate-report`. Renderer-side header lifting: [gate-report-preamble-contract.md](../../.claude/memory/contracts/gate-report-preamble-contract.md).

### What context Plain Speak sees

The Plain Speak LLM call uses two payloads: a **system message** carrying the built-in rule prompt verbatim, and a **user message** carrying a structured XML envelope. Splitting them this way lets the model treat the rule (its own instructions) as trusted while treating everything inside the envelope (operator + agent text) as untrusted data.

The system message is the built-in `DEFAULT_OVERLAY_PROMPT`, sent unmodified. **As of 2026-05-30 the prompt is locked to this default** — the overlay execution path (`runOverlay` in [`ai-manager/index.ts`](../../src/main/services/ai-manager/index.ts)) passes `DEFAULT_OVERLAY_PROMPT` regardless of any value stored in `aiManager.overlay.prompt` (see _Prompt: locked to the built-in default_ below). The user message looks like this:

```
<context untrusted="true">
[operator] <text the operator said>
[agent] <text the agent said>
[operator] <text the operator said>
[agent] <text the agent said — THIS LAST AGENT TURN IS WHAT YOU SUMMARISE>
</context>
```

The transcript is a flat sequence of `[operator]` and `[agent]` lines in chronological order. Content rules:

- **Slice point.** If the session has been compacted (the agent emits a compaction summary and continues with reduced context), the transcript starts at the most recent compaction divider. Earlier history before the divider is dropped, mirroring what the agent itself can see.
- **System messages are dropped** from the transcript (other than the compaction divider itself, which passes through as the prefix line). The agent never sees them either.
- **Sub-agent (sidechain) messages are excluded** so a parallel forked exploration doesn't leak into the main thread's overlay.
- **JSON content blocks are unwrapped to plain text.** When an agent message is stored as a content-block array (`[{"type":"text","text":"..."}, ...]`), only the `text` blocks are extracted and concatenated; tool-call blocks and other types are dropped. Plaintext rows pass through unchanged.
- **Front-truncation to fit `MAX_CHARS`.** If the full transcript exceeds the prompt budget (`CONTEXT_TOKEN_CAP * 4` chars), the oldest lines are dropped from the front first. The latest agent message is always retained.
- The harness at [`tools/plain-speak/`](../../tools/plain-speak/) reproduces this envelope shape byte-for-byte; the lockstep test at [`tests/unit/services/overlay-prompt-lockstep.test.ts`](../../tests/unit/services/overlay-prompt-lockstep.test.ts) guards that `tools/plain-speak/prompts-source.md` and `DEFAULT_OVERLAY_PROMPT` in [`src/shared/ai-manager-defaults.ts`](../../src/shared/ai-manager-defaults.ts) stay byte-for-byte identical (the `.gitattributes` file pins the source file to LF so cross-platform contributors don't break the comparison).

The transcript is built by `buildOverlayTranscript()` in [`src/main/services/ai-manager/context-builder.ts`](../../src/main/services/ai-manager/context-builder.ts) and the envelope by `buildEnvelope()` in [`src/main/services/ai-manager/overlay-evaluator.ts`](../../src/main/services/ai-manager/overlay-evaluator.ts).

### Prompt: locked to the built-in default

**As of 2026-05-30 the Plain Speak prompt is locked — it is no longer user-editable.** The settings rule card dropped its prompt textbox, its _Reset to default_ link, and its _Preview_ button (only the title, description, and model line remain), and the runtime always rewrites with the built-in `DEFAULT_OVERLAY_PROMPT`, ignoring whatever is stored in `aiManager.overlay.prompt`. Anyone who had customised the prompt before this change falls back to the standard prompt. The `aiManager.overlay.prompt` field still exists in settings (it is shared with Inbox Pilot via `ruleConfigSchema`), but nothing reads it on the overlay path and nothing in the UI writes it. Inbox Pilot's prompt stays fully editable — this lock is **overlay-only**. The prompt-lock invariants (`plain-speak-prompt-editor` through `inbox-pilot-keeps-its-prompt`) in [plain-speak-settings-contract.md](../../.claude/memory/contracts/plain-speak-settings-contract.md) pin it.

The one-time boot migration below still runs, but it now only touches the stored `aiManager.overlay.prompt` value that the runtime ignores — it no longer affects what the model receives.

The Plain Speak prompt is stored in `aiManager.overlay.prompt` in your settings. When Omniscio starts up, a one-time migration (`migrateOverlayPrompt()` in [`src/main/services/ai-manager/migrate-overlay-prompt.ts`](../../src/main/services/ai-manager/migrate-overlay-prompt.ts)) checks the `overlayPromptMigratedV8` settings flag. If unset (or `false`), the migration runs once and flips the flag to `true`; once `true`, it short-circuits on every subsequent boot. **V8 is a carve-out, not a force-overwrite** — it preserves user customisations:

- If your stored prompt **equals the bundled v4.2 default verbatim** (i.e. you never customised it), the migration upgrades it to the v4.3 `DEFAULT_OVERLAY_PROMPT` and snapshots the pre-upgrade prompt to `aiManager.overlayPromptPreV8Backup` so you can restore the v4.2 text from Settings if v4.3 misbehaves.
- If your stored prompt is **anything else** (you customised it at any point), the migration leaves your prompt untouched and just flips the flag — no backup is taken since there's nothing to undo.

V8 ships the v4.3 prompt, which adds explicit `<question_shape>` rules to the v4.2 "Sectioned Markdown" prompt: every `## Questions` entry MUST have a numbered or bolded header AND ≥2 lettered options (e.g. `- A. ...`); orphan `- Or:` bullets and bare-question shapes are forbidden. The deterministic post-processor (V8 stage in the cascade — `normalizeQuestionShape`) backstops these rules at the pipeline boundary, so even users on customised prompts get the QW-shape fix. The carve-out approach is intentional for v4.3: it's a small refinement (3 paragraphs of new text), not a structural overhaul, and wiping a user's tuned prompt for a minor refinement would be hostile.

The V8 cap bump (50,000 → 75,000 chars on `ruleConfigSchema.prompt` and the textarea's `maxLength`) accommodates the new bundled v4.3 default — ~51,200 chars, slightly over the old 50k ceiling. The Zod cap and the renderer textarea cap MUST stay in lockstep so settings saves never silently fail; both regression tests in [`tests/unit/shared/ai-manager-types.test.ts`](../../tests/unit/shared/ai-manager-types.test.ts) and [`tests/unit/rule-editor.test.tsx`](../../tests/unit/rule-editor.test.tsx) pin them at 75,000.

Earlier rule-set bumps: V7 introduced the v4.2 "Sectioned Markdown" prompt with the model switch from Haiku to **Qwen 3 32B via OpenRouter** and a tightened brevity contract; V6 reoriented the `Latest` section to verb-led imperative form (banned "You asked me to..." wrappers via `<latest_rules>`); V5 reoriented the entire prompt to first-person voice (the persona is the agent itself rewriting its own message, with a `<voice_rules>` section forbidding "the agent" / "the assistant" anywhere in the output); V4 added the **View original** action plus the explicit-recommendation marking rule. The legacy V1–V7 flags are kept in the schema for forward-compat with old `config.json` files but are no longer read. **The prompt is no longer user-editable** (locked to `DEFAULT_OVERLAY_PROMPT` as of 2026-05-30 — see _Prompt: locked to the built-in default_ above), so there is no Settings textbox to tune it; the future-migration flags (`overlayPromptMigratedV9`, etc.) would only ever rewrite a stored value the runtime ignores.

### Cost and master-toggle controls

> **Inline-only note (2026-07-20):** the live inline card is **free**, so the daily cost cap and the API-key allowance below no longer apply under normal operation and were **removed from the Settings panel**. They now govern only the archived cascade (behind `AMC_ENABLE_OVERLAY_CASCADE`). The one live control is the master on/off toggle; the fields (`overlayDailyCapUsd`, `overlayAllowApiKey`) remain in the schema for back-compat.

Plain Speak has its **own** independent set of cost and on/off knobs — they are not shared with Inbox Pilot. There are three layers of control:

- **Daily cost cap (default $10.00) — cascade-only now.** Plain Speak has its own running daily total in USD, persisted in `aiManager.overlayDailyCapUsd`. When the cap is crossed, the (archived) cascade stops firing for the rest of the day; Inbox Pilot (which has its own separate cap, defaulting to $1.00) is unaffected. Capped fires are recorded in the decisions log with outcome `capped`. (The inline card never bills, so it never caps.)
- **Master on/off toggle (single flag — `overlay.enabled`).** A single **Enable Plain Speak** toggle sits at the very top of Settings → Plain Speak. It flips `aiManager.overlay.enabled` directly — there is no separate "pause" flag any more. When the toggle is **off**, Plain Speak does not fire for any session AND the rest of the panel — the read-only rule card and the recent-activity feed — is hidden entirely so the page collapses to just the toggle. Flip it back on and everything reappears in the same state it was left in. The execution gate at [`ai-manager/index.ts`](../../src/main/services/ai-manager/index.ts) reads `cfg.overlay.enabled && toggles.aiOverlayEnabled` — those are the only two booleans that gate the overlay branch. This is independent of Inbox Pilot's on/off state. **Default is ON for new installs** as of 2026-08-24 — fresh installs get `true` straight from `DEFAULT_AI_MANAGER_SETTINGS.overlay.enabled` (a seed-only flip; no migration touches existing installs, so they keep their state), and new users get a one-time "Meet Plain Speak" intro inbox card ([`plain-speak-intro-alert.ts`](../../src/main/services/app/plain-speak-intro-alert.ts)). It had been OFF since 2026-05-23 (the paid-cascade era); the live card is free now, so default-on carries no per-turn cost. Historical boot migrations still shape older installs:
  - The 2026-05-16 default-on migration (`migrateOverlayDefaultOn()`) was retired the same date and its source file deleted; the vestigial `aiManager.overlayDefaultOnMigrated` flag is preserved in the Zod schema so older `config.json` files still parse without errors.
  - The 2026-08-25 ship-on migration (`migrateOverlayShipOnV1()`) flips `overlay.enabled` false → true for existing installs and writes `overlayDefaultToPlainSpeak = false` on exactly the installs it turns on, so nobody has their message bodies replaced by a card they never asked for. Its sentinel `aiManager.overlayShipOnMigratedV1` is deliberately DISTINCT from `overlayDefaultOnMigrated` and `overlayForcedOffMigratedV2` — both are already true on every existing install, so a pass gated on either would short-circuit and never run. It is registered AFTER `overlay-pause-retirement-v1` (which only ever forces off) so the new default is the final word.
  - The 2026-05-25 force-off reset (`migrateOverlayForcedOffV2()`) was **retired on 2026-08-25** and its source file deleted, for the same reason in reverse: it force-offs any install whose `overlay.enabled` is true and whose sentinel is absent, which after the default-ON flip is every fresh install on first boot. Its `aiManager.overlayForcedOffMigratedV2` / `overlayResetBannerPending` flags stay in the Zod schema so existing `config.json` files still parse. A lint guard ([`plain-speak-default-on.test.ts`](../../tests/unit/lint/plain-speak-default-on.test.ts)) pins it retired.
  - On 2026-05-25, `migrateOverlayForcedOffV2()` reset `overlay.enabled` back to `false` for the installs the retired migration had switched on. Gated on a distinct sentinel (`overlayForcedOffMigratedV2`), runs once, never auto-enables. When it actually flips a live `true → false`, it also sets `aiManager.overlayResetBannerPending = true`, which drives the one-time reset banner described under _What you see_; the banner is cleared by the `ai-manager:dismiss-reset-banner` IPC (`handleAiManagerDismissResetBanner`), which writes `overlayResetBannerPending = false`.
  - On 2026-05-26, [`migrateOverlayPauseRetirementV1()`](../../src/main/services/config-store/migrations/migrate-overlay-pause-retirement-v1.ts) collapsed the previous two-flag model (`overlay.enabled` + `overlayPaused`) into a single master. Gated on `overlayPauseRetirementMigratedV1`. Honours historical pause intent on first run only: if `overlayPaused === true`, or if the legacy global `paused === true` and no per-feature override is set, the migration forces `overlay.enabled = false`. It NEVER auto-enables anyone — a user who re-enables Plain Speak after the migration keeps it on. The `overlayPaused` field stays in the Zod schema for back-compat but no runtime code reads it for the overlay branch. See [plain-speak-two-flag-drift postmortem](../../.claude/memory/postmortems/plain-speak-two-flag-drift-postmortem.md) for why the retirement was structural rather than another band-aid.
- **No per-session opt-out (removed 2026-05-14).** The per-session Plain Speak checkbox in each session's three-dot menu was removed to reduce menu clutter. Plain Speak now runs on **every** session whenever the master toggle is on and the daily cap allows it — there is no UI affordance to opt a single session out. The underlying store API still supports a per-session flag, but it is no longer surfaced or written to from the UI.

The **inline card is free**, so it runs on **any** account — OAuth or API-key — with no allowance to set (the `Allow on API-key accounts` toggle was removed from the panel under inline-only). The `overlayAllowApiKey` field remains in the schema and still gates the archived cascade on the paid path (behind `AMC_ENABLE_OVERLAY_CASCADE`), where an API-key account would be billed directly.

### Decisions log

Every Plain Speak invocation writes a row to the decisions log with `feature = 'overlay'` (the internal feature identifier — Plain Speak is the user-facing label). The row captures: the action taken, the LLM's reasoning, input character count, prompt hash, model used, tokens in/out, cost in USD, outcome (`in_flight` / `applied` / `cancelled` / `failed` / `skipped` / `capped`), error string if any, and whether it was a dry-run preview.

Retention runs on a daily prune:

- **Real (non-dry-run) decisions** are kept for **90 days**.
- **Dry-run preview decisions** are kept for **7 days** (they're noisy and short-lived).
- **Per-session cap** of **1000 most recent rows per session** — older rows beyond that count get pruned even within the 90-day window, so a single very-busy session can't fill the table.
- Orphaned `in_flight` rows older than 5 minutes are reaped at startup and marked `failed` with `error = "orphan reaped at startup"`.

You can see decisions in two places: the inline **Plain Speak** header on the affected agent bubble (click to flip between rewrite and original), and the full list in Settings → Plain Speak (most recent first). A **View in decisions log** button on each inline card jumps to that exact row.

### Privacy

A few hard rules about what Plain Speak does and does not store:

- **Your rule prompt text is NEVER logged in feature events.** The five tracked feature events (`ai_manager_router`, `ai_manager_overlay`, `ai_manager_session_enabled`, `ai_manager_preview`, `ai_manager_pause_all`) carry only a small allow-list of metadata: `outcome`, `action`, `model`, `was_dry_run`, `feature`, `paused`. The prompt itself, the agent transcript, and the LLM's reasoning text are never serialized into analytics events.
- **Prompt content is hashed, not stored, in the decisions log.** Each row carries a `promptHash` (SHA-256 of the system + user message) so you can prove a given decision came from a given prompt without storing the prompt itself in the log.
- **The agent transcript is never duplicated.** Plain Speak reads from the existing `conversation_messages` table on each call; it does not maintain a separate copy of your session content. Deleting a session removes the source data Plain Speak was reasoning about.
- **No extra provider sees your transcript.** Under inline-only, the agent that already ran your session writes the card in the same turn — Plain Speak makes **no separate model call**, so your transcript reaches no additional provider (Qwen, OpenRouter, extra Anthropic calls) the way the old cascade did. Only if the archived cascade is re-enabled (`AMC_ENABLE_OVERLAY_CASCADE`) would it call Qwen 3 32B / Haiku 4.5 / gpt-oss-20b via OpenRouter (or a single Sonnet / Qwen call) as before. Inbox Pilot is unrelated and still runs on Anthropic Haiku.

## Related

- [Plain Speak](plain-speak.md) — part 1.
- [Plain Speak part 2](plain-speak-part-2.md) — what you see while it works.
- [dev-pipeline.md](dev-pipeline.md) — the pipeline whose gate reports this page describes.

