Omniscio documentation
Browse all documentation
  1. Getting Started13
  2. Sessions & Agents115
  3. Inbox & Notifications59
  4. Projects & Tasks95
  5. Automation & Scheduling75
  6. Knowledge & Memory26
  7. AI Features60
  8. Integrations100
  9. Plugins & Marketplace33
  10. Cloud & Teams56
  11. Settings & Customization58
  12. Account & Billing28
  13. Troubleshooting84
  14. CLI & API Reference22
  15. Legal & Policies4
  16. Uncategorised22

Plain Speak — gates, cost and privacy (part 3)

Part 3 of the Plain Speak page: the reports a build pipeline gets, exactly what the rewrite is allowed to see, the built-in prompt it is locked to, the master switch and the spending controls behind it, the log of decisions it made, and what leaves your machine.

What it is

This is part 3 of the Plain Speak page — the controls and the boundaries: what it may look at, what it costs, what it records and what it sends.

Where to find it

The switches are in Settings → Plain Speak; the decisions log is read from the settings surface, and the gate reports arrive on a build pipeline's own output.

How it behaves

Dev-pipeline gate reports

The one message class with its own rules — and the answer, in one place, to "does a gate report get a Plain Speak card?"

Yes — the agent's own card, free. Never a paid or synthesized one. A dev-pipeline gate report (a message leading with one of the six byte-exact gate headers — ## 🔍 🟢 Plan Ready, ## ⚔️ 🟢 Red Team Complete, ## 🔨 🟢 Build Complete, ## 💎 🟢 Elegance Pass Complete, ## 📝 🟢 Docs Complete, ## 🚀 🟢 Ready to Merge, including the 🟡 needs-you / 🔴 blocked variants) is detected in evaluate() at GATE 4.6 via isGateReport. That branch does two things:

  • it skips overlay-TEXT generation entirely — no paid cascade, ever (a gate report is never billable), and the old client-side per-phase synthesis is RETIRED and must never be rebuilt;
  • it then calls applyGateReportInlineCard, which takes the agent's own [[AMC_PLAIN_SPEAK]] card (written below the [DEV-PIPELINE | …] marker) and applies it through the shared applyInlineCard, persisting overlayText.

Persisting overlayText is what makes the pill render, so the result is: the Plain Speak ⇄ Original pill appears, Plain Speak shows the agent's card, and Original keeps the native gate report (its header + the pipeline stepper the renderer already draws). If the agent wrote no valid card, a skipped 'gate report: no inline card' row is recorded and the bubble renders natively with no pill.

Two deterministic touches on the way through. The TLDR is guaranteed to lead with the gate's phase emoji + status circle (e.g. 🔨 🟢), mirrored from the report's own header when the agent left them out — a two-glyph prepend, idempotent, never a rewrite (gateHeaderGlyphs + ensureGatePhaseGlyphsInTldr, invariant a-dev-pipeline-gate-report). And on a green gate a routine "approve to proceed?" widget is stripped from the card's ## Questions, because the gate auto-advances and the toggle would be a dead control; a 🟡/🔴 gate keeps its real question.

Timing. Unlike the short-deferral and content skips, this branch runs before the needs-you alert is held back — a gate card IS a needs-you message, so its chime / badge / inbox row fire immediately and are never deferred. OVERLAY_PENDING never engages. Inbox Pilot still classifies the turn if enabled.

History, so the dates stop confusing people. A 2026-07-19 experiment machine-synthesized a card; the 2026-07-22 owner directive reversed it to a blanket "no overlay at all"; 2026-07-23 superseded that with the current design — no paid/synthesized card, but the agent's own card applied free. Prose written against the 2026-07-22 state is stale.

Mechanics: ai-manager/index.ts (the GATE 4.6 branch) → run-overlay.ts (applyGateReportInlineCard); grammar in gate-header-grammar.ts. Contract: plain-speak-overlay-gating-contract.md evaluate-gate-a-dev-pipeline / dev-pipeline-gate-report-shows / a-dev-pipeline-gate-report. Renderer-side header lifting: gate-report-preamble-contract.md.

What context Plain Speak sees

The Plain Speak LLM call uses two payloads: a system message carrying the built-in rule prompt verbatim, and a user message carrying a structured XML envelope. Splitting them this way lets the model treat the rule (its own instructions) as trusted while treating everything inside the envelope (operator + agent text) as untrusted data.

The system message is the built-in DEFAULT_OVERLAY_PROMPT, sent unmodified. As of 2026-05-30 the prompt is locked to this default — the overlay execution path (runOverlay in ai-manager/index.ts) passes DEFAULT_OVERLAY_PROMPT regardless of any value stored in aiManager.overlay.prompt (see Prompt: locked to the built-in default below). The user message looks like this:

<context untrusted="true">
[operator] <text the operator said>
[agent] <text the agent said>
[operator] <text the operator said>
[agent] <text the agent said — THIS LAST AGENT TURN IS WHAT YOU SUMMARISE>
</context>

The transcript is a flat sequence of [operator] and [agent] lines in chronological order. Content rules:

  • Slice point. If the session has been compacted (the agent emits a compaction summary and continues with reduced context), the transcript starts at the most recent compaction divider. Earlier history before the divider is dropped, mirroring what the agent itself can see.
  • System messages are dropped from the transcript (other than the compaction divider itself, which passes through as the prefix line). The agent never sees them either.
  • Sub-agent (sidechain) messages are excluded so a parallel forked exploration doesn't leak into the main thread's overlay.
  • JSON content blocks are unwrapped to plain text. When an agent message is stored as a content-block array ([{"type":"text","text":"..."}, ...]), only the text blocks are extracted and concatenated; tool-call blocks and other types are dropped. Plaintext rows pass through unchanged.
  • Front-truncation to fit MAX_CHARS. If the full transcript exceeds the prompt budget (CONTEXT_TOKEN_CAP * 4 chars), the oldest lines are dropped from the front first. The latest agent message is always retained.
  • The harness at tools/plain-speak/ reproduces this envelope shape byte-for-byte; the lockstep test at tests/unit/services/overlay-prompt-lockstep.test.ts guards that tools/plain-speak/prompts-source.md and DEFAULT_OVERLAY_PROMPT in src/shared/ai-manager-defaults.ts stay byte-for-byte identical (the .gitattributes file pins the source file to LF so cross-platform contributors don't break the comparison).

The transcript is built by buildOverlayTranscript() in src/main/services/ai-manager/context-builder.ts and the envelope by buildEnvelope() in src/main/services/ai-manager/overlay-evaluator.ts.

Prompt: locked to the built-in default

As of 2026-05-30 the Plain Speak prompt is locked — it is no longer user-editable. The settings rule card dropped its prompt textbox, its Reset to default link, and its Preview button (only the title, description, and model line remain), and the runtime always rewrites with the built-in DEFAULT_OVERLAY_PROMPT, ignoring whatever is stored in aiManager.overlay.prompt. Anyone who had customised the prompt before this change falls back to the standard prompt. The aiManager.overlay.prompt field still exists in settings (it is shared with Inbox Pilot via ruleConfigSchema), but nothing reads it on the overlay path and nothing in the UI writes it. Inbox Pilot's prompt stays fully editable — this lock is overlay-only. The prompt-lock invariants (plain-speak-prompt-editor through inbox-pilot-keeps-its-prompt) in plain-speak-settings-contract.md pin it.

The one-time boot migration below still runs, but it now only touches the stored aiManager.overlay.prompt value that the runtime ignores — it no longer affects what the model receives.

The Plain Speak prompt is stored in aiManager.overlay.prompt in your settings. When Omniscio starts up, a one-time migration (migrateOverlayPrompt() in src/main/services/ai-manager/migrate-overlay-prompt.ts) checks the overlayPromptMigratedV8 settings flag. If unset (or false), the migration runs once and flips the flag to true; once true, it short-circuits on every subsequent boot. V8 is a carve-out, not a force-overwrite — it preserves user customisations:

  • If your stored prompt equals the bundled v4.2 default verbatim (i.e. you never customised it), the migration upgrades it to the v4.3 DEFAULT_OVERLAY_PROMPT and snapshots the pre-upgrade prompt to aiManager.overlayPromptPreV8Backup so you can restore the v4.2 text from Settings if v4.3 misbehaves.
  • If your stored prompt is anything else (you customised it at any point), the migration leaves your prompt untouched and just flips the flag — no backup is taken since there's nothing to undo.

V8 ships the v4.3 prompt, which adds explicit <question_shape> rules to the v4.2 "Sectioned Markdown" prompt: every ## Questions entry MUST have a numbered or bolded header AND ≥2 lettered options (e.g. - A. ...); orphan - Or: bullets and bare-question shapes are forbidden. The deterministic post-processor (V8 stage in the cascade — normalizeQuestionShape) backstops these rules at the pipeline boundary, so even users on customised prompts get the QW-shape fix. The carve-out approach is intentional for v4.3: it's a small refinement (3 paragraphs of new text), not a structural overhaul, and wiping a user's tuned prompt for a minor refinement would be hostile.

The V8 cap bump (50,000 → 75,000 chars on ruleConfigSchema.prompt and the textarea's maxLength) accommodates the new bundled v4.3 default — ~51,200 chars, slightly over the old 50k ceiling. The Zod cap and the renderer textarea cap MUST stay in lockstep so settings saves never silently fail; both regression tests in tests/unit/shared/ai-manager-types.test.ts and tests/unit/rule-editor.test.tsx pin them at 75,000.

Earlier rule-set bumps: V7 introduced the v4.2 "Sectioned Markdown" prompt with the model switch from Haiku to Qwen 3 32B via OpenRouter and a tightened brevity contract; V6 reoriented the Latest section to verb-led imperative form (banned "You asked me to..." wrappers via <latest_rules>); V5 reoriented the entire prompt to first-person voice (the persona is the agent itself rewriting its own message, with a <voice_rules> section forbidding "the agent" / "the assistant" anywhere in the output); V4 added the View original action plus the explicit-recommendation marking rule. The legacy V1–V7 flags are kept in the schema for forward-compat with old config.json files but are no longer read. The prompt is no longer user-editable (locked to DEFAULT_OVERLAY_PROMPT as of 2026-05-30 — see Prompt: locked to the built-in default above), so there is no Settings textbox to tune it; the future-migration flags (overlayPromptMigratedV9, etc.) would only ever rewrite a stored value the runtime ignores.

Cost and master-toggle controls

Inline-only note (2026-07-20): the live inline card is free, so the daily cost cap and the API-key allowance below no longer apply under normal operation and were removed from the Settings panel. They now govern only the archived cascade (behind AMC_ENABLE_OVERLAY_CASCADE). The one live control is the master on/off toggle; the fields (overlayDailyCapUsd, overlayAllowApiKey) remain in the schema for back-compat.

Plain Speak has its own independent set of cost and on/off knobs — they are not shared with Inbox Pilot. There are three layers of control:

  • Daily cost cap (default $3.00) — cascade-only now. Plain Speak has its own running daily total in USD, persisted in aiManager.overlayDailyCapUsd. When the cap is crossed, the (archived) cascade stops firing for the rest of the day; Inbox Pilot (which has its own separate cap, defaulting to $1.00) is unaffected. Capped fires are recorded in the decisions log with outcome capped. (The inline card never bills, so it never caps.)
  • Master on/off toggle (single flag — overlay.enabled). A single Enable Plain Speak toggle sits at the very top of Settings → Plain Speak. It flips aiManager.overlay.enabled directly — there is no separate "pause" flag any more. When the toggle is off, Plain Speak does not fire for any session AND the rest of the panel — the read-only rule card and the recent-activity feed — is hidden entirely so the page collapses to just the toggle. Flip it back on and everything reappears in the same state it was left in. The execution gate at ai-manager/index.ts reads cfg.overlay.enabled && toggles.aiOverlayEnabled — those are the only two booleans that gate the overlay branch. This is independent of Inbox Pilot's on/off state. Default is ON for new installs as of 2026-08-24 — fresh installs get true straight from DEFAULT_AI_MANAGER_SETTINGS.overlay.enabled (a seed-only flip; no migration touches existing installs, so they keep their state), and new users get a one-time "Meet Plain Speak" intro inbox card (plain-speak-intro-alert.ts). It had been OFF since 2026-05-23 (the paid-cascade era); the live card is free now, so default-on carries no per-turn cost. Historical boot migrations still shape older installs:
    • The 2026-05-16 default-on migration (migrateOverlayDefaultOn()) was retired the same date and its source file deleted; the vestigial aiManager.overlayDefaultOnMigrated flag is preserved in the Zod schema so older config.json files still parse without errors.
    • The 2026-08-25 ship-on migration (migrateOverlayShipOnV1()) flips overlay.enabled false → true for existing installs and writes overlayDefaultToPlainSpeak = false on exactly the installs it turns on, so nobody has their message bodies replaced by a card they never asked for. Its sentinel aiManager.overlayShipOnMigratedV1 is deliberately DISTINCT from overlayDefaultOnMigrated and overlayForcedOffMigratedV2 — both are already true on every existing install, so a pass gated on either would short-circuit and never run. It is registered AFTER overlay-pause-retirement-v1 (which only ever forces off) so the new default is the final word.
    • The 2026-05-25 force-off reset (migrateOverlayForcedOffV2()) was retired on 2026-08-25 and its source file deleted, for the same reason in reverse: it force-offs any install whose overlay.enabled is true and whose sentinel is absent, which after the default-ON flip is every fresh install on first boot. Its aiManager.overlayForcedOffMigratedV2 / overlayResetBannerPending flags stay in the Zod schema so existing config.json files still parse. A lint guard (plain-speak-default-on.test.ts) pins it retired.
    • On 2026-05-25, migrateOverlayForcedOffV2() reset overlay.enabled back to false for the installs the retired migration had switched on. Gated on a distinct sentinel (overlayForcedOffMigratedV2), runs once, never auto-enables. When it actually flips a live true → false, it also sets aiManager.overlayResetBannerPending = true, which drives the one-time reset banner described under What you see; the banner is cleared by the ai-manager:dismiss-reset-banner IPC (handleAiManagerDismissResetBanner), which writes overlayResetBannerPending = false.
    • On 2026-05-26, migrateOverlayPauseRetirementV1() collapsed the previous two-flag model (overlay.enabled + overlayPaused) into a single master. Gated on overlayPauseRetirementMigratedV1. Honours historical pause intent on first run only: if overlayPaused === true, or if the legacy global paused === true and no per-feature override is set, the migration forces overlay.enabled = false. It NEVER auto-enables anyone — a user who re-enables Plain Speak after the migration keeps it on. The overlayPaused field stays in the Zod schema for back-compat but no runtime code reads it for the overlay branch. See plain-speak-two-flag-drift postmortem for why the retirement was structural rather than another band-aid.
  • No per-session opt-out (removed 2026-05-14). The per-session Plain Speak checkbox in each session's three-dot menu was removed to reduce menu clutter. Plain Speak now runs on every session whenever the master toggle is on and the daily cap allows it — there is no UI affordance to opt a single session out. The underlying store API still supports a per-session flag, but it is no longer surfaced or written to from the UI.

The inline card is free, so it runs on any account — OAuth or API-key — with no allowance to set (the Allow on API-key accounts toggle was removed from the panel under inline-only). The overlayAllowApiKey field remains in the schema and still gates the archived cascade on the paid path (behind AMC_ENABLE_OVERLAY_CASCADE), where an API-key account would be billed directly.

Decisions log

Every Plain Speak invocation writes a row to the decisions log with feature = 'overlay' (the internal feature identifier — Plain Speak is the user-facing label). The row captures: the action taken, the LLM's reasoning, input character count, prompt hash, model used, tokens in/out, cost in USD, outcome (in_flight / applied / cancelled / failed / skipped / capped), error string if any, and whether it was a dry-run preview.

Retention runs on a daily prune:

  • Real (non-dry-run) decisions are kept for 90 days.
  • Dry-run preview decisions are kept for 7 days (they're noisy and short-lived).
  • Per-session cap of 1000 most recent rows per session — older rows beyond that count get pruned even within the 90-day window, so a single very-busy session can't fill the table.
  • Orphaned in_flight rows older than 5 minutes are reaped at startup and marked failed with error = "orphan reaped at startup".

You can see decisions in two places: the inline Plain Speak header on the affected agent bubble (click to flip between rewrite and original), and the full list in Settings → Plain Speak (most recent first). A View in decisions log button on each inline card jumps to that exact row.

Privacy

A few hard rules about what Plain Speak does and does not store:

  • Your rule prompt text is NEVER logged in feature events. The five tracked feature events (ai_manager_router, ai_manager_overlay, ai_manager_session_enabled, ai_manager_preview, ai_manager_pause_all) carry only a small allow-list of metadata: outcome, action, model, was_dry_run, feature, paused. The prompt itself, the agent transcript, and the LLM's reasoning text are never serialized into analytics events.
  • Prompt content is hashed, not stored, in the decisions log. Each row carries a promptHash (SHA-256 of the system + user message) so you can prove a given decision came from a given prompt without storing the prompt itself in the log.
  • The agent transcript is never duplicated. Plain Speak reads from the existing conversation_messages table on each call; it does not maintain a separate copy of your session content. Deleting a session removes the source data Plain Speak was reasoning about.
  • No extra provider sees your transcript. Under inline-only, the agent that already ran your session writes the card in the same turn — Plain Speak makes no separate model call, so your transcript reaches no additional provider (Qwen, OpenRouter, extra Anthropic calls) the way the old cascade did. Only if the archived cascade is re-enabled (AMC_ENABLE_OVERLAY_CASCADE) would it call Qwen 3 32B / Haiku 4.5 / gpt-oss-20b via OpenRouter (or a single Sonnet / Qwen call) as before. Inbox Pilot is unrelated and still runs on Anthropic Haiku.

Related

Last verified 2026-10-06