Omniscio documentation
Browse all documentation
  1. Getting Started13
  2. Sessions & Agents115
  3. Inbox & Notifications59
  4. Projects & Tasks95
  5. Automation & Scheduling75
  6. Knowledge & Memory26
  7. AI Features60
  8. Integrations100
  9. Plugins & Marketplace33
  10. Cloud & Teams56
  11. Settings & Customization58
  12. Account & Billing28
  13. Troubleshooting84
  14. CLI & API Reference22
  15. Legal & Policies4
  16. Uncategorised22

Spoken Narration

Spoken Narration is the audio sibling of Plain Speak. Where Plain Speak shows a rewritten glance-card on screen, Spoken Narration plays a short, friendly spoken recap of each completed agent reply — in the agent's own words ("here's what I did …"), not the raw message text and not the general read-aloud summary. It reuses the existing TTS pipeline (Fish Audio / Grok / ElevenLabs).

What it is

In development — hidden by default. Spoken Narration ships behind the spoken-narration unreleased-feature gate (setting narrationEnabled, default off; reveal via Settings → Lab or AMC_SHOW_SPOKEN_NARRATION=1). It is not visible to users until it is flipped to shipped.

Spoken Narration is the audio sibling of Plain Speak. Where Plain Speak shows a rewritten glance-card on screen, Spoken Narration plays a short, friendly spoken recap of each completed agent reply — in the agent's own words ("here's what I did …"), not the raw message text and not the general read-aloud summary. It reuses the existing TTS pipeline (Fish Audio / Grok / ElevenLabs).

Every engine can produce it (2026-07-31). The recap is written by the agent itself, inline, after a [[OMNISCIO_NARRATION]] marker at the very end of its reply (AMC strips it and routes it to the narration overlay / TTS). A Claude session is nudged to write it via the CLI --append-system-prompt; a non-Claude engine (Codex, Gemini, OpenCode, Cursor, Kimi Code, Hermes, Devin) — which has no such flag — gets the SAME nudge invisibly through the shared external first-message context plus a short per-turn reminder, the same generalized channel that delivers the Plain Speak card and the Final-Message marker to those engines (see plain-speak.md → "Every engine gets the directive"). Codex (2026-08-11) on a recent version instead carries the standing narration directive in Codex's own developerInstructions system-prompt slot (an older Codex falls back to first-message; its per-turn reminder still rides first-message). The narration extraction is engine-agnostic, so a Codex recap plays exactly like a Claude one.

A single narration mode (narrationMode, default auto) — set from the control in the top toolbar (on a phone, in the session header) or in Settings — picks the behavior:

  • Auto-play (auto) — when a reply finishes in the session you're viewing, its recap plays automatically. New replies only: switching away and back does not replay, and opening a session does not auto-play its last message.
  • Click to play (click) — every agent message that carries a narration shows a ▶ play button to play that message's recap on demand; nothing auto-plays.
  • Don't narrate (off) — the agent is not asked to write a recap at all (no [[OMNISCIO_NARRATION]] script), nothing is captured, and no play button appears. The mode control itself stays visible so you can switch back on.

The reply text streams live and is never blocked or delayed by narration. If there's no script or TTS fails, the text is unaffected and audio simply doesn't play.

On mobile (the phone / web client): tap a message's play button to hear its recap — the audio plays right on your phone. Auto-play of new replies is best-effort: because mobile browsers block sound until you've interacted with the page, a new reply only auto-plays after your first tap on a narration play button (that tap unlocks audio for the session). Before that first tap, new replies stay silent until you tap one.

Where to find it

A setting in Settings picks the behaviour, and the Narration tab appears on the same glance-card surface as Plain Speak. It is off by default and can be revealed early from Settings → Lab.

How it behaves

The mini-player

When nothing is playing, a narrated message shows just a ▶ play button. Press it and the button expands in place into a little player — pause/resume, a go-to-beginning ⏮, and a stop ✕ — right where the ▶ was. When the recap finishes (or you press stop) the player collapses back to just ▶. So the extra controls are only there while you're actually listening, never cluttering the header the rest of the time.

The play/pause control also stays in sync with your keyboard's media keys and your computer's media controls — pause or resume the recap from outside AMC and the button flips to match, so it never shows the wrong state.

If narration can't play right now, the ▶ is greyed out and hovering it tells you why — your daily narration limit is reached (resets tomorrow), no narration voice is set up, you're offline, or read-aloud is turned off — instead of a button that silently does nothing.

  • Go Back to Beginning (⏮, inside the player) — restart the recap from the top. It stops whatever is playing first (so it never overlaps or picks up mid-sentence), then replays the same cached audio — no new synthesis, no extra cost.

Reading it

  • Narration view — the message's view switch (the pill that flips Plain Speak ⇄ Original) gains a third Narration tab that shows the spoken script as text, so you can READ the recap instead of, or as well as, hearing it. The text loads the moment you open the tab; if it can't be found you get a short "not available to read" note rather than a spinner. The default view is never Narration — it's an explicit pick, one click away.

Both work on desktop and the phone / web client.

How it works

  1. The agent writes the script inline — no extra AI call. When the feature is on AND the mode isn't off, spawned sessions are nudged to end their reply with a [[OMNISCIO_NARRATION]] marker followed by a short spoken script. Because it rides the agent's normal turn, there is $0 of additional generation cost. Choosing "Don't narrate" removes the nudge, so the agent is never asked to narrate.
  2. AMC captures + hides it. The marker and its tail are stripped so they never appear as visible message text (even a partial marker mid-stream). The script is stored on an ai_manager_decisions feature='narration' row (its own narration_text column), keyed by message id — the same split-storage pattern Plain Speak uses, so the immutable message row is untouched.
  3. The renderer learns a message has audio via a presence-only signal (a boolean hasNarration on history load + a text-free availability push). The auto-play / availability path never carries the script text — audio is synthesized on demand in the main process and streamed as sound. The one time the text crosses to the renderer is when you deliberately open the Narration read-view (below): it lazily fetches THAT message's script over a read-only channel so you can read it — an explicit, per-message read of your own content, never a bulk load.
  4. Playback is metered under its own daily cap (narrationDailyCapUSD, default $2, separate from the general read-aloud cap) with a distinct tts-narration cost source — no double-counting against the read-aloud path.

Relationship to Plain Speak

Plain Speak (screen card) and Spoken Narration (audio) are independent: a narration is captured whenever narration is active (narrationEnabled on and the mode isn't off), even if the Plain Speak overlay is off. When an agent emits both, the narration is the canonical last tail, and AMC splits the two cleanly so neither swallows the other.

Engineering invariants (dual-marker parse order, presence-only delivery, the availability-keyed auto-play race fix, the cost-cap discipline) are pinned by spoken-narration-contract.md.

For agents

Settings

Settings → Voice Control → Spoken Narration (visible only while the feature is revealed): the reveal toggle (narrationEnabled, in Lab), the narration mode picker (narrationMode — Auto-play / Click to play / Don't narrate; the same control also sits in the top toolbar, and on a phone in the session header), a Playback speed picker (narrationPlaybackRate, 0.5×–2×, default 1× Normal) that sets how fast the recap plays — applied at play time, so it's instant and free (no re-recording) and changes only the pace, not the voice — and an advanced Spoken-script prompt override (narrationPromptOverride) to tune how the recap is written.

Per-session CLI opt-out: a session spawned via the CLI (/project/:name/new or /agent/sessions) can pass suppressOverlayPrompts: true to skip the narration prompt entirely for that session — a headless spawn then never writes a [[OMNISCIO_NARRATION]] script, so nothing is synthesized. Off by default; the same flag also suppresses the Plain Speak card. See cli-server-gating.md.

Related

Last verified 2026-09-23