---
title: Spoken Narration
---
# Spoken Narration

## What it is

> **In development — hidden by default.** Spoken Narration ships behind the
> `spoken-narration` unreleased-feature gate (setting `narrationEnabled`, default
> off; reveal via Settings → Lab or `AMC_SHOW_SPOKEN_NARRATION=1`). It is not visible
> to users until it is flipped to `shipped`.

**Spoken Narration** is the audio sibling of [Plain Speak](plain-speak.md). Where
Plain Speak shows a rewritten glance-card on screen, Spoken Narration plays a short,
friendly **spoken recap** of each completed agent reply — in the agent's own words
("here's what I did …"), not the raw message text and not the general read-aloud
summary. It reuses the existing TTS pipeline (Fish Audio / Grok / ElevenLabs).

**Every engine can produce it (2026-07-31).** The recap is written by the agent itself, inline, after a `[[OMNISCIO_NARRATION]]` marker at the very end of its reply (AMC strips it and routes it to the narration overlay / TTS). A **Claude** session is nudged to write it via the CLI `--append-system-prompt`; a **non-Claude** engine (Codex, Gemini, OpenCode, Cursor, Kimi Code, Hermes, Devin) — which has no such flag — gets the SAME nudge invisibly through the shared external first-message context plus a short per-turn reminder, the same generalized channel that delivers the Plain Speak card and the Final-Message marker to those engines (see [plain-speak.md](plain-speak.md) → "Every engine gets the directive"). **Codex (2026-08-11)** on a recent version instead carries the standing narration directive in Codex's own `developerInstructions` system-prompt slot (an older Codex falls back to first-message; its per-turn reminder still rides first-message). The narration extraction is engine-agnostic, so a Codex recap plays exactly like a Claude one.

A single **narration mode** (`narrationMode`, default `auto`) — set from the control in
the top toolbar (on a phone, in the session header) or in Settings — picks the behavior:

- **Auto-play** (`auto`) — when a reply finishes in the session you're viewing, its recap
  plays automatically. New replies only: switching away and back does not replay, and
  opening a session does not auto-play its last message.
- **Click to play** (`click`) — every agent message that carries a narration shows a
  ▶ play button to play that message's recap on demand; nothing auto-plays.
- **Don't narrate** (`off`) — the agent is not asked to write a recap at all (no
  `[[OMNISCIO_NARRATION]]` script), nothing is captured, and no play button appears. The mode
  control itself stays visible so you can switch back on.

The reply **text streams live and is never blocked or delayed** by narration. If
there's no script or TTS fails, the text is unaffected and audio simply doesn't play.

**On mobile** (the phone / web client): tap a message's play button to hear its recap
— the audio plays right on your phone. Auto-play of new replies is best-effort: because
mobile browsers block sound until you've interacted with the page, a new reply only
auto-plays **after your first tap** on a narration play button (that tap unlocks audio
for the session). Before that first tap, new replies stay silent until you tap one.

## Where to find it

A setting in **Settings** picks the behaviour, and the **Narration** tab appears on the same glance-card surface as Plain Speak. It is off by default and can be revealed early from **Settings → Lab**.

## How it behaves

### The mini-player

When nothing is playing, a narrated message shows just a **▶ play button**. Press it and the
button **expands in place into a little player** — pause/resume, a go-to-beginning ⏮, and a stop
✕ — right where the ▶ was. When the recap finishes (or you press stop) the player **collapses back
to just ▶**. So the extra controls are only there while you're actually listening, never cluttering
the header the rest of the time.

The play/pause control also stays in sync with **your keyboard's media keys and your computer's
media controls** — pause or resume the recap from outside AMC and the button flips to match, so it
never shows the wrong state.

If narration **can't play right now**, the ▶ is **greyed out** and hovering it tells you why —
**your daily narration limit is reached** (resets tomorrow), no narration **voice is set up**,
you're **offline**, or **read-aloud is turned off** — instead of a button that silently does nothing.

- **Go Back to Beginning** (⏮, inside the player) — restart the recap from the top. It stops
  whatever is playing first (so it never overlaps or picks up mid-sentence), then replays the same
  cached audio — no new synthesis, no extra cost.

### Reading it

- **Narration** view — the message's view switch (the pill that flips **Plain Speak** ⇄
  **Original**) gains a third **Narration** tab that shows the spoken script **as text**, so you
  can READ the recap instead of, or as well as, hearing it. The text loads the moment you open
  the tab; if it can't be found you get a short "not available to read" note rather than a
  spinner. The default view is never Narration — it's an explicit pick, one click away.

Both work on desktop and the phone / web client.

### How it works

1. **The agent writes the script inline** — no extra AI call. When the feature is on AND
   the mode isn't `off`, spawned sessions are nudged to end their reply with a
   `[[OMNISCIO_NARRATION]]` marker followed by a short spoken script. Because it rides the
   agent's normal turn, there is **$0 of additional generation cost**. Choosing "Don't
   narrate" removes the nudge, so the agent is never asked to narrate.
2. **AMC captures + hides it.** The marker and its tail are stripped so they never
   appear as visible message text (even a partial marker mid-stream). The script is
   stored on an `ai_manager_decisions` `feature='narration'` row (its own
   `narration_text` column), keyed by message id — the same split-storage pattern
   Plain Speak uses, so the immutable message row is untouched.
3. **The renderer learns a message has audio** via a presence-only signal (a boolean
   `hasNarration` on history load + a text-free availability push). The auto-play /
   availability path **never carries the script text** — audio is synthesized on demand in
   the main process and streamed as sound. The **one** time the text crosses to the renderer
   is when you deliberately open the **Narration read-view** (below): it lazily fetches THAT
   message's script over a read-only channel so you can read it — an explicit, per-message
   read of your own content, never a bulk load.
4. **Playback is metered** under its own daily cap (`narrationDailyCapUSD`, default $2,
   separate from the general read-aloud cap) with a distinct `tts-narration` cost
   source — no double-counting against the read-aloud path.

### Relationship to Plain Speak

Plain Speak (screen card) and Spoken Narration (audio) are independent: a narration
is captured whenever narration is active (`narrationEnabled` on and the mode isn't
`off`), even if the Plain Speak overlay is off.
When an agent emits both, the narration is the canonical **last** tail, and AMC splits
the two cleanly so neither swallows the other.

Engineering invariants (dual-marker parse order, presence-only delivery, the
availability-keyed auto-play race fix, the cost-cap discipline) are pinned by
[spoken-narration-contract.md](../../.claude/memory/contracts/spoken-narration-contract.md).

## For agents

### Settings

Settings → Voice Control → **Spoken Narration** (visible only while the feature is
revealed): the reveal toggle (`narrationEnabled`, in Lab), the **narration mode** picker
(`narrationMode` — Auto-play / Click to play / Don't narrate; the same control also sits
in the top toolbar, and on a phone in the session header), a **Playback speed** picker (`narrationPlaybackRate`, 0.5×–2×,
default 1× Normal) that sets how fast the recap plays — applied at play time, so it's
instant and free (no re-recording) and changes only the pace, not the voice — and an
advanced **Spoken-script prompt** override (`narrationPromptOverride`) to tune how the
recap is written.

**Per-session CLI opt-out:** a session spawned via the CLI (`/project/:name/new` or
`/agent/sessions`) can pass `suppressOverlayPrompts: true` to skip the narration *prompt*
entirely for that session — a headless spawn then never writes a `[[OMNISCIO_NARRATION]]`
script, so nothing is synthesized. Off by default; the same flag also suppresses the Plain
Speak card. See [cli-server-gating.md](../../.claude/memory/cli-server-gating.md).

## Related

Plain Speak is the on-screen half of the same idea, and the voice settings that pick a speaking voice govern both of them.
