---
title: Reply by Voice (talk back to a spoken recap)
---

# Reply by Voice

## What it is

> **In development — hidden by default.** Reply by Voice ships behind the
> `narration-voice-reply` unreleased-feature gate (setting `narrationVoiceReplyEnabled`,
> default off; reveal via Settings → Lab). It also **requires Voice Input** (`voiceEnabled`)
> to capture. Not visible to users until it is flipped to `shipped`.

### What it is

**Reply by Voice** is the reply sibling of [Spoken Narration](spoken-narration.md). Where
narration plays the AI's spoken recap, Reply by Voice lets you **talk back**: a **mic button**
sits beside each message's narration play button, and tapping it opens a window where you
speak, watch a **live transcript** appear, **edit** it if the speech-to-text slipped, and
**send** it to that session as your reply — voice instead of typing, right at the recap moment.

It reuses the app's existing voice input end to end (mic capture, local Whisper or cloud
speech-to-text, and the normal send path) — there's no new voice engine.

**On mobile** (the phone / web client): the mic button and window work on a phone too, over
the same secure connection. Mobile mic capture is **best-effort** — it depends on the phone
browser allowing microphone access, and isn't guaranteed on every browser until tested on a
real device.

## Where to find it

The **mic button** sits beside each message's narration play button, on messages that carry a spoken recap. Its own controls are in **Settings → Voice & Speech → Reply by Voice**, visible only while the feature is revealed.

## How it behaves

### How it works

1. **Talk.** Opening the window takes over the microphone and starts dictation; your words
   appear live and each settled phrase is added to an editable box. Opening the window also
   **stops the recap audio** so the mic doesn't hear the AI talking over you.
2. **It cleanly owns the mic.** While the window is open it is the sole owner of the mic, so
   the app's always-listening voice system stands down — your spoken reply can never be
   double-sent or misrouted into the wrong session.
3. **Review or auto-send (your choice).** By default you **review and edit**, then tap Send.
   You can switch to **auto-send** in settings, which sends after you pause (with a short,
   cancelable countdown, and only for a long-enough phrase so a stray noise can't fire a paid
   reply).
4. **Send.** Your (edited) words go to that session exactly as if you'd typed them.

If Voice Input is off, the window shows a clear "turn on Voice Input" prompt instead of
failing silently.

### Settings

Settings → Voice & Speech → **Reply by Voice** (visible only while the feature is revealed):
the enable toggle (`narrationVoiceReplyEnabled`) and the send-behavior toggle
**Auto-send when you pause** (`narrationVoiceReplyAutoSend`, default off = review-first).

### Relationship to Spoken Narration

The mic button appears exactly where a narration play button does — on agent messages that
carry a spoken recap — so Reply by Voice only shows up where there's something to reply to.
Engineering invariants (the mic-ownership fence, the release drain, capture reuse, the
narration-echo stop, the send-behavior guards) are pinned by
[narration-voice-reply-contract.md](../../.claude/memory/contracts/narration-voice-reply-contract.md);
the shared voice-routing rules live in
[voice-capture-routing-contract.md](../../.claude/memory/contracts/voice-capture-routing-contract.md).

## Related

- [spoken-narration.md](spoken-narration.md) — playing the AI's spoken recap, which this feature replies to.
- [reply-to-start-nudge.md](reply-to-start-nudge.md) — the other way a reply starts something.
- [voice-and-tts.md](voice-and-tts.md) — the voice and speech surface this feature reuses to capture your words.

