Voice Report Back (TTS read-back of voice-command answers)
Voice Report Back speaks the prose answer to a voice command aloud through whichever TTS provider you have selected, so the press-talk-listen loop never needs the screen. Off by default, it speaks only LLM-layer replies, clips long answers behind a Speak more button, and stops at a daily cost cap. Also lists what deliberately stays silent and what this first slice cannot do yet.
What it is
After Omniscio successfully parses a voice command and produces an answer that's a real prose reply from the LLM — not a deterministic regex/embedding dispatch like "open inbox" or "snooze for one hour" — Voice Report Back speaks that answer back to you through whichever TTS provider you've selected (Fish Audio / Grok / ElevenLabs). It exists so the loop "press Alt+V → talk → hear the answer" works without ever looking at the screen.
The feature is off by default. Turning it on adds one extra step at the end of an LLM-dispatched voice command: the same text that would otherwise show up only in the floating voice-result card is also spoken aloud. Regex-layer commands ("archive this session", "open inbox") and embedding-layer commands are intentionally not spoken — they already have the existing toast + chime feedback path, and re-speaking "Session archived" on every clipped command would be noise.
This is Slice 1: one-way only. You hear the response; you can't interrupt it by speaking ("barge-in" is deferred to Slice 2, research lives at docs/superpowers/specs/2026-05-29-voice-barge-in-design.md). Only LLM-layer prose replies speak — tool confirmations ("Snoozed this session for 1 hour") stay silent. There is no in-flight cancel — the only kill switch is the toggle itself, which doesn't stop audio that's already playing.
Where to find it
Open Settings → Voice Control and expand the Voice Report Back card — the third of that section's three accordion cards, beside Voice input and Text-to-speech output. The card holds the Enable Voice Report Back switch, and Advanced settings inside it reveals the Max characters and Daily cost cap (USD) knobs. Outside Settings you trigger the feature by giving a voice command (Alt+V, the wake word, or the Test voice commands button), and the spoken answer plus its [Speak more] button appear on the floating voice-result card.
How it behaves
How to turn it on
Open Settings → Voice Control. Voice Control is a top-level Settings section with three accordion cards: Voice input, Text-to-speech output, and Voice Report Back. Expand the third card.
Flip Enable Voice Report Back (voiceReportBackEnabled) on. This is the only switch you need — the feature is now armed and will speak the next LLM-layer voice-command response.
Click Advanced settings to expose two tuning knobs:
- Max characters (
voiceReportBackMaxChars, default600, range50 – 10000, step50) — how many characters of a response are spoken before the rest is held back as a[Speak more]button. Long LLM answers get clipped here. 600 chars is roughly 90 seconds of speech in most voices. - Daily cost cap (USD) (
voiceReportBackDailyCapUSD, default1.0, range0 – 100, step0.1) — a hard spending ceiling per calendar day. When today's cumulative Voice Report Back TTS spend reaches the cap, the next eligible response is silently skipped (a[voice-report-back] daily cap reached …line is written to the main log) and speaking resumes tomorrow. The cap is checked pre-flight, before any TTS call goes out, so capacity overruns cost nothing.
Voice Report Back uses whichever TTS provider is selected in Settings → Voice Control → Text-to-speech output. There is no separate provider picker for read-back — Fish Audio / Grok / ElevenLabs all work, and the provider you hear is whatever TTS uses for any other speak action.
What you hear
Trigger a voice command via Alt+V (or the wake word, or the "Test voice commands" button), say something the regex and embedding layers can't pattern-match (genuinely novel phrasings — "summarize what's on my plate this afternoon", "what's blocking the auth refactor", "tell me what changed in this session"). The LLM layer answers in prose, and that prose is now spoken aloud.
If the answer fits within your Max characters budget, the whole thing speaks and the floating voice-result card looks the same as before.
If the answer is longer than the budget, Omniscio speaks the first N characters, appends " … continued", and the floating voice-result card grows a new [Speak more] button. Tapping it speaks the next N characters from where you left off. If the remainder is still over budget, the button reappears after the chunk; you can chain through a long answer one chunk at a time. The button uses the same TTS provider and the same daily cost cap as the original utterance.
The cost cap blocks BOTH the initial speak AND every subsequent [Speak more] click — once you've hit the daily ceiling, no further audio plays from this feature for the rest of the day.
What does NOT speak
Voice Report Back is deliberately narrow. Each rule below is enforced in the gate cascade described under "How it works":
- Off-by-default master toggle. Until you flip
voiceReportBackEnabled, nothing speaks. - Regex-layer dispatches. "archive this session", "snooze for 30 minutes", "open inbox", "pause" — every command the L0 regex parser matches deterministically. These already chime + toast; speaking them would be noise.
- Embedding-layer dispatches. Phrasing variations the L1 HuggingFace-similarity layer maps to a canonical command. Same reasoning.
- Fallback / unknown. Anything the parser can't classify.
- Tool confirmations. "Snoozed for 1 hour" / "Archived" / "Session ended" — these are action acknowledgements, not LLM prose. They have their own existing speech path via
speakResponse()in the voice service and are not routed through Voice Report Back. - Preview events. While the user is mid-utterance in buffered-mode dictation, the voice service emits "preview" intent results so the UI can show a hint of what's about to fire. These are explicitly excluded — Voice Report Back is wired only into the
executeAction()andexecuteChain()post-execute push points, not into the preview path. - Empty / whitespace text. A response whose trimmed length is zero is skipped.
- Daily cost cap reached. Pre-flight check; logs and skips.
Slice 1 limitations
This page is the present-tense truth for what Voice Report Back actually does today; everything below ships in later slices or is intentionally out of scope.
- No barge-in. You cannot interrupt an in-flight read-back by speaking. The user-facing impact is small for sub-600-char chunks (the default is ~90 seconds of audio) but real for long answers. Research notes for Slice 2 (push-to-talk barge-in vs VAD barge-in vs polite-only):
docs/superpowers/specs/2026-05-29-voice-barge-in-design.md. - No tool-confirmation speech. When you say "snooze this session", the resulting "Snoozed this session for 1 hour" goes through the existing
speakResponse()TTS path, not through Voice Report Back. Only LLM-layer prose replies speak. - No mid-speech cancel. Toggling Enable Voice Report Back OFF does not stop audio that's already playing — it only prevents the next utterance. There is no "shut up" command, hotkey, or button.
- No voice cloning / custom ElevenLabs voices. The feature uses whatever ElevenLabs Voice ID the user enters in Settings → Voice Control → Text-to-speech output. There's no per-feature override and no voice-cloning hook.
- Cost estimate is an estimate. The hardcoded per-char rates are approximations of public pricing. If a provider quietly raises rates, the cap will under-block until the rates here are refreshed. The cap remains the backstop.
For agents
How it works
When the voice service finishes parsing a command and is about to execute the matching action, it does two things in order: (1) emits an IPC.VOICE_INTENT_RESULT push to the renderer so the floating voice-result card can paint, and (2) calls voiceReportBackService.handleIntentResult({ source, text }) directly — no separate push-bus subscription, no IPC round-trip. Both executeAction() (single command) and each step of executeChain() (chained "X and then Y" commands) make this call. The voice service lives at src/main/services/voice/voice-service.ts; the Voice Report Back service lives at src/main/services/voice/voice-report-back-service.ts.
handleIntentResult then walks four gates in sequence; failing any one quietly returns:
voiceReportBackEnabled === true— master switch.event.source === 'llm'— the parse layer must bellm. Sourcesregex,embedding,fallback, andunknownare dropped here.- Non-empty text after
trim()— empty payloads return. queryTodaysCostBySource('voice-report-back-tts') < voiceReportBackDailyCapUSD— pre-flight cap check against theapi_cost_logtable. Today's cumulative spend across every Voice Report Back TTS call is summed; if it has reached the cap, a[voice-report-back] daily cap reached …line is logged viaelectron-logand the call returns.
If all four gates pass, the text is sliced. When text.length > voiceReportBackMaxChars, the spoken text becomes text.slice(0, maxChars).trim() + ' … continued' AND the service emits an IPC.VOICE_REPORT_BACK_TRUNCATED push with { fullText, spokenCharCount: maxChars }. The renderer's VoiceResultCard (inside src/renderer/src/components/ui/VoiceInput.tsx) subscribes to that channel via useIpcListener and renders the [Speak more] button. When text.length <= voiceReportBackMaxChars, the full text is spoken with no push.
The actual speak call is ttsService.speakText(spokenText), which routes to whichever provider the user picked in Settings → Voice Control → Text-to-speech output. The TTS service handles caching, chunking, and per-utterance provider lock; Voice Report Back doesn't see any of that. The speakText call is awaited inside a try / catch / finally — any TTS error is logged with log.warn('[voice-report-back] speak failed:', …) and swallowed (no toast, no inbox card, no crash). Whether the speak succeeds OR throws, the finally block always logs a row to api_cost_log via:
trackApiCostRaw(
'voice-report-back', // synthetic accountId
'voice-report-back-tts', // source label
provider, // model field (fish-audio | grok | elevenlabs)
charCount, // inputTokens slot (TTS has no real tokens)
0, // outputTokens
costUsd,
provider // routedVia
)
The per-character cost rates baked into the service are:
fish-audio: $0.000015/chargrok: $0.00003/charelevenlabs: $0.00025/char
These rates power BOTH the pre-flight cap check (via the log) AND the cost-log row written after each speak. The cost estimate is a flat per-char multiply — there's no token-counting because TTS APIs charge by characters, and the daily cap is the safety net for any rate drift versus the provider's actual billing.
[Speak more] flow
Clicking [Speak more] in the voice-result card invokes IPC.VOICE_REPORT_BACK_SPEAK_MORE with { fullText, offset }, where offset is the previously-spoken character count carried by the truncation push. The IPC handler (registered in src/main/ipc/voice-report-back-handlers.ts) calls voiceReportBackService.speakMore(...), which runs the same pipeline starting from the new offset: settings re-read, daily cap re-checked, fresh slice, fresh speak, fresh cost-log row. If the remainder is still over the max-chars budget, a follow-up VOICE_REPORT_BACK_TRUNCATED push fires with the updated spokenCharCount so the [Speak more] button can re-appear and you can chain through.
Boot wiring
registerVoiceReportBackHandlers() is called once at app boot from src/main/ipc/register.ts, which mounts the VOICE_REPORT_BACK_SPEAK_MORE IPC handler. The service itself has no boot step — it's a plain singleton imported by voice-service.ts, instantiated lazily on first use.
Spec
Design spec (rationale, alternatives considered, sign-off history): docs/superpowers/specs/2026-05-27-voice-report-back-design.md.
Related
- voice-and-tts.md — the underlying Voice Control + TTS stack Voice Report Back depends on (provider selection, voice IDs, chunking, caching all live there); also documents the sibling "Omni Session Briefing" feature, a different "speak this aloud" surface that summarizes a session rather than answering a command
Last verified 2026-09-23