Omniscio documentation
Browse all documentation
  1. Getting Started13
  2. Sessions & Agents115
  3. Inbox & Notifications59
  4. Projects & Tasks95
  5. Automation & Scheduling75
  6. Knowledge & Memory26
  7. AI Features60
  8. Integrations100
  9. Plugins & Marketplace33
  10. Cloud & Teams56
  11. Settings & Customization58
  12. Account & Billing28
  13. Troubleshooting84
  14. CLI & API Reference22
  15. Legal & Policies4
  16. Uncategorised22

Voice Report Back (TTS read-back of voice-command answers)

Voice Report Back speaks the prose answer to a voice command aloud through whichever TTS provider you have selected, so the press-talk-listen loop never needs the screen. Off by default, it speaks only LLM-layer replies, clips long answers behind a Speak more button, and stops at a daily cost cap. Also lists what deliberately stays silent and what this first slice cannot do yet.

What it is

After Omniscio successfully parses a voice command and produces an answer that's a real prose reply from the LLM — not a deterministic regex/embedding dispatch like "open inbox" or "snooze for one hour" — Voice Report Back speaks that answer back to you through whichever TTS provider you've selected (Fish Audio / Grok / ElevenLabs). It exists so the loop "press Alt+V → talk → hear the answer" works without ever looking at the screen.

The feature is off by default. Turning it on adds one extra step at the end of an LLM-dispatched voice command: the same text that would otherwise show up only in the floating voice-result card is also spoken aloud. Regex-layer commands ("archive this session", "open inbox") and embedding-layer commands are intentionally not spoken — they already have the existing toast + chime feedback path, and re-speaking "Session archived" on every clipped command would be noise.

This is Slice 1: one-way only. You hear the response; you can't interrupt it by speaking ("barge-in" is deferred to Slice 2, research lives at docs/superpowers/specs/2026-05-29-voice-barge-in-design.md). Only LLM-layer prose replies speak — tool confirmations ("Snoozed this session for 1 hour") stay silent. There is no in-flight cancel — the only kill switch is the toggle itself, which doesn't stop audio that's already playing.

Where to find it

Open Settings → Voice Control and expand the Voice Report Back card — the third of that section's three accordion cards, beside Voice input and Text-to-speech output. The card holds the Enable Voice Report Back switch, and Advanced settings inside it reveals the Max characters and Daily cost cap (USD) knobs. Outside Settings you trigger the feature by giving a voice command (Alt+V, the wake word, or the Test voice commands button), and the spoken answer plus its [Speak more] button appear on the floating voice-result card.

How it behaves

How to turn it on

Open Settings → Voice Control. Voice Control is a top-level Settings section with three accordion cards: Voice input, Text-to-speech output, and Voice Report Back. Expand the third card.

Flip Enable Voice Report Back (voiceReportBackEnabled) on. This is the only switch you need — the feature is now armed and will speak the next LLM-layer voice-command response.

Click Advanced settings to expose two tuning knobs:

  • Max characters (voiceReportBackMaxChars, default 600, range 50 – 10000, step 50) — how many characters of a response are spoken before the rest is held back as a [Speak more] button. Long LLM answers get clipped here. 600 chars is roughly 90 seconds of speech in most voices.
  • Daily cost cap (USD) (voiceReportBackDailyCapUSD, default 1.0, range 0 – 100, step 0.1) — a hard spending ceiling per calendar day. When today's cumulative Voice Report Back TTS spend reaches the cap, the next eligible response is silently skipped (a [voice-report-back] daily cap reached … line is written to the main log) and speaking resumes tomorrow. The cap is checked pre-flight, before any TTS call goes out, so capacity overruns cost nothing.

Voice Report Back uses whichever TTS provider is selected in Settings → Voice Control → Text-to-speech output. There is no separate provider picker for read-back — Fish Audio / Grok / ElevenLabs all work, and the provider you hear is whatever TTS uses for any other speak action.

What you hear

Trigger a voice command via Alt+V (or the wake word, or the "Test voice commands" button), say something the regex and embedding layers can't pattern-match (genuinely novel phrasings — "summarize what's on my plate this afternoon", "what's blocking the auth refactor", "tell me what changed in this session"). The LLM layer answers in prose, and that prose is now spoken aloud.

If the answer fits within your Max characters budget, the whole thing speaks and the floating voice-result card looks the same as before.

If the answer is longer than the budget, Omniscio speaks the first N characters, appends " … continued", and the floating voice-result card grows a new [Speak more] button. Tapping it speaks the next N characters from where you left off. If the remainder is still over budget, the button reappears after the chunk; you can chain through a long answer one chunk at a time. The button uses the same TTS provider and the same daily cost cap as the original utterance.

The cost cap blocks BOTH the initial speak AND every subsequent [Speak more] click — once you've hit the daily ceiling, no further audio plays from this feature for the rest of the day.

What does NOT speak

Voice Report Back is deliberately narrow. Each rule below is enforced in the gate cascade described under "How it works":

  • Off-by-default master toggle. Until you flip voiceReportBackEnabled, nothing speaks.
  • Regex-layer dispatches. "archive this session", "snooze for 30 minutes", "open inbox", "pause" — every command the L0 regex parser matches deterministically. These already chime + toast; speaking them would be noise.
  • Embedding-layer dispatches. Phrasing variations the L1 HuggingFace-similarity layer maps to a canonical command. Same reasoning.
  • Fallback / unknown. Anything the parser can't classify.
  • Tool confirmations. "Snoozed for 1 hour" / "Archived" / "Session ended" — these are action acknowledgements, not LLM prose. They have their own existing speech path via speakResponse() in the voice service and are not routed through Voice Report Back.
  • Preview events. While the user is mid-utterance in buffered-mode dictation, the voice service emits "preview" intent results so the UI can show a hint of what's about to fire. These are explicitly excluded — Voice Report Back is wired only into the executeAction() and executeChain() post-execute push points, not into the preview path.
  • Empty / whitespace text. A response whose trimmed length is zero is skipped.
  • Daily cost cap reached. Pre-flight check; logs and skips.

Slice 1 limitations

This page is the present-tense truth for what Voice Report Back actually does today; everything below ships in later slices or is intentionally out of scope.

  • No barge-in. You cannot interrupt an in-flight read-back by speaking. The user-facing impact is small for sub-600-char chunks (the default is ~90 seconds of audio) but real for long answers. Research notes for Slice 2 (push-to-talk barge-in vs VAD barge-in vs polite-only): docs/superpowers/specs/2026-05-29-voice-barge-in-design.md.
  • No tool-confirmation speech. When you say "snooze this session", the resulting "Snoozed this session for 1 hour" goes through the existing speakResponse() TTS path, not through Voice Report Back. Only LLM-layer prose replies speak.
  • No mid-speech cancel. Toggling Enable Voice Report Back OFF does not stop audio that's already playing — it only prevents the next utterance. There is no "shut up" command, hotkey, or button.
  • No voice cloning / custom ElevenLabs voices. The feature uses whatever ElevenLabs Voice ID the user enters in Settings → Voice Control → Text-to-speech output. There's no per-feature override and no voice-cloning hook.
  • Cost estimate is an estimate. The hardcoded per-char rates are approximations of public pricing. If a provider quietly raises rates, the cap will under-block until the rates here are refreshed. The cap remains the backstop.

For agents

How it works

When the voice service finishes parsing a command and is about to execute the matching action, it does two things in order: (1) emits an IPC.VOICE_INTENT_RESULT push to the renderer so the floating voice-result card can paint, and (2) calls voiceReportBackService.handleIntentResult({ source, text }) directly — no separate push-bus subscription, no IPC round-trip. Both executeAction() (single command) and each step of executeChain() (chained "X and then Y" commands) make this call. The voice service lives at src/main/services/voice/voice-service.ts; the Voice Report Back service lives at src/main/services/voice/voice-report-back-service.ts.

handleIntentResult then walks four gates in sequence; failing any one quietly returns:

  1. voiceReportBackEnabled === true — master switch.
  2. event.source === 'llm' — the parse layer must be llm. Sources regex, embedding, fallback, and unknown are dropped here.
  3. Non-empty text after trim() — empty payloads return.
  4. queryTodaysCostBySource('voice-report-back-tts') < voiceReportBackDailyCapUSD — pre-flight cap check against the api_cost_log table. Today's cumulative spend across every Voice Report Back TTS call is summed; if it has reached the cap, a [voice-report-back] daily cap reached … line is logged via electron-log and the call returns.

If all four gates pass, the text is sliced. When text.length > voiceReportBackMaxChars, the spoken text becomes text.slice(0, maxChars).trim() + ' … continued' AND the service emits an IPC.VOICE_REPORT_BACK_TRUNCATED push with { fullText, spokenCharCount: maxChars }. The renderer's VoiceResultCard (inside src/renderer/src/components/ui/VoiceInput.tsx) subscribes to that channel via useIpcListener and renders the [Speak more] button. When text.length <= voiceReportBackMaxChars, the full text is spoken with no push.

The actual speak call is ttsService.speakText(spokenText), which routes to whichever provider the user picked in Settings → Voice Control → Text-to-speech output. The TTS service handles caching, chunking, and per-utterance provider lock; Voice Report Back doesn't see any of that. The speakText call is awaited inside a try / catch / finally — any TTS error is logged with log.warn('[voice-report-back] speak failed:', …) and swallowed (no toast, no inbox card, no crash). Whether the speak succeeds OR throws, the finally block always logs a row to api_cost_log via:

trackApiCostRaw(
  'voice-report-back',          // synthetic accountId
  'voice-report-back-tts',      // source label
  provider,                     // model field (fish-audio | grok | elevenlabs)
  charCount,                    // inputTokens slot (TTS has no real tokens)
  0,                            // outputTokens
  costUsd,
  provider                      // routedVia
)

The per-character cost rates baked into the service are:

  • fish-audio: $0.000015/char
  • grok: $0.00003/char
  • elevenlabs: $0.00025/char

These rates power BOTH the pre-flight cap check (via the log) AND the cost-log row written after each speak. The cost estimate is a flat per-char multiply — there's no token-counting because TTS APIs charge by characters, and the daily cap is the safety net for any rate drift versus the provider's actual billing.

[Speak more] flow

Clicking [Speak more] in the voice-result card invokes IPC.VOICE_REPORT_BACK_SPEAK_MORE with { fullText, offset }, where offset is the previously-spoken character count carried by the truncation push. The IPC handler (registered in src/main/ipc/voice-report-back-handlers.ts) calls voiceReportBackService.speakMore(...), which runs the same pipeline starting from the new offset: settings re-read, daily cap re-checked, fresh slice, fresh speak, fresh cost-log row. If the remainder is still over the max-chars budget, a follow-up VOICE_REPORT_BACK_TRUNCATED push fires with the updated spokenCharCount so the [Speak more] button can re-appear and you can chain through.

Boot wiring

registerVoiceReportBackHandlers() is called once at app boot from src/main/ipc/register.ts, which mounts the VOICE_REPORT_BACK_SPEAK_MORE IPC handler. The service itself has no boot step — it's a plain singleton imported by voice-service.ts, instantiated lazily on first use.

Spec

Design spec (rationale, alternatives considered, sign-off history): docs/superpowers/specs/2026-05-27-voice-report-back-design.md.

Related

  • voice-and-tts.md — the underlying Voice Control + TTS stack Voice Report Back depends on (provider selection, voice IDs, chunking, caching all live there); also documents the sibling "Omni Session Briefing" feature, a different "speak this aloud" surface that summarizes a session rather than answering a command

Last verified 2026-09-23