---
title: Voice, TTS & Wake Word (hands-free Omniscio)
---

# Voice, TTS & Wake Word (hands-free Omniscio)

## What it is

> **Desktop app only.** Microphone capture, STT, TTS, and wake-word processing all run inside the Electron desktop app (the AudioWorklet captures audio in the renderer; the main process drives the STT/TTS providers and wake-word engines). There is no CLI route family and no mobile/web surface for any of these features. If you use Omniscio in a browser or on mobile, the voice stack is unavailable.

Omniscio ships a three-piece voice stack so you can run it without touching keyboard or mouse. **TTS** (Text-to-Speech) reads agent messages aloud — either on demand or auto-play as soon as an agent replies. Six TTS providers ship: **Fish Audio** (default), **Grok (xAI)**, **ElevenLabs**, **Speechify (Simba)**, **Pika (Pika Speech)**, and **Soniox (TTS v2)** — pick one in Settings → Voice Control → **TTS Provider**. **Voice dictation** captures what you say and delivers the transcript directly to the active session's composer — no intent parsing, no command matching. **Voice commands** turn spoken sentences into app actions: "open inbox", "answer yes", "create a session in the Api project to refactor the auth middleware", and 40+ more. STT (speech-to-text) runs on Deepgram Nova-3 (~250 ms round-trip), ElevenLabs Scribe v2 (~80 ms), or Groq Whisper v3 Turbo (batch: the full transcript arrives in one piece shortly after you stop speaking, with no live as-you-talk text). **Wake word** ("Hey Omniscio" or "Computer", or your choice) runs 100% on-device so the mic is always listening but nothing leaves your machine until you say the word — the **built-in free engine (openWakeWord)** (no key, no account, fully offline; reads model files from a per-user folder) is — as of 2026-07-22 — the ONLY user-selectable engine. The **Picovoice Porcupine** engine is retired (hidden from the picker, its proprietary models unbundled, its code retained and reversible). Pick in Settings → Voice → Wake Word → **Detection engine**.

## Where to find it

Settings → **Voice Control** holds the master voice stack: the TTS enable toggle, the TTS provider
picker with each provider's voice and API-key fields, dictation delivery, push-to-talk hold keys,
and the Omni Session Briefing controls. The wake word is configured under Settings → Voice → Wake
Word, and the speech recognition service is picked under Settings → Voice. Within the app itself,
the microphone button and the Alt+V command hotkey sit with the composer, and Settings → Voice
Control → **Test voice commands** exercises the command parser.

## How it behaves

### How to use it

1. **Turn on TTS.** Settings → **Voice Control** → flip `ttsEnabled`. Then pick a provider in **TTS Provider** (`ttsProvider`) — **Fish Audio** (default), **Grok (xAI)**, **ElevenLabs**, **Speechify (Simba)**, **Pika (Pika Speech)**, or **Soniox (TTS v2)**. Tune volume (`ttsVolume`); speed (`ttsSpeed`) applies to Fish only — none of Grok's, ElevenLabs', nor Speechify's API accepts a speed parameter, so the slider is ignored when Grok, ElevenLabs, or Speechify is selected. Flip `ttsAutoRead` to auto-play every agent reply, or leave it off to play on demand via the speaker button on any message. Long messages are chunked; anything past `ttsMaxChars` is truncated with an ellipsis.
   - **Fish Audio flow.** Pick a `ttsVoiceId` (Fish Audio voice), or add a custom one via **Add custom voice** → paste a fish.audio voice URL or model ID. **A personal Fish Audio key is NOT required** — custom voices, like the presets, synthesize and preview through the company relay, so a keyless user can add one, preview it, and hear it in narration. The built-in voices are **metered** (2026-09-02): each synthesis draws on your plan’s monthly voice time, then on your prepaid credit balance, and when both are empty playback stops with a top-up/upgrade prompt — never a message about an API key you never entered. Free plans include **no** voice time, so a free account runs entirely on credit. Adding your own key removes the metering entirely: you are billed by Fish Audio directly, with no limit from us. A personal key is optional: paste it in the **Fish Audio API Key** field and tap Save (Omniscio probes Fish to confirm it) only if you want the live voice-name lookup while adding a custom voice — without a key, Omniscio accepts a well-formed voice ID and confirms it the first time it plays.
   - **Grok flow.** Pick a `grokVoiceId` from the five preset voice tiles: **Ara**, **Eve**, **Leo**, **Rex**, **Sal** — **Eve** is the default. Tap a tile to select. Grok tiles ship no bundled preview clips and no per-tile Preview button — you'll first hear the voice when an agent reply auto-plays or when you tap the speaker button on a message. Paste your xAI key in the **xAI API Key** field and tap Save — Omniscio saves the key and then probes xAI to confirm it works (validation status appears next to the field). Same encrypted-storage UX as the Fish key (`safeStorage` + `enc:` prefix in `config.json`).
   - **ElevenLabs flow.** Pick an `elevenlabsVoiceId` (ElevenLabs Voice ID). Paste your ElevenLabs key in the **ElevenLabs API Key** field and tap Save — Omniscio saves the key and probes ElevenLabs to confirm it works. Like Grok, ElevenLabs accepts no speed parameter, so the speed slider is ignored when ElevenLabs is selected. Same encrypted-storage UX as the Fish key.
   - **Speechify (Simba) flow.** Paste your Speechify key and tap Save — Omniscio saves the key first, then validates it against `GET https://api.sws.speechify.com/v1/voices` (`validateApiKey('speechify', key)`). Once the key is saved, the **Speechify Voice** field becomes a **picker populated live from your own Speechify voice library** (fetched from that same `/v1/voices` endpoint via `TTS_LIST_SPEECHIFY_VOICES` — only voice metadata crosses IPC, never the key, which stays main-process only), and a **Test voice** button beside it auditions the selected voice through the normal preview path (`TTS_PREVIEW_VOICE`, cost logged as `tts-preview`). Pick a voice (`speechifyVoiceId`, default `alec`) or choose **Enter a custom voice ID…** to type an unlisted one; if the library can't be loaded (no key, offline, or empty) the field **falls back to a plain text box** so you're never worse off than before. Speechify is **user-key-only** (no company relay, like Grok / ElevenLabs), so it stays unavailable until you add your own key. The model is pinned to **simba-3.0** and, like ElevenLabs, Speechify accepts no speed parameter, so the speed slider is ignored when Speechify is selected. Same encrypted-storage UX as the Fish key (`safeStorage` + `enc:` prefix in `config.json`, main-process only). Note: on the free tier Speechify rate-limits to **1 request/second** (a 429 surfaces the friendly "rate-limited — wait a moment" notice).
   - **Pika (Pika Speech) flow.** Pick a `pikaVoiceId` from the eight preset voices (Documentary Narrator [default], News Anchor, Radio Alto, and more). Paste your Pika key in the **Pika API Key** field (from dev.pika.art/keys) and tap Save — Pika is **user-key-only** (no company relay, like Grok / ElevenLabs / Speechify), so it stays unavailable until you add your own key, and there is **no live key validation** (Pika has no validation endpoint, so the key is accepted as entered). Unlike the streaming providers, Pika Speech is an **async job API** — it POSTs the text (`script` + `voice_preset`) to `api.dev.pika.art`, polls until the clip is ready, then downloads the mp3, so a synth takes a few seconds rather than streaming instantly (best suited to narration clips). Like ElevenLabs / Speechify it ignores the speed slider. If Pika errors or runs out of credits, spoken narration auto-falls back to the free Fish voice (the shared out-of-credits safety net). Same encrypted-storage UX as the Fish key.
   - **Soniox (TTS v2) flow.** Pick a `sonioxVoiceId` from the preset voices (**Emma** [default], **Daniel**, **Grace**, and more) or **Enter a custom voice name** to type any of Soniox's ~70 named voices. Like Fish, Soniox is **company-funded** — a personal key is **optional**: presets and the **Test voice** button synthesize through the company relay (`sonioxTtsRelay`) with no key, so a keyless user can pick a voice and hear it (no key is ever required to preview, unlike Speechify/Pika). Like Fish, that keyless path is **metered** (2026-09-02) — plan voice time, then prepaid credit, then a top-up/upgrade prompt. Add your own key (from console.soniox.com) to bill your OWN Soniox account instead and drop the metering. Add your own key (from console.soniox.com) in the optional **Soniox API Key** field only if you want to bill your OWN Soniox account instead of the company one. Soniox v2 supports **emotion tags** ([excited], [calm], [chuckles], …), so spoken narration can direct delivery — the narration prompt teaches the agent the active voice's tags automatically. Like ElevenLabs/Speechify it ignores the speed slider at synth (Soniox speed is applied per request). Same encrypted-storage UX as the Fish key.
2. **Try a voice command.** Press **Alt+V** (the dedicated command hotkey), say something like "open the inbox" or "snooze this session for one hour". Release → Omniscio transcribes via your chosen STT provider, parses intent, and runs the matching action. On Mac the **Alt+V** chord works everywhere EXCEPT while the cursor is in a text field, so typing can never false-trigger it; from a text field use the wake word, the mic button, or rebind **Voice command** to a Cmd combo. You can also trigger commands via the **"Test voice commands"** button in Settings → Voice Control. (Note: the toolbar microphone button and the wake word trigger **dictation**, not commands — see step 6.)
3. **Dictate to a session or Team Chat.** Make sure a session or Team Chat channel is open and active, then either say your wake word ("Hey Omniscio" or "Computer") or click the **microphone button** in the input toolbar. Speak — Omniscio captures the transcript and delivers it to whichever composer is active (session or Team Chat). What happens next depends on the **Dictation delivery** setting:
   - **Send immediately** (default, `auto-send`) — the transcript is sent automatically.
   - **Review before sending** (`review-first`) — the transcript is placed in the message box so you can read and edit before sending.
     If neither a session nor Team Chat is open when you dictate, a toast prompts you to open one.
     > **Push-to-talk** (hold a key anywhere in the OS, release to stop) ships as two optional system-wide hold keys: **Push to Command** (hold, speak a command, release to run it, with the normal chime + result toast) and **Push to Dictate** (hold, speak, release to drop the transcript into the focused session input, honoring this same delivery setting). Configure them in **Voice Control** settings under **Push to talk**: both are off by default with no key pre-filled (settings `voicePttCommandEnabled`/`voicePttCommandKey` + `voicePttDictateEnabled`/`voicePttDictateKey`); the picker suggests layout-safe keys (Right Option / Right Command on US Mac layouts, Right Ctrl / Right Shift on AltGr layouts) and **Advanced (bind any key)** records anything, including bare modifiers. Presses shorter than 100 ms are ignored (accidental-brush guard), and PTT requires the master **Enable Voice Commands** toggle. macOS: the controls stay locked behind a **Grant Accessibility Permission** banner until the grant lands (the same grant FlowVoice uses). Right Windows is never suggested on Windows because releasing it would open the Start Menu. If the system key listener dies, the group shows a restart notice (hook supervisor states: ok / restarting / dead). The hold keys reuse the same engine, providers, and feedback surfaces as the in-app mic button.
4. **Chain commands.** Use "and then" / "then" / "after that" to run multiple in one breath: _"archive this session and then open the api project"_.
5. **Enable wake word for hands-free dictation.** Flip `wakeWordEnabled`, then pick a **Detection engine** (`wakeWordEngine`, default `openwakeword`): the **built-in free engine** (`openwakeword`, the default) needs no key and works out of the box with two built-in wake words (Hey Omniscio, Computer) whose model files install automatically on first use into `<userData>/wake-word-models/` (the Settings card shows the exact folder, per-keyword model status, and a "Check again" button); **Picovoice** needs a paid access key (paste it in the key field, which only shows for this engine). The default keyword is **Hey Omniscio**; change in `wakeWordKeywords`. Tune `wakeWordSensitivity` (0.0–1.0, default 0.7 — lower if you get false triggers, higher if it misses you; on the free engine this maps linearly to a score threshold, 0.7 → 0.60). `wakeWordFocusOnly` (default true) only listens when Omniscio is focused, which saves CPU and feels less creepy. When the wake word fires, dictation mode activates automatically — the transcript is delivered to the active session's composer per your **Dictation delivery** setting (see step 3 above).
6. **Tune the Dictation delivery setting.** In **Voice Control** settings, find **Dictation delivery** (`voiceDictationDelivery`). Choose **Send immediately** (auto-sends as soon as transcription finishes) or **Review before sending** (drops the transcript in the message box for you to edit first). The default is **Send immediately**.

> **Watch out — "pause" now pauses, not interrupts.** Saying a bare **"pause"** sets the active session's status to `paused` (mutually exclusive with snooze). It used to be a synonym for interrupt — that wiring was removed. To interrupt the agent mid-turn, say **"stop"**, **"interrupt"**, **"cancel"**, or **"halt"**.

> **New Tier 1 session controls.** **"snooze"** / **"snooze this session"** / **"snooze for 30 minutes"** snoozes the active session (default 60 min; "for N minutes" or "for N hours" is parsed — "snooze for 2 hours" → 120 min). **"approve"** / **"approved"** / **"yes"** / **"yeah"** / **"yep"** / **"yup"** / **"sure"** / **"go ahead"** / **"do it"** / **"sounds good"** / **"looks good"** / **"permission granted"** / **"grant permission"** all send a fixed affirmative text reply (literally `'Yes, go ahead.'`) to the active session — it **sends a chat message** and does **not** click any approval button or answer a tool-permission prompt; structured in-chat approval is a future capability.

> **New Tier 2–3 navigation and inbox commands.** Three commands were added in the 2026-05-22 voice-control-plane Plan 2. **Switch to session** — say **"switch to the \<name\> session"** or **"go to session \<name\>"** to jump directly to a named session using fuzzy name matching. **Approve inbox item** — say **"approve the request"**, **"accept the request"**, or **"approve the cron job"** to confirm the currently selected inbox approval (cron job, automation, recipe step, or similar). **Reject inbox item** — say **"reject the request"**, **"decline the cron job"**, or **"deny the request"** to dismiss the currently selected inbox approval. Both inbox commands require an approval card to be selected in the Inbox view; if nothing is selected the command is a no-op. Voice snooze of inbox items and reading inbox items aloud are not yet available.

## For agents

### How it works

TTS is orchestrated by [/src/main/services/tts/tts-service.ts](/src/main/services/tts/tts-service.ts), which delegates to one of six pluggable providers: [/src/main/services/tts-providers/fish-provider.ts](/src/main/services/tts-providers/fish-provider.ts) calls `POST https://api.fish.audio/v1/tts` (model `s2-pro`), [/src/main/services/tts-providers/grok-provider.ts](/src/main/services/tts-providers/grok-provider.ts) calls xAI's TTS endpoint using `getXaiApiKey()` and the configured `grokVoiceId`, [/src/main/services/tts-providers/elevenlabs-provider.ts](/src/main/services/tts-providers/elevenlabs-provider.ts) calls ElevenLabs using the configured `elevenlabsVoiceId`, and [/src/main/services/tts-providers/speechify-provider.ts](/src/main/services/tts-providers/speechify-provider.ts) POSTs the raw-audio streaming endpoint `https://api.sws.speechify.com/v1/audio/stream` (which returns mp3 bytes directly) with `getSpeechifyApiKey()`, the configured `speechifyVoiceId`, and the pinned model `simba-3.0` — and, like ElevenLabs, ignores `opts.speed` (Speechify's endpoint accepts no speed parameter), and [/src/main/services/tts-providers/pika-provider.ts](/src/main/services/tts-providers/pika-provider.ts) implements the `TtsProvider` interface DIRECTLY (not via `TtsProviderBase` — Pika Speech is an async job API, not a single streaming fetch): it POSTs `{ script, voice_preset }` to `https://api.dev.pika.art/v1/media/pika/pika-audio/pika-speech`, polls `GET /v1/media/jobs/{id}` until `completed`, then downloads the result mp3 (all three calls carry the `X-API-Key` header). It attaches the HTTP `status` to its errors so the spoken-narration Fish fallback + out-of-credits alert fire exactly as for the other providers, and it retries ONLY the pre-accept submit so a network blip can't double-bill an accepted job. Pika is user-key-only and has **no** `validateApiKey` entry (no validation endpoint), so its key field skips the post-save probe. [/src/main/services/tts-providers/soniox-provider.ts](/src/main/services/tts-providers/soniox-provider.ts) POSTs `https://tts-rt.soniox.com/tts` (model `tts-rt-v2` carried in the request BODY) and — like Fish — routes an empty key through the company relay (`sonioxTtsRelay`, `/t/soniox-tts`, METERED per caller ) at the relay — plan voice allowance, then prepaid credit, then `402`; `use_company_tts` is `minTier: 'free'` and the cached client check now only asks whether anyone is signed in to meter), so it is the second company-relayed voice; it has no `validateApiKey` entry (the optional key is accepted as entered). **Per-utterance provider lock**: each `speak*()` entry calls `resolveProvider()` once at the start, snapshots `{ provider, apiKey, voiceId }` from current settings, and threads that snapshot through every chunk of the utterance — flipping `ttsProvider` mid-playback does NOT swap voice; the in-flight audio finishes in its original voice and only the next utterance picks up the change. The Save button on an API key field saves the key first, then calls `validateApiKey(provider, key)` from [/src/main/services/api-key-validator.ts](/src/main/services/api-key-validator.ts) — keyed on the `ValidatableProvider` for that field (`'fish-audio'`, `'xai-grok'`, `'elevenlabs'`, `'speechify'`) — and surfaces the result next to the field; `validateApiKey` is a generic dispatcher (Speechify validates against `GET https://api.sws.speechify.com/v1/voices`). Long replies are chunked, and audio is cached on disk (SHA-256 keyed by `v2:${provider}:${voiceId}:${speed}:${text}`, LRU, 100 MB max, 7-day TTL) so repeated messages don't re-cost an API call. The cache key includes the provider name (and the cache version was bumped to `v2`) so a Grok request can never replay a stale Fish entry — switching providers cannot replay the wrong voice. Voice commands live in [/src/main/services/voice/voice-service.ts](/src/main/services/voice/voice-service.ts) — a 3-layer intent parser runs in order: (1) regex (<1 ms, handles "open X", "snooze N hours", etc.), (2) HuggingFace embeddings similarity against 41 canonical commands (~20 ms, catches phrasing variation), (3) Groq-Llama or Claude Haiku fallback (~300 ms, handles genuinely novel phrasings). The three Tier 1 actions added in the 2026-05-20 voice-control-plane plan (`session:snooze`, `session:pause`, `session:approve`) are reachable via the **L0 regex layer only** — embedding/LLM-layer training data for them lands in a later plan, so the canonical-command count stays at 41 until that refresh. STT is pluggable: Deepgram Nova-3 and ElevenLabs Scribe v2 stream live partials over WebSockets, while Groq Whisper v3 Turbo (`GroqWhisperSttProvider` in [/src/main/services/voice/voice-stt-providers.ts](/src/main/services/voice/voice-stt-providers.ts)) buffers the recording (10 MB cap) and uploads it as one multipart POST to `api.groq.com` when you stop, firing a single final transcript and never a partial; `SttProvider.disconnect()` is async so that final lands before voice-service resets its trigger mode. Groq STT reuses the same Groq API key as the voice-command LLM fallback. All three are wired via [/src/renderer/src/components/ui/VoiceInput.tsx](/src/renderer/src/components/ui/VoiceInput.tsx). Wake word is [/src/main/services/wake-word-service.ts](/src/main/services/wake-word-service.ts) — dual-pipeline: an AudioWorklet in the renderer sends 16 kHz PCM frames over IPC to the main process, which hands them to the selected engine behind a `WakeWordEngine` seam ([/src/main/services/wake-word/engine.ts](/src/main/services/wake-word/engine.ts)). `PorcupineEngine` wraps the Picovoice Porcupine v4.x native addon (512-sample frames, sync); `OpenWakeWordEngine` runs a three-model ONNX chain (melspectrogram → Google speech embedding → per-keyword head) on `onnxruntime-node` 1.21.0 — the exact runtime Omniscio already ships for semantic search — consuming 1280-sample (80 ms) chunks with a 2 s per-keyword refractory so one utterance fires once (~2.7 ms inference per chunk, ~3.4% of one core while actively listening). Switching engines disposes the old engine before the new one consumes a frame; ORT sessions are released on stop/shutdown. State machine pauses the wake engine during STT so you don't hear your own word re-trigger. CPU usage stays under 4%. 55 voice action types are declared in the `VoiceActionType` union in [/src/shared/types/voice.ts](/src/shared/types/voice.ts), spanning 9 categories (Session / Project / UI / Gmail / Settings / TTS / Notifications / Inbox / Special). The 12 Tier 1 actions (active session control) plus the 5 Tier 2–3 actions (navigation and inbox approvals — `navigate`, `session:select`, `project:select`, `inbox:approve`, `inbox:reject`) — 17 registry entries in total — are declared and routed through a typed risk-tier action registry at [/src/shared/voice-action-registry.ts](/src/shared/voice-action-registry.ts), with renderer handlers in [/src/renderer/src/features/voice/voice-dispatch-handlers.ts](/src/renderer/src/features/voice/voice-dispatch-handlers.ts) dispatched via [/src/renderer/src/features/voice/voice-dispatch.ts](/src/renderer/src/features/voice/voice-dispatch.ts). Each registry entry carries a `risk` classification (`read` / `low` / `mid` / `destructive`); the dispatcher refuses to execute any action tagged `destructive`. **No Tier 1 or Tier 2–3 action is destructive** — the gate is a guarded seam so higher tiers (global / project-wide) can opt into it without re-plumbing. Settings split across [/src/renderer/src/features/settings/sections/voice/VoiceSettings.tsx](/src/renderer/src/features/settings/sections/voice/VoiceSettings.tsx) (TTS + voice-command) and [/src/renderer/src/features/settings/sections/voice/WakeWordSettings.tsx](/src/renderer/src/features/settings/sections/voice/WakeWordSettings.tsx) (relocated into `sections/voice/` and split into per-control sub-components in the 2026-07-22 Tier-4 settings decomposition). Picovoice access key is encrypted in config.

## Related

- [notifications-and-silence.md](notifications-and-silence.md) — silence suppresses TTS auto-read the same way it suppresses sounds
- [use-quick-responses.md](use-quick-responses.md) — Alt+Z / Alt+1 / Alt+2 / Alt+3 are the keyboard parallels to voice replies
- [voice-l4-ask-agent.md](voice-l4-ask-agent.md) - Voice L4 managed ready-queue: ask several running agents at once, hear answers one at a time (go ahead / next / skip / what's ready?)

For the engines and the rest of the stack, see
[Voice, TTS & Wake Word (part 2)](voice-and-tts-part-2.md), which covers the wake-word engines,
the speech-to-text providers, the Omni Session Briefing, spoken output that follows the content's
language, third-party dictation and screen readers, and spoken thinking fillers.
