---
title: Voice, TTS & Wake Word (hands-free Omniscio) (part 2)
---

# Voice, TTS & Wake Word (hands-free Omniscio) (part 2)

## What it is

This is part 2 of the
[Voice, TTS & Wake Word (hands-free Omniscio)](voice-and-tts.md) page. It covers the engines
underneath the voice stack: the two wake-word engines and the retired Picovoice option, the eight
speech-to-text providers with their keys, costs and capabilities, the Omni Session Briefing,
spoken output that follows the content's language, third-party dictation and screen readers, and
spoken thinking fillers.

## Where to find it

The surfaces are the same ones the parent page names. Settings → Voice → Wake Word carries the
wake-word engine picker and its key field, Settings → Voice carries the **Speech recognition
service** picker that chooses among the speech-to-text providers, and Settings → **Voice Control**
holds the TTS and voice-command settings that the Omni Session Briefing controls sit under. Each
provider's API key is entered in the key field beside its own picker.

## How it behaves

### Wake word engines (Picovoice vs the built-in free engine)

> **Picovoice RETIRED (2026-07-22).** The Picovoice engine is now hidden from the picker behind the
> `picovoice-wake-word` unreleased-feature gate, and its proprietary models are no longer bundled
> (`build.files` excludes the `.ppn` keyword files + `porcupine_params.pv`). A persisted `picovoice`
> selection falls back to the free engine (`effectiveWakeWordEngine`), so **the built-in free engine
> below is the only user-selectable engine now.** The enum, `PorcupineEngine`, and the
> `@picovoice/porcupine-node` dependency are DELIBERATELY kept — reversible by flipping the gate to
> shipped/on and dropping the packaging exclusions. See wake-word-dual-engine-contract.md **I16**. The
> two-engine description below is the architecture (still true in code) and the behavior when re-enabled.

Two engines, one picker (Settings → Voice → Wake Word → **Detection engine**, `wakeWordEngine`):

- **Built-in free engine** (`openwakeword`, the default) — an in-repo TypeScript port of the openWakeWord streaming pipeline (Apache-2.0 spec) on `onnxruntime-node` 1.21.0. No key, no account, no online validation: works on a plane, and works out of the box. Omniscio bundles the model files under `resources/wake-word-models/` and installs them copy-if-absent into `<userData>/wake-word-models/` on first use, so there is nothing to download and nothing to train. The set is two shared preprocessing models (`melspectrogram.onnx`, `embedding_model.onnx`, Apache-2.0) plus one head per built-in keyword named `<keyword>.onnx` (`computer.onnx`, `hey-omniscio.onnx`). A user can still drop their own model into the folder; the install never overwrites it. Missing files are a friendly config state, not a crash: keywords without a head are skipped with a log, zero usable models surfaces "needs model files" copy + the Settings card shows per-keyword status with a help link.
- **Picovoice Porcupine** (`picovoice`) — needs a personal access key from console.picovoice.ai. The key is validated ONLINE at every engine start. **Picovoice ended its free tier on 2026-06-30**: free keys no longer validate (paid keys keep working), which is why the free built-in engine is the default. When a saved key is rejected at start, Omniscio classifies it (`WakeWordBadKeyError` via `mapWakeWordInitError`), the START failure envelope carries `data.reason='picovoice-bad-key'`, and the renderer latches (`picovoiceKeyDead` in the wake-word store — no duty-cycle retry storm) and shows a ONE-per-run prompt offering a one-click switch to the free engine. The switch never happens without the user's click. Saving a fresh key re-arms the engine. Existing installs that still carried the old picovoice default WITHOUT a saved key are moved to the free engine once by a one-shot config migration (`migrate-wake-word-engine-default`, sentinel `wakeWordEngineDefaultMigrated`); a saved key marks a deliberate setup and is never touched.

**Model licensing (why these models ship, and which never do):** the two bundled head models are self-trained with the HeyBuddy pipeline (`scripts/wake-word-training/`, Apache-2.0 code, CC-BY-4.0 data), so they are cleared for commercial use, and the two preprocessing models are Apache-2.0. That is why they can ship. openWakeWord's PRETRAINED models (incl. `hey_jarvis`) are CC-BY-NC-SA (non-commercial) and are never committed or bundled. Evaluation models can still be fetched dev-time via `scripts/dev/fetch-oww-eval-models.mjs` into a gitignored fixtures dir. The exact set of four committed files is pinned by SHA-256 in a guard test, so a non-commercial model cannot slip into the bundle. Self-training is also the only way to get the "Computer" and "Hey Omniscio" heads (no pretrained equivalents exist). See `resources/wake-word-models/NOTICE.md` for provenance.

Wake phrase vs bare name: the built-in free engine's `hey-omniscio` head is a self-trained "Hey Omniscio" PHRASE model (see `resources/wake-word-models/NOTICE.md`), so say the full phrase "Hey Omniscio" to wake it. The bare name "Omni" alone is not guaranteed to fire. The free engine is English-centered; Porcupine's custom words cover more languages. Porcupine is lighter (sub-1% CPU vs ~3.4% of one core while listening), immaterial at Omniscio's duty-cycled usage.

### Speech-to-text engines (cloud keys vs the free Built-in engine)

> **Status: three engines still hidden, one now on by default.** The Built-in (free) speech engine is registered as the `local-stt` unreleased feature (Lab toggle `localSttEnabled`, dev reveal `AMC_SHOW_LOCAL_STT=1`), **Grok (xAI)** is registered as the `grok-stt` unreleased feature (Lab toggle `grokSttEnabled`, dev reveal `AMC_SHOW_GROK_STT=1`), and **Google Gemini 3.5** is registered as the `gemini-stt` unreleased feature (Lab toggle `geminiSttEnabled`, dev reveal `AMC_SHOW_GEMINI_STT=1`). **Meta Muse Voice Transcribe** is registered as the `meta-stt` feature (Lab toggle `metaSttEnabled`, dev reveal `AMC_SHOW_META_STT=1`), but it now declares **`defaultOn: true`** — it is VISIBLE to a default user, and turning its Lab toggle off hides it again. It was held back while the model was in development; that reason expired when Meta shipped it as a priced product (launched 2026-09-01), and nothing is billed to Omniscio either way since the engine uses the user's own Meta key. It is `defaultOn` rather than `status: 'shipped'` precisely so the toggle stays live instead of becoming inert. The other three remain hidden until flipped, and keyless behavior is unchanged throughout.

Eight STT providers share one picker (Settings → Voice → **Speech recognition service**, `voiceSttProvider`). Each engine is described ONCE — label, capture format, billing rate, credential, meetings capability — as a row in `STT_REGISTRY` (`src/shared/stt/stt-registry.ts`); the picker labels, the cost tables, the paid-source list, the spend unit, the encryption manifest, and the meetings capability table all DERIVE from that row, so adding an engine cannot leave one of them behind:

- **Deepgram Nova-3** (default), **ElevenLabs Scribe v2**, **Groq Whisper v3 Turbo**: cloud services, each needs its own API key. Deepgram and ElevenLabs stream live partial text; Groq is batch (one final transcript after you stop).
- **Grok (xAI)** (`grok`, gated behind the `grok-stt` flag): xAI's batch speech-to-text (`POST api.x.ai/v1/stt`, model `grok-stt`, ~$0.10/hr — cheaper than Deepgram). Batch like Groq — the transcript arrives after you stop. **It reuses the existing xAI key** (the same key as Grok TTS, via `getXaiApiKey` and the `TTS_SET_XAI_API_KEY` channel + `xai-grok` validation), so there is no separate key to add. Provider file `voice/stt-providers/grok.ts` (extends `BatchSttBase`); response parsing is deliberately tolerant (`text`|`transcript`, `duration`|last-segment-`end`) because xAI's live response contract is unverified, and the streaming WebSocket endpoint is a deferred follow-up. Billed under `stt-grok`. NOTE: `grok` (xAI) is a DISTINCT id from `groq` (Groq Whisper) — the one-letter difference is real.
- **Google Gemini 3.5** (`gemini`, gated behind the `gemini-stt` flag): Google's Gemini 3.5 Transcribe Live — a genuinely STREAMING WebSocket provider (unlike the Grok/Groq batch path), so it emits live interim + final partial text like Deepgram/ElevenLabs. Connects to `generativelanguage.googleapis.com` (BidiGenerateContent), streams raw PCM16 audio (the same `pcm16` AudioWorklet capture path as the Built-in engine, NOT WebM), and **reuses the existing `geminiApiKey`** (the same key as the Gemini CLI engine, via `getGeminiApiKey`) — no separate Voice key field, so the picker points to Settings → Accounts → Gemini (a Google sign-in alone is not enough; the Live endpoint needs an API key). Provider file `voice/stt-providers/gemini.ts` (extends `StreamingSttBase`, mirrors ElevenLabs' raw-WS lifecycle: pre-connect audio buffer, 6 s WS-ping keepalive, finalize-before-close). SECURITY: the key rides in the connect URL, so the URL is never logged and every ws/provider error is key-redacted (`redactGeminiKey`). Billed under `stt-gemini` (~$0.009/min audio-in, computed from streamed PCM bytes ÷ 32000). Ships gated behind `gemini-stt` for a live-verification pass before an unflagged release. **In FlowVoice (OS-wide dictation) this provider now has a SECOND credential mode** — a user with no key of their own dictates through the company gateway relay instead, which holds the key server-side; the provider gains one optional `relay` field on `SttConnectConfig` (auth token in a HEADER, never the URL) and suppresses its local cost row for a relayed session because the gateway already billed it. The in-app mic here is unchanged and still key-only; see [flowvoice.md](flowvoice.md) and the FlowVoice contract §16.
- **Meta Muse Voice Transcribe** (`meta`, `defaultOn` in the `meta-stt` feature — visible by default, still switchable off in Settings → Lab): Meta's `muse-voice-transcribe-1.0` over a WebSocket (`wss://api.meta.ai/v1/asr/realtime`) — streaming, so it emits live partial text. Streams raw PCM16 (the same AudioWorklet capture path as Gemini and the Built-in engine, NOT WebM). Billed under `stt-meta` at $3.00 per 1,000 audio-minutes ($0.18/hr), computed from streamed PCM bytes ÷ 32000. It needs its **own** Meta Model API key (`metaApiKey`) — unlike Grok and Gemini there is no existing Meta credential in the app to reuse — pasted at Settings → Voice → **Meta Model API key**. Provider file `voice/stt-providers/meta.ts` (extends `StreamingSttBase`), with the wire protocol split into a pure, socket-free `meta-muse-protocol.ts` so its awkward parts are unit-testable without a key.
  - **Three protocol differences worth knowing**, each of which otherwise produces a silent "connects but transcribes nothing": (1) **auth is in the first frame**, not the `Authorization` header — the header is ignored on this endpoint, so the connection is only considered ready on the server's ACK, never on socket open; (2) **turns overlap** — a later turn can open before an earlier turn's `speechComplete`, so all state is keyed by `turnId` and `speechEnd` is a boundary that carries no transcript; (3) **idle audio closes the stream** and a WebSocket ping does not reset that timer, so this provider runs no keep-alive pinger unlike its siblings.
  - **Limitations, stated rather than papered over:** turn-level timestamps only — no word-level timings and no confidence scores. Nothing in Omniscio consumes either today, so this serves every existing consumer, but a future feature needing word timings cannot use it. It is also **the engine Meetings uses for speaker-labelled transcripts**: its `DIARIZATION` mode labels speakers itself on ONE mixed stream, which is what picking Meta as the meeting engine selects. Note the mode is mutually exclusive with push-to-talk — Meta fixes it for the life of the session, and warns DIARIZATION "is not tuned for low-latency, voice-command use", which is fine for a meeting and wrong for dictation. Its diarization covers **non-overlapping speech only**: people talking over each other are not attributed.
- **Meta Muse Voice Transcribe via OpenRouter** (`openrouter`): the SAME engine as `meta` above, reached through a different transport. OpenRouter carries `meta/muse-voice-transcribe-1.0` at the same $0.18/hr, but only over a **batch HTTP** route (`POST https://openrouter.ai/api/v1/audio/transcriptions`) — it exposes **no realtime/WebSocket endpoint** for it at all (probed, not assumed: `/api/v1/realtime` and `/api/v1/audio/realtime` both 404, while `/api/v1/audio/transcriptions` answers 401 unauthenticated). A WebSocket client cannot talk to a batch endpoint, so this is a batch provider by construction and behaves like Groq/Grok: no interim text, one transcript after you stop. **It exists to remove the Meta-account requirement** — it rides the OpenRouter key you may already have. `captureFormat: 'pcm16'` — and this is NOT the webm its sibling batch providers upload. **Muse accepts only a WAV at 16 kHz or 24 kHz** and rejects everything else; the live API says so in its own words: `"Meta transcription requires a 16000 Hz or 24000 Hz WAV sample rate (received 22050 Hz)"`. Omniscio's pcm16 capture path is already 16 kHz mono Int16, so the provider prepends a 44-byte WAV header and sends that — no transcoding, nothing for a decoder to do. **It also CANNOT diarize:** measured on two distinct voices, the response is one flat `text` string with no speaker labels, and `response_format: "verbose_json"` is refused outright (`"The selected model does not support response_format \"verbose_json\""`). It transcribes accurately and fast — ~20 s of speech in ~4.9 s — but it cannot tell you WHO spoke. Billed under `stt-openrouter`, computed from the response's own `usage.seconds` — a real measured duration, so a successful call never records $0. **WHICH KEY MATTERS:** the app has two OpenRouter getters and they are not interchangeable — `getOpenrouterApiKey()` returns the BUNDLED DEFAULT (AMC's own key, used by Plain Speak/Qwen routing) while `getOpenrouterProviderApiKey()` returns the **user's**. This provider uses the **user's**, declared `shared-vendor` → `userOpenrouterApiKey`; billing the bundled key for a user's dictation would spend company money on their audio. Provider file `voice/stt-providers/openrouter.ts` (extends `CloudWhisperBatchBase`), overriding `responseFormat` to plain `json` because OpenRouter documents a **400** for `verbose_json` when the routed provider returns no structured output.
  - **Why two rows for one engine:** they are two transports, not two models. Keeping them separate makes the picker honest about what you are choosing (live vs after-stop) and keeps an auth failure unambiguous — `meta` fails on a Meta key, `openrouter` on an OpenRouter key.
  - **Not offered for meetings**, like `meta`, though for a different reason: there is no realtime route through this transport, so a live meeting transcript is not achievable here at all.
- **Built-in (free)** (`local`): Whisper tiny.en running on your own computer through `@huggingface/transformers` + `onnxruntime-node`, the exact runtime Omniscio already ships for semantic search and wake word. No key, no account, and no audio leaves the machine. Batch like Groq: the transcript arrives in one piece after you stop speaking, in about half a second on a modern machine and one to a few seconds on a low-spec one. English only in v1. Good for commands and short notes; the cloud keys stay the upsell for speed, noise, and accents.

How the Built-in engine works:

- **One-time model download.** The first use downloads about 42 MB (Whisper tiny.en, q8) from HuggingFace into `<userData>/models/Xenova/whisper-tiny.en/<revision>/`, next to the semantic-search embedding model. The revision is pinned to an immutable commit and every file is verified against a committed SHA-256 manifest (`stt-local-model-manifest.ts`); a mismatch wipes the cache and blocks use (fail closed, never load unverified weights). The Settings row shows Installed / downloading progress / a retry, and the picker's key area shows a Download button instead of a key field.
- **Runs in its own separate process.** Inference lives in a dedicated Electron `utilityProcess` child (`stt-local-worker-utility-entry.ts`, bundled by `scripts/build-stt-worker-utility.js`, mirroring the isolated embedding engine via the shared `services/ml/ml-child-host` adapter) so the app never freezes while transcribing **and** a hard native onnx crash kills only that helper — transcription degrades to an error and the app stays up, instead of the crash taking the whole app down (CID 9bf615c2). The child spawns on first voice use (never at boot) and tears itself down after 90 seconds idle to give back the roughly half a GB of model memory; the next use respawns it in well under a second. (`AMC_DISABLE_ML_SERVICE=1` reverts it to the legacy in-process worker thread.)
- **Keyless auto-fallback.** While the feature is revealed, a voice start whose selected cloud provider has no key from anywhere lands on the Built-in engine instead of erroring, with a one-time toast pointing at the faster paid options. Every keyed path is untouched, an explicit Built-in pick never checks keys, and a configured cloud key failing at runtime keeps today's error path (never a silent accuracy downgrade).
- **PCM capture.** For the Built-in engine **and Gemini Live** the renderer streams raw 16 kHz PCM16 through the same AudioWorklet the wake word uses (shared module `src/renderer/src/lib/voice-audio-capture.ts`); the cloud container providers (Deepgram / ElevenLabs / Groq / Grok) keep the WebM/Opus MediaRecorder path. `VOICE_START` returns which one to use (`captureFormat`, decided once by `sttCaptureFormat` in `voice/stt-capture-format.ts`).
- **Zero-cost usage events.** Local transcriptions are recorded through the normal voice cost path at $0 (`stt-local` source), so usage statistics stay complete.

**Daily spend cap on paid STT (`sttDailyCapUSD`, default `$5`, range `0–100`).** Paid STT (Deepgram / ElevenLabs / Groq / Grok / Gemini) bills per minute of audio, so a runaway or long-idle socket could otherwise cost unbounded. `doStartListening()` runs a pre-connect check: when today's summed `stt-deepgram` + `stt-elevenlabs` + `stt-groq` + `stt-grok` + `stt-gemini` spend (via `queryTodaysCostBySource`) reaches the cap, a new listen refuses with a plain "spending cap reached" notice instead of opening a metered socket. `0` disables paid STT entirely, and the free Built-in (`local`) engine is always exempt. Mirrors the read-aloud `ttsDailyCapUSD` and the per-layer voice caps; the cap query fails OPEN, so a telemetry hiccup never blocks voice (F004). It also has a client-side idle deadline in every capture mode — including `single`, which previously had none, so an idle socket could stream indefinitely (F045). Settable via the settings API/import; there is no Settings slider yet (backend guardrail with a safe default). **Deepgram cost tracks streamed audio, not wall-clock (F128):** `DeepgramSttProvider.disconnect()` bills from the COUNT of audio chunks actually streamed (chunk count × the 250 ms WebM capture timeslice `WEBM_CAPTURE_TIMESLICE_MS`, via `deepgramBilledSecondsFromChunks` in `voice/stt-cost.ts`), not connect→disconnect wall-clock — so the 6 s KeepAlive that holds the socket open through silence no longer inflates the estimate or trips this cap early. **Company-side backstop (F050):** the same `STT_PAID_COST_SOURCES` aggregate also feeds a daily-spend budget alert on the spend monitor (`AMC_STT_DAILY_SPEND_ALERT_USD`, default `$10`; see [ai-spend-alerts.md](ai-spend-alerts.md)) — the per-user cap here does not roll up across a fleet keyless-billing the shared built-in key.

### Omni Session Briefing

Spoken, AI-generated summary of a single session's conversation — "catch-me-up" style. Lead-with-accomplishments, 75–120 words, read aloud via TTS.

**Turn it on.** Settings → **Voice Control** → **Omni Session Briefing** → flip **Enable Omni Briefing**. Pick model (**Haiku** = fast, **Sonnet** = more nuanced). Optional: flip **Auto-play on session open** to hear a briefing whenever you open a session (guarded against re-play on navigate-away-and-back within the same app lifetime).

**Three ways to trigger**:

1. **Speaker button** in any session header (only visible when Omni is enabled)
2. **Voice command** — say "omni briefing", "brief me", "catch me up", or "summarize this session"
3. **Global hotkey** (default **Ctrl+Shift+J**) — fires even when Omniscio is minimized or in the background. Configure under Settings → Voice Control → Omni Session Briefing → **Global hotkey**: flip the toggle off to stop the OS-level interception entirely, or click the X next to the key combination field to clear it. Re-binding accepts any modifier+key combo (Ctrl/Alt/Shift required).

**What it summarizes**: operator messages + final agent text only. Tool-use lines (anything prefixed with ▸) are stripped before the LLM sees the transcript, keeping the spoken output free of file paths and code snippets.

**Editing the voice/tone**: the super prompt lives as `DEFAULT_BRIEFING_PROMPT` in [/src/main/services/session/session-briefing-service.ts](/src/main/services/session/session-briefing-service.ts). Edit in source and restart. A non-empty `jarvisBriefingPromptOverride` in settings wins over the source default when set.

**Caching**: briefings cache per session by message count. Re-triggering after no new activity replays the cached text (no new LLM call, but TTS still plays). Any new agent or operator message invalidates the cache and the next trigger regenerates.

**Spoken in the conversation's language**: the briefing is generated in the language of the conversation, not always English (see "Spoken output follows the content's language" below). If you set your own `jarvisBriefingPromptOverride`, yours is used verbatim — the language directive rides only on the default prompt.

### Spoken output follows the content's language

Both spoken AI features — the **read-aloud summary** of a long agent reply (`ttsSummarizationMode`) and the **Omni briefing** — generate their spoken text in the language of the content, not always English. This is part of Omniscio's "AI answers in your language" behavior: the language is auto-detected from the content, per item, with nothing to configure.

They also **match the voice to that language when a matching-language voice exists**, else keep your configured voice (the owner-approved coupling). The rule is fail-safe: it never overrides a custom voice it can't classify, never switches on an uncertain detection, never churns between two same-language voices, and never throws. Because Fish `s2-pro`, ElevenLabs, and Grok are multilingual, the current voice _speaks_ the detected language regardless.

**In practice today the voice never changes** — every preset voice in Omniscio's catalog is English, so there is no non-English voice to switch to, and Omniscio keeps your voice and speaks the language with it. The voice-switch wiring is complete and activates automatically the day a non-English preset is added to the catalog (a preset must carry a `language` tag to participate).

The language directive is the same shared building block the text AI features use ([/src/shared/ai-output-language.ts](/src/shared/ai-output-language.ts)); the voice matcher is [/src/shared/tts-voice-language.ts](/src/shared/tts-voice-language.ts) (`selectVoiceForLanguage` + `voiceLanguageCandidates`). RT-F016 resolves the voice _after_ the synchronous filler/TTS pre-emption in `speakLastMessage`, so a barge-in ("stop") still interrupts instantly. No new paid API calls — detection is a fast local pass.

### External dictation & screen readers

Omniscio's chat textareas also work with third-party dictation apps (**Wispr Flow**, **Dragon NaturallySpeaking**, Windows Speech Recognition) and screen readers (**NVDA**, **JAWS**, **Narrator**) — none of these are Omniscio features, but they share a Windows-level dependency that can fail silently if not handled. They all rely on Windows UI Automation (UIA) to read text out of and inject text into other apps. Electron disables its accessibility tree by default for performance, and without that tree Omniscio's textareas are invisible to UIA — Wispr's transcript or its paste-last-transcript shortcut produces no text. Chromium's auto-detect normally turns the tree on when it sees assistive tech probe, but the auto-detect can race with dictation tools on cold start, leaving one specific user broken while everyone else "just works." Omniscio enables the tree unconditionally at startup via `app.setAccessibilitySupportEnabled(true)` in [/src/main/index.ts](/src/main/index.ts) so the tree is up before any UIA probe — always on, no setting toggle, on every machine.

#### Why dictation tools may not reach Omniscio — admin elevation (Windows)

The single most common reason dictation breaks in Omniscio on Windows is **admin / non-admin elevation mismatch**. Users instinctively right-click → **Run as administrator** when an app misbehaves; this is the wrong instinct for Omniscio. Omniscio does not need administrator privileges for any normal operation (its data lives in `%APPDATA%`, the CLI server binds to a localhost port, and CLI processes spawn at user level). Running Omniscio elevated will silently break every UIA-based dictation tool unless that tool is also running elevated — and most are not, by design.

**Why this happens.** Windows enforces UIPI (User Interface Privilege Isolation): a process at lower integrity (medium / "normal user") cannot read or write input to a higher-integrity (admin) process's windows. Wispr Flow's hotkey listener typically runs in a per-user background process at normal integrity. When you press its global hotkey while Omniscio is focused, Wispr's UIA probe asks Windows "is there a focused text input here?" — and Windows answers "no, you can't see across the integrity boundary." Wispr concludes there's no text target and silently does nothing. The same applies to UIA `SetValue` calls when Wispr tries to inject a transcript.

**Symptoms when this hits a user.**

- The dictation tool's global hotkey (e.g. **Ctrl+Win** for Wispr Flow) does **nothing** in Omniscio's chat box. No error, no prompt — completely silent.
- When you manually engage dictation through the tool's own toolbar / icon, transcription happens, but the tool either fails to paste back or shows a generic "Update Admin Settings to Paste" / "Allow accessibility access" message.
- Dictation tools' "paste last transcript" shortcut (e.g. Wispr Flow's **Alt+Shift+Z**) silently fails to paste into Omniscio.
- Manual **Ctrl+V** still works (raw keyboard input, not subject to UIPI). This false-positive is what tricks users into thinking the dictation tool is broken rather than the integration.

If only manual **Ctrl+V** works and every other paste / dictation path silently fails, the elevation mismatch is almost certainly the cause.

**How to verify.** Open Task Manager → Details tab → right-click any column header → **Select columns** → check **Elevated** → OK. Locate the Omniscio process (`Omniscio.exe`, or `electron.exe` if running dev) and the dictation tool's processes (e.g. all `Wispr Flow*.exe`). For dictation to work, **both must show the same value** in the Elevated column — typically both **No**.

**The fix.** Don't run Omniscio as administrator. If you previously launched it via right-click → Run as administrator (or pinned a shortcut that does), close it fully and launch it normally instead. The dictation tools at user-level can then talk to Omniscio at user-level via UIA.

#### Other things to check if dictation still doesn't reach Omniscio

- The tool's per-app allowlist should include "Omniscio."
- Windows Settings → Privacy → Accessibility should show the tool as permitted.
- Some dictation tools require their own background service to be running — check the tool's tray icon.

### Spoken thinking fillers (slow voice turns)

When the voice assistant runs a slow lookup (asking about a thread, searching, or
an ask-about-a-thread answer), Omniscio can speak a brief filler such as "One sec." or
"Let me check." so you are not left wondering whether it heard you - useful
eyes-free or while driving.

- Off by default. Fillers play only when the voice response style is set to
  "Conversational" (Settings -> Voice Control), the relevant voice layer (L2
  conversation or L3 ask-about-a-thread) is enabled, and text-to-speech is on.
  Low Power Mode suppresses them.
- A filler only plays when the lookup is genuinely slow (about 1.2 seconds), at
  most once per question, and the real answer always cuts the filler short - the
  two never talk over each other.
- Phrases are fixed and reused from the speech cache, so fillers add no new AI
  cost and count against the existing per-layer daily voice spending caps.

Implementation: `speakFiller` in [/src/main/services/tts/tts-service.ts](/src/main/services/tts/tts-service.ts)
plus one-shot timers at the two slow call-sites (the L2 read-tool loop and the L3
answer handler); invariants in `.claude/memory/contracts/voice-thinking-fillers-contract.md`.

## Related

This is part 2 of a two-part page. Part 1 is
[Voice, TTS & Wake Word (hands-free Omniscio)](voice-and-tts.md), which covers the three-piece
voice stack itself, how to turn TTS, dictation and voice commands on, and how the whole stack is
wired.

- [notifications-and-silence.md](notifications-and-silence.md) — silence suppresses TTS auto-read the same way it suppresses sounds
- [use-quick-responses.md](use-quick-responses.md) — Alt+Z / Alt+1 / Alt+2 / Alt+3 are the keyboard parallels to voice replies
- [voice-l4-ask-agent.md](voice-l4-ask-agent.md) - Voice L4 managed ready-queue: ask several running agents at once, hear answers one at a time (go ahead / next / skip / what's ready?)
