---
title: FlowVoice (OS-wide dictation + history)
---

# FlowVoice (OS-wide dictation + history)

## What it is

FlowVoice is Omniscio's **OS-wide dictation** feature: hold a global hotkey, speak, release, and the cleaned transcript is typed at the OS cursor in **whatever app is focused** (browser, editor, Slack, Gmail), not just Omniscio's own message box. It is a **self-contained port** of the standalone FlowVoice app, living entirely under `src/main/services/flowvoice/` (main) and `src/renderer/src/features/flowvoice/` (renderer), with its own `flowvoice:*` IPC channels and its own JSON settings/history stores; it does not touch Omniscio's existing voice stack. See the contract at `.claude/memory/contracts/flowvoice.contract.md`. The feature is **off by default** and is not gated behind an unreleased-feature flag: the settings card is visible to every user.

## Where to find it

**Settings → Voice Control**, under the **Dictation (FlowVoice)** heading. Without leaving a session you can also reach the same card by **right-clicking the mic button** in the message box, which opens it in a dialog over your composer rather than sending you to Settings.

### Where to configure it

**Settings -> Voice Control**, under the **Dictation (FlowVoice)** heading (`VoiceSpeechSettings.tsx`) — or, without leaving the composer, right-click the mic button in the message box, which opens that same card in a dialog over your session (flowvoice.contract.md `settings-open-where-the-user-is`). Progressive disclosure: only the master **Enable FlowVoice** toggle is visible until it is on; the card then shows Server / Microphone / Accessibility / Automation status pills with a Recheck action, both hotkey recorders, cleanup level, language, microphone picker, the **Sound cues** / **Send automatically** / **Show floating pill** toggles, and the history list. macOS requires Microphone, Accessibility **and Automation**. Accessibility covers the hold-to-talk global hotkey and the synthetic keystroke; Automation (System Settings -> Privacy & Security -> Automation) is what lets Omniscio drive System Events to perform the paste, and it needs the `com.apple.security.automation.apple-events` entitlement in the packaged build before macOS will even offer it. Naming only Accessibility is a known trap: the toggle-mode hotkey is an Electron `globalShortcut` that needs NO grant at all, so dictation can appear to work perfectly while the cross-app paste is dead. Granting cannot be automated, so the card deep-links to the System Settings pane. The Automation pill is OBSERVED rather than polled, because no read of that grant is free of a consent prompt: it stays neutral until the first dictation, turns green once an Apple event goes through, and turns red only when macOS refuses one for lack of consent — which also shows a hint naming the System Settings path and noting that dictating into Omniscio itself keeps working without it (`noteAppleEventOutcome` in `permissions.ts`). Windows/Linux need no explicit grants (`permissions.ts`).

## How it behaves

### Hotkeys (two independent slots)

Settings keep two global hotkeys (`hotkeyHold` + `hotkeyToggle` in `src/shared/flowvoice/types.ts`); either or both can be bound, and either may be `''` (unbound):

- **Hold to talk**: default `CommandOrControl+Shift+D` (Ctrl+Shift+D on Windows/Linux, Cmd+Shift+D on macOS). Bound via uiohook-napi (keydown starts, keyup stops), the only way to drive press-and-hold; a bare held modifier (e.g. Right Cmd, stored as `key:NNNNN`) or a modifier-only combo (e.g. Ctrl+Win as `Control+Meta`) is valid in this slot only. A quick tap is still safe: a key-up that races ahead of the async start is remembered and stops the capture as soon as it begins (`controller.ts`).
- **Toggle on / off**: unbound by default. Bound via Electron `globalShortcut` (press flips recording on/off); it refuses bare modifiers and combos already taken by another Omniscio shortcut, and the settings card warns when a bare modifier is recorded into this slot.

Both are rebindable in-card via the click-to-capture `HotkeyRecorder`. (The earlier single-hotkey + `mode` shape was migrated to this pair; `FlowVoiceMode` survives only for the settings-store migration.)

**The macOS `fn` / Globe key cannot be used as the dictation hotkey.** It is unsupported at all three layers independently — Chromium fires no keyboard event for it (so the recorder cannot even see the press), `uiohook-napi`'s key table has no `fn` entry (Hold mode cannot match it), and Electron accelerators have none (Toggle mode cannot register it). The recorder used to sit on "Listening…" forever with no explanation; it now says the key did not reach the app and names `fn` as an example, and says the same for any key that arrives but maps to no bindable accelerator (a media key, F25+). Pick a normal key or combination instead.

### How dictation works (pipeline)

Hotkey down: remember the frontmost app, then a hidden recorder-host window captures the mic and streams audio live to the FlowVoice `/stream` WebSocket. Hotkey up: the mic is held live for a further 500 ms before it closes (the stop tail — a key-up lands on the last syllable, and without it that syllable arrived clipped), then the stream finishes and the cleaned text is pasted at the focused app's cursor via the clipboard-paste path (a constant Cmd/Ctrl+V; the transcript is never interpolated into an osascript/PowerShell string; the previous clipboard is restored on a deferred timer; max 5,000 characters per injection - see `inject.ts`). Transcription streams live, and every captured frame is ALSO written to a file on disk as it arrives, so a failed stream is recovered from that saved audio instead of being lost (see *When transcription fails* below). When Omniscio itself is frontmost the text lands in Omniscio's focused composer (session or Team Chat, whichever is active) through the clipboard too: it is written to the clipboard and pasted by the app's own Paste command (`pasteIntoWebContents`, no PowerShell/osascript hop), which is what puts every dictation into clipboard history where it can be re-pasted. The trade is that a paste replaces whatever is selected in the composer. **Send automatically** (`autoSubmitInAmc`, default on) submits it through Omniscio's real send path (honoring submit-key mode), never an OS-level keystroke. Which app is frontmost is read via `lsappinfo` (LaunchServices), which needs no macOS permission — deliberately, because that answer decides WHERE the text goes. Omniscio recognises its own window under every name it can be reported as — the bundle display name `Omniscio`, the lowercase `app.getName()` value `omniscio`, and `Electron` in dev — compared case-insensitively (`isSelfAppName` in `context.ts`). Those names differ by one capital letter, and matching only `app.getName()` once classified every in-app dictation as another app, so it took the osascript paste that macOS refused without Automation consent instead of the permission-free in-app paste. If the frontmost app genuinely cannot be identified, the target is `unknown`, which is NOT treated as "Omniscio": the transcript is pasted at the OS cursor and **never** auto-submitted. That distinction exists because it once collapsed — an unreadable frontmost app meant "Omniscio is the target", so every dictation was pasted into the composer and auto-sent while the user was typing in another app entirely. The FlowVoice server URL + token are baked into the build (`config.ts`, committed while the repo is private; env vars `AMC_FLOWVOICE_API_URL`/`_TOKEN` override for self-hosting), so a default user configures no API keys.

Cleanup happens per the `cleanupLevel` setting: `off` (verbatim — every word as spoken), `light` (strip fillers and false starts, keep wording), `strong` (adds punctuation, casing and paragraph breaks — the default). On the Gemini engine the speech model does the cleaning itself, so `light` runs no extra pass and `off` switches the model into verbatim mode (see below). Language is selectable (10 choices, default English) and the microphone is pickable (`microphoneId`, default system default).

### Speech engine — hosted (default) or Gemini (your key, or ours)

`sttEngine` (`src/shared/flowvoice/types.ts`, default `hosted`) chooses which engine transcribes dictation:

- **`hosted`** (default, unchanged): the first-party hosted `/stream` service — streaming STT plus a server-side LLM cleanup pass, company-funded (metered under the `flowvoice` cost source at `FLOWVOICE_USD_PER_MIN`).
- **`gemini`**: Google **Gemini 3.5 Transcribe Live**, reusing the same `GeminiLiveSttProvider` as the in-app mic (`src/main/services/voice/stt-providers/gemini.ts`) wrapped in a StreamSession adapter (`flowvoice/gemini-stream.ts`). **Which credential it runs on is decided per dictation** (`flowvoice/gemini-credential-resolution.ts`):
  - **You have your own Gemini key** → the socket goes DIRECT to Google on your key, billed to you under `stt-gemini`, and it respects the shared per-install STT daily cap (`sttDailyCapUSD`, $5 default). Your key always wins; the company path never overrides it.
  - **You do not** → the socket goes to our own **gateway relay** (`/v1/stt/gemini` on jls-gateway), which holds the company Gemini key server-side and proxies to Google. You must be signed in, because the relay bills that dictation to your identity against the company's caps. The company key is never sent to your machine in any form.
  - **Neither** → dictation refuses with the message that fits: signed out, or no key.

Only ONE side ever bills a dictation. On the relay path the gateway bills it, and the desktop deliberately writes no local `stt-gemini` row — billing it in both places would charge the same audio twice and throttle you against your own cap for spend you never made. Neither Gemini path touches the company `flowvoice` meter, which belongs to the hosted engine.

The Gemini option is **gated** behind the in-development `gemini-stt` flag (the SAME flag as the in-app Gemini STT — off by default; reveal it under **Settings → Lab → Voice & Speech**). That flag is also what keeps company spend bounded to people who deliberately opted in. Either way the audio goes to Google; on the relay path it travels via our gateway rather than straight from your machine.

> **Worth knowing before a long session.** Relayed dictation draws on the same company AI allowance as every other company-funded feature, so roughly half an hour of dictation in a day can exhaust a user's daily allowance and pause their other company-funded AI until midnight. Adding your own Gemini key removes that entirely. A single relayed dictation is also bounded at about four minutes — Cloud Run hangs up at its own request deadline, so the relay closes first with a reason rather than letting you be cut off mid-sentence.

Two engine-specific mechanics: **(1) Audio format** — Gemini consumes raw PCM16 16 kHz mono, so when it is selected the recorder-host captures via the shared AudioWorklet at a forced 16 kHz `AudioContext` (`voice-audio-capture.ts`) instead of the WebM/Opus `MediaRecorder` the hosted path uses. **(2) Cleanup** — fidelity is chosen at the ENGINE. `off` asks Gemini for `mode: VERBATIM`, so you get every word exactly as spoken; skipping only the local pass would not have been enough, because Gemini cleans upstream and `off` would still have returned edited text. Otherwise Gemini is asked for `mode: SMART`, so it returns an ALREADY-CLEANED transcript (Google removes fillers, stutters and false starts, resolves self-corrections, and punctuates). So `light` runs NO local pass on this engine — it would be a second disfluency pass over text with none left, which deleted real words ("Testing, testing, testing." came back "Testing."). `strong` still runs locally via the shared lightweight-LLM lane (`llmProviderService.chat`), but ADDITIVE-ONLY: it may add punctuation, casing and paragraph breaks and may never remove or reword anything, because grammar-fixing turned "Can you push the deploy?" into "We can push the deploy." Fail-soft to the raw text on any error (`flowvoice/transcript-cleanup.ts`). The `hosted` engine's server-side cleanup is unchanged, and delivery (paste/insert at the cursor) is identical for both engines.

### What dictation costs, and where you see it

The FlowVoice card (Settings → Voice Control → Dictation) carries a **Usage & spend** block. It deliberately shows **two figures that are not the same figure**, side by side and labelled as such, because dictation is billed on two different ledgers depending on which engine ran it:

- **Today's metered dictation — your machine's own ledger.** This is what the **daily cap** governs, and the cap is editable right in the block, next to the number it limits (a meter you can see but not adjust would be a dead end). Setting it to 0 turns the cap off; the block then says so rather than painting a full red bar at a user who simply disabled it. The figure is an **estimate**, and the block says so, showing the per-minute rate it meters at — nothing here is a bill.
- **Company-funded relay — our cost, not yours.** Dictation that runs on Omniscio's own key (the relay path used when you have no Gemini key of your own) is billed to the company, recorded on the gateway, and shown here as covered by Omniscio. It never touches your bill and it is **not** what the daily cap meters. This one is a **monthly** figure, unlike the daily one above it — the block states each period rather than implying a single window.

**An unreadable figure is never rendered as $0.00.** If the gateway is unreachable, you are signed out, or the reply carries no figure for dictation, the line reads **"Unknown — we could not read this from your account right now."** Zero because nothing was used and "we could not find out" are different facts, and only one of them would be honest. The same rule governs the history list below: a dictation row shows a cost only where the local ledger genuinely metered one, and shows **duration only** otherwise — an unmetered row (any Gemini dictation, which bills elsewhere by design) carries no cost field at all, so nothing can mistake absence for a zero.

This is also where you can see the mechanism described above working: relayed dictation now draws on your monthly company AI allowance, exactly as the note below describes — until 2026-09-16 that record was written to the company-wide and daily counters only, so the allowance meter reported none of it. Company money was being spent and the number that reports it was invisible.

### Cleanup runs while you speak, not after

On the Gemini path the cleanup pass used to be a single LLM call over the entire transcript, fired only once dictation stopped — the one post-speech step whose cost grew with how long you talked, so a five-minute dictation waited far longer than a thirty-second one before anything appeared ("it sends the whole block at once at the end and then does the translation").

Gemini finalises transcript segments continuously mid-speech, so those segments now feed a sentence-boundary chunker (`flowvoice/transcript-chunker.ts`) and each completed chunk is cleaned in the background while the user keeps talking; stopping only has to clean the small remaining tail. The post-speech wait is therefore roughly **constant** instead of proportional to dictation length.

Delivery is deliberately unchanged — still ONE complete cleaned block inserted at the cursor. This is a latency change, not incremental delivery: no partial text, and no text that visibly rewrites itself. Chunks split only between complete sentences (so a self-correction is never cut in half) with a word-boundary fallback for unpunctuated speech, run at most two at a time (the lightweight-LLM lane shares a circuit breaker app-wide), join by index rather than completion order, and fail soft per chunk to their raw text. The cleanup token budget scales with input length — it was a flat 2,048, which silently discarded the cleaned text and delivered the raw transcript on very long dictations.

### HUD + composer mic button

A transparent always-on-top HUD pill (`FlowVoiceHud.tsx`, rendered in the `flowvoice-hud.html` overlay window) floats while FlowVoice is enabled. It **follows your theme** — surfaces, accent, body font and app text size all come from the same tokens as the rest of Omniscio, like the recording and meeting HUDs (until 2026-09-15 it kept FlowVoice's own warm-dark + amber palette and its own font, so it looked like a different product parked on top of the app). The pill shows: an idle pill (hover reveals Dictate / History / Settings; clicking it starts dictation), a live waveform with Cancel and Accept while recording, an accent-coloured sweep while processing, and a dot + message on error. The error pill **clears itself after about 8 seconds** and clicking it retries — it deliberately floats even for a hidden-pill user, because dictation is driven by a global hotkey from inside other apps, so an in-app surface would not reach you at the moment a dictation fails. (It used to have no way down at all: the pre-flight failures — signed out, no Gemini key, daily cap reached — stop before dictation starts, so "Dictation is unavailable — sign in to enable it." stayed on screen until some later dictation happened to succeed.) The **Show floating pill** toggle (`showIdleHud`, default on) hides the idle pill; the HUD still appears during active dictation. **Sound cues** (`playSoundFeedback`, default on) plays a soft chime at the recording start and stop boundaries, synthesised in the recorder-host window. The stop chime sounds on the key-up itself — not when the mic actually closes — so the confirmation never lags half a second behind the release while the stop tail (above) runs. **The start chime is the honest "you may speak now" signal, and it is worth waiting for**: the mic cannot open instantly (opening the OS capture device measured 49-196 ms on a dev box, occasionally seconds), and the chime is only dispatched once capture is genuinely live. Words spoken before it are not clipped or buffered — there is no audio yet, so they never existed. Everything that does NOT need the microphone is already built ahead of the press (flowvoice.contract.md `nothing-sits-ahead-of-the-microphone`), so that wait is the device open and nothing else; `startDictation` logs the measured gap (`[flowvoice] mic live Nms after start`) on every dictation.

The session composer, the council panel, and the **Team Chat composer** each get a FlowVoice mic button in their left tool zone (`FlowVoiceMicButton.tsx`): click to start, click again to stop, a small pulsing red dot while recording; right-click opens the FlowVoice settings in a dialog over the composer — you tune dictation where you are, never bounced out of the session to the Settings page (flowvoice.contract.md `settings-open-where-the-user-is`) — and the tooltip shows both configured hotkeys. The dialog hosts the same settings card the Settings page shows, imported rather than forked, and it is a SIBLING of the button rather than nested inside it, so switching FlowVoice off from within the card hides the button without slamming the dialog shut. It focuses the composer textarea before starting so the OS-level paste at stop has a guaranteed sink. The button self-hides when FlowVoice is disabled or on mobile (FlowVoice is desktop-only), and the **Show mic button in the message box** toggle (`showComposerMicButton`, default on) hides it for a user who drives dictation entirely from the hotkeys — dictation stays fully live, only the button goes. Together with `showIdleHud` that means every always-visible FlowVoice affordance can be dismissed WITHOUT turning the feature off (flowvoice.contract.md `every-visible-affordance-has-its-own-opt-out`); both are presentation-only and are read at the RENDER site, never on a start/stop path. Visibility reaches the button by being mirrored into `FlowVoiceState` and fed to the renderer's `flowvoice-store` by one app-level `FLOWVOICE_STATE_CHANGED` listener — there is one button per keep-alive SessionPanel, so none of them subscribe individually.

### When transcription fails — your recording is saved

Every dictation is saved to disk while you speak (`saved-recordings.ts`, under `userData/flowvoice-recordings/`), so no failure can lose what you said:

- **The connection drops mid-dictation, or the result comes back broken or empty.** On release, the app transcribes the saved audio again by itself: first a fresh session of the same engine (the audio replayed at 8× real time), then the built-in offline transcriber (free, no network, English only, for the Gemini engine's audio). The recovered text is pasted where you were, like a normal dictation. A transcript from a stream that errored part-way is never pasted as if it were complete.
- **Nothing could transcribe it right now** (for example, you're offline and the offline model isn't available). The recording is kept. The pill says *"Couldn't transcribe right now — your recording is saved. Click to try again."*, and clicking it retries that recording and pastes at your cursor.
- **Saved recordings list.** Settings → Voice Control → Dictation (FlowVoice) → History shows a **Saved recordings** group (only when there is something in it): each recording that failed or was recovered, with **Transcribe again** (the text goes into history and is copied to your clipboard) and **Delete**.
- **Quitting or crashing mid-dictation** no longer throws the recording away: it appears in that list on the next start. Pressing ✕ Cancel still discards it.
- **The paste itself fails** after a good transcript: the pill says the text is on your clipboard and in your history.
- **Privacy:** a successful dictation's audio is deleted as soon as its text is safely in history. Saved recordings are cleared by erase-all and by signing out, and are aged out by your data-retention window like dictation history.

### History (inline in the settings card)

The standalone FlowVoice sidebar tile + `FlowVoiceView` panel were **retired**; `FLOWVOICE_PROJECT_ID` (`__flowvoice__`) survives in `src/shared/virtual-project-ids.ts` only so stored references still resolve. History now renders inline in the settings card: the 20 most recent dictations (the store caps at 50 in `flowvoice-history.json` under `userData`, `history-store.ts`), each showing target app, window title, relative time, and duration, plus the cleaned text, a raw-transcript disclosure when it differs, per-row Copy / Remove, and a Clear all action. The list live-updates on the `flowvoice:result` push.

**History is age-purged as well as capped.** The 50-record cap is only the first limit — what is kept is also deleted once it is older than your **data-retention window** (30 days by default, set at Settings → Data & Storage), so a dictation you made months ago is gone whether or not the 50 slots filled up. The purge covers everything a dictation record holds: the **raw transcript**, the **cleaned text**, and **the app and window title you dictated into** — that last one is why the window matters if you dictate into something sensitive. It runs on the app's periodic retention sweep rather than on a timer of its own, and the **Clear all** action and the per-row **Remove** clear a record immediately if you would rather not wait. There is a kill switch for the sweep (`AMC_DISABLE_FLOWVOICE_HISTORY_PURGE=1`).

## For agents

### Trigger from the CLI

Three POST routes on the CLI control server (`127.0.0.1:19519`, bearer-authed, 10/min mutation budget) drive dictation headlessly — the parity twins of the `FLOWVOICE_START` / `_STOP` / `_CANCEL` command channels, calling the **same** `controller.ts` entry points as the hotkey and the composer mic button (`cli-server-flowvoice-routes.ts`):

- **`POST /flowvoice/start`** — begin a dictation (`startDictation('manual')`). Returns `{ state }`; the returned phase reflects what happened (`recording` on success, `idle`/`error` otherwise).
- **`POST /flowvoice/stop`** — finish, transcribe via the selected engine, deliver at the OS cursor, and return `{ result, state }` so a script can read the transcript back. `result` is the just-completed history record, or `null` when the capture produced nothing (a freshness window keeps a stale earlier row from being returned).
- **`POST /flowvoice/cancel`** — abort the in-flight capture without transcribing. Returns `{ state }`.

Saved recordings (dictations whose transcription failed or was recovered) have their own routes:

- **`GET /flowvoice/recordings`** — list them (plain read).
- **`POST /flowvoice/recordings/:id/retry`** — transcribe one again through the same recovery chain; returns `{ result }` (`{ ok: true, text }` or `{ ok: false, reason: 'failed' }`), `404` for an unknown id, `409` while it is already being transcribed. A CLI retry never pastes at the OS cursor.
- **`DELETE /flowvoice/recordings/:id`** — delete one.

`GET /flowvoice/state` is the read twin (plain bearer read, 60/min). Because delivery still types at the OS cursor, a CLI-triggered dictation lands in whatever app is focused, exactly like the hotkey.

Two settings routes make the whole feature drivable headlessly:

- **`GET /flowvoice/settings`** — read the persisted settings (plain read). Behavioral config only; FlowVoice's store holds no credential.
- **`PATCH /flowvoice/settings`** — update them. Partial, and `.strict()`: anything omitted keeps its current value, while an unknown or ill-typed key is a `400`.

**Turning dictation on headlessly is `PATCH /flowvoice/settings` with `{"enabled":true}`** — `startDictation` no-ops while FlowVoice is off, so a start against a disabled install returns an `idle` phase rather than an error. That PATCH goes through the same `applyFlowVoiceSettings` path the settings card uses, so enabling also binds the global hotkey, records the F037 audio-egress consent, mirrors the runtime state the composer mic button reads, floats the HUD, and warms the recorder host. A store-only write would skip all five and leave dictation looking on but inert.

## Related

The in-app microphone — talking to a session instead of typing at the OS cursor — is a different feature on the [Voice and TTS](voice-and-tts.md) page. Language and accent handling for both is on [Language support](language-support.md), and the cost surfaces FlowVoice meters into are described on [Cost control](cost-control.md).
