Omniscio documentation
Browse all documentation
  1. Getting Started13
  2. Sessions & Agents115
  3. Inbox & Notifications59
  4. Projects & Tasks95
  5. Automation & Scheduling75
  6. Knowledge & Memory26
  7. AI Features60
  8. Integrations100
  9. Plugins & Marketplace33
  10. Cloud & Teams56
  11. Settings & Customization58
  12. Account & Billing28
  13. Troubleshooting84
  14. CLI & API Reference22
  15. Legal & Policies4
  16. Uncategorised22

FlowVoice (OS-wide dictation + history)

Omniscio's OS-wide dictation: hold a hotkey, speak, and the cleaned transcript is typed at the cursor in whatever app is focused — a browser, an editor, Slack — not just Omniscio's own message box. Covers the two hotkey slots, the hosted and Gemini speech engines and what each costs, the floating HUD and composer mic button, where to configure it, and the promise that a dictation is delivered whole or its audio is kept.

What it is

FlowVoice is Omniscio's OS-wide dictation feature: hold a global hotkey, speak, release, and the cleaned transcript is typed at the OS cursor in whatever app is focused (browser, editor, Slack, Gmail), not just Omniscio's own message box. It is a self-contained port of the standalone FlowVoice app, living entirely under src/main/services/flowvoice/ (main) and src/renderer/src/features/flowvoice/ (renderer), with its own flowvoice:* IPC channels and its own JSON settings/history stores; it does not touch Omniscio's existing voice stack. See the contract at .claude/memory/contracts/flowvoice.contract.md. The feature is off by default and is not gated behind an unreleased-feature flag: the settings card is visible to every user.

Where to find it

Settings → Voice Control, under the Dictation (FlowVoice) heading. Without leaving a session you can also reach the same card by right-clicking the mic button in the message box, which opens it in a dialog over your composer rather than sending you to Settings.

The one-time offer to turn it on

Because it ships off and lives behind Settings, most people would never learn it exists. So Omniscio offers it to you once: a card in your inbox reading Dictation, anywhere, with a single Turn on dictation button. Pressing it enables the feature there and then; dismissing the card, or just leaving it, means it never comes back. Anyone who has already turned dictation on is never offered it. On a phone the offer never appears — dictation types at a desktop cursor.

Where to configure it

Settings -> Voice Control, under the Dictation (FlowVoice) heading (VoiceSpeechSettings.tsx) — or, without leaving the composer, right-click the mic button in the message box, which opens that same card in a dialog over your session (flowvoice.contract.md settings-open-where-the-user-is). Progressive disclosure: only the master Enable FlowVoice toggle is visible until it is on; the card then shows Microphone / Accessibility / Automation status pills (and a Server pill as well on the hosted engine — it reports the hosted service, so it is absent when you dictate on Gemini) with a Recheck action, both hotkey recorders, the speech engine picker, cleanup level, language, microphone picker, the Sound cues / Send automatically / Show floating pill toggles, and the history list. macOS requires Microphone, Accessibility and Automation. Accessibility covers the hold-to-talk global hotkey and the synthetic keystroke; Automation (System Settings -> Privacy & Security -> Automation) is what lets Omniscio drive System Events to perform the paste, and it needs the com.apple.security.automation.apple-events entitlement in the packaged build before macOS will even offer it. Naming only Accessibility is a known trap: the toggle-mode hotkey is an Electron globalShortcut that needs NO grant at all, so dictation can appear to work perfectly while the cross-app paste is dead. Granting cannot be automated, so the card deep-links to the System Settings pane. The Automation pill is OBSERVED rather than polled, because no read of that grant is free of a consent prompt: it stays neutral until the first dictation, turns green once an Apple event goes through, and turns red only when macOS refuses one for lack of consent — which also shows a hint naming the System Settings path and noting that dictating into Omniscio itself keeps working without it (noteAppleEventOutcome in permissions.ts). Windows/Linux need no explicit grants (permissions.ts).

How it behaves

Hotkeys (two independent slots)

Settings keep two global hotkeys (hotkeyHold + hotkeyToggle in src/shared/flowvoice/types.ts); either or both can be bound, and either may be '' (unbound):

  • Hold to talk: default Control+Meta — hold Ctrl+Windows (Ctrl+Cmd on macOS). It replaced CommandOrControl+Shift+D, which also fires Chrome/Edge/Firefox "bookmark all tabs" and VS Code's Run view, and the key hook only OBSERVES keystrokes, so both fired. Press Ctrl first: the hook starts on the Windows key's keydown with Ctrl already held, and that same "a key went down before Win" is what keeps the Start menu from opening when the key is released (the same reason Win-key hotkey tools synthesise a leading Ctrl). Releasing either key ends the dictation. Ctrl+Windows is a live Windows chord: the OS desktop shortcuts (Ctrl+Win+Left/Right, Ctrl+Win+D) start a hold as well. Nothing is typed from one — a recording that catches no speech is retained and the phase returns to idle without a paste (MIN_RECOVERABLE_MS in controller.ts). Bound via uiohook-napi (keydown starts, keyup stops), the only way to drive press-and-hold; a bare held modifier (e.g. Right Cmd, stored as key:NNNNN) or a modifier-only combo is valid in this slot only — the Toggle slot cannot bind one, and the settings card warns when you try. A quick tap is still safe: a key-up that races ahead of the async start is remembered and stops the capture as soon as it begins (controller.ts).
  • Toggle on / off: unbound by default. Bound via Electron globalShortcut (press flips recording on/off); it refuses bare modifiers and combos already taken by another Omniscio shortcut, and the settings card warns when a bare modifier is recorded into this slot.

Both are rebindable in-card via the click-to-capture HotkeyRecorder. Because both are global, neither can be a key you type normally — a letter, digit, Space, Backspace, an arrow, punctuation, or one of those with only Shift: the recorder refuses it and asks you to add Ctrl, Alt or the Windows/Command key (a function key, or modifiers alone, are fine), the settings route refuses to save one, and a value an older version saved is simply not bound. The recorder also stops listening the moment you click or tab away, so a key typed in another field is never captured. (The earlier single-hotkey + mode shape was migrated to this pair; FlowVoiceMode survives only for the settings-store migration.)

The macOS fn / Globe key cannot be used as the dictation hotkey. It is unsupported at all three layers independently — Chromium fires no keyboard event for it (so the recorder cannot even see the press), uiohook-napi's key table has no fn entry (Hold mode cannot match it), and Electron accelerators have none (Toggle mode cannot register it). The recorder used to sit on "Listening…" forever with no explanation; it now says the key did not reach the app and names fn as an example, and says the same for any key that arrives but maps to no bindable accelerator (a media key, F25+). Pick a normal key or combination instead.

How dictation works (pipeline)

Hotkey down: remember the frontmost app, then a hidden recorder-host window captures the mic and streams audio live to the FlowVoice /stream WebSocket. Hotkey up: the mic is held live for a further 500 ms before it closes (the stop tail — a key-up lands on the last syllable, and without it that syllable arrived clipped), then the stream finishes and the cleaned text is pasted at the focused app's cursor via the clipboard-paste path (a constant Cmd/Ctrl+V; the transcript is never interpolated into an osascript/PowerShell string; max 5,000 characters per paste, and a longer dictation is pasted in pieces - see inject.ts). The dictation stays on your clipboard afterwards - nothing puts what you had copied before back - so whatever happens to the paste, Ctrl/Cmd+V gives you the words; a dictation pasted in pieces has the whole text put back on the clipboard once the last piece lands. Transcription streams live, and every captured frame is ALSO written to a file on disk as it arrives, so a failed stream is recovered from that saved audio instead of being lost (see When transcription fails below). When Omniscio itself is frontmost the text lands in Omniscio's focused composer (session or Team Chat, whichever is active) through the clipboard too: it is written to the clipboard and pasted by the app's own Paste command (pasteIntoWebContents, no PowerShell/osascript hop), which is what puts every dictation into clipboard history where it can be re-pasted. The trade is that a paste replaces whatever is selected in the composer. Send automatically (autoSubmitInAmc, default OFF) submits it through Omniscio's real send path (honoring submit-key mode, and the after-sending setting — which, in a project view, keeps you on a session that was not waiting on you, so dictating into a running session never moves you off it), never an OS-level keystroke. Left off, a dictation into the message box simply waits there for you to read it and press Enter — and only when the text actually landed in a session's message box: dictating into the search box or a notes editor pastes there and sends nothing, and a dictation that finishes after you switched to another app is pasted at that app's cursor, unsent. Which app is frontmost is read via lsappinfo (LaunchServices), which needs no macOS permission — deliberately, because that answer decides WHERE the text goes. Omniscio recognises its own window under every name it can be reported as — the bundle display name Omniscio, the lowercase app.getName() value omniscio, and Electron in dev — compared case-insensitively (isSelfAppName in context.ts). Those names differ by one capital letter, and matching only app.getName() once classified every in-app dictation as another app, so it took the osascript paste that macOS refused without Automation consent instead of the permission-free in-app paste. If the frontmost app genuinely cannot be identified, the target is unknown, which is NOT treated as "Omniscio": the transcript is pasted at the OS cursor and never auto-submitted. That distinction exists because it once collapsed — an unreadable frontmost app meant "Omniscio is the target", so every dictation was pasted into the composer and auto-sent while the user was typing in another app entirely. The FlowVoice server URL + token are resolved AT RUNTIME (config.ts, cloud-credential.ts), and only while the hosted engine is selected — the shared credential is fetched solely for that engine: a signed-in fetch, held in memory and re-fetched on a short TTL, so a server-side rotation propagates without an app restart. Nothing is baked into a shipped build — the release guard fails a build that carries a real credential. Env vars AMC_FLOWVOICE_API_URL/_TOKEN (or a gitignored dev credentials file) cover npm run dev and self-hosting. Either way a default user configures no API keys.

Cleanup happens per the cleanupLevel setting: off (verbatim — every word as spoken), light (strip fillers and false starts, keep wording), strong (adds punctuation, casing and paragraph breaks — the default). On the Gemini engine the speech model does the cleaning itself, so light runs no extra pass and off switches the model into verbatim mode (see below). Language is selectable (10 choices, default English) and the microphone is pickable (microphoneId, default system default).

Speech engine — Gemini (default) or hosted

sttEngine (src/shared/flowvoice/types.ts, default gemini) chooses which engine transcribes dictation. The picker is in the FlowVoice card for every user; it is no longer behind an in-development flag. A saved engine is always honoured — only the default moved, so an install that already holds one keeps it (a stored hosted stays on hosted).

The new default applies to installs that have no engine stored, and an install that saved FlowVoice settings on an earlier build does have one. Older builds wrote the whole settings object on every save, so enabling FlowVoice was enough to store "sttEngine": "hosted" — the default then — even though the picker was behind a Lab flag and most of those users never saw it. Those installs therefore stay on the hosted engine and keep the shared company credential; the flip reaches new installs. This is a decision on record (2026-10-05), not a gap left open: reinterpreting a pre-flip hosted as un-chosen would move those users onto the relay and start billing their own company allowance, which is a separate call.

  • gemini (default): Google Gemini 3.5 Transcribe Live, reusing the same GeminiLiveSttProvider as the in-app mic (src/main/services/voice/stt-providers/gemini.ts) wrapped in a StreamSession adapter (flowvoice/gemini-stream.ts). Which credential it runs on is decided per dictation (flowvoice/gemini-credential-resolution.ts):
    • You have your own Gemini key → the socket goes DIRECT to Google on your key, billed to you under stt-gemini, and it respects the shared per-install STT daily cap (sttDailyCapUSD, $5 default). Your key always wins; the company path never overrides it.
    • You do not → the socket goes to our own gateway relay (/v1/stt/gemini on jls-gateway), which holds the company Gemini key server-side and proxies to Google. You must be signed in, because the relay bills that dictation to your identity against the company's caps. The company key is never sent to your machine in any form.
    • Neither → dictation refuses with the message that fits: signed out ("Sign in to use Gemini dictation without your own key — or add your Gemini API key in Settings → Accounts → Gemini"), or no key.
  • hosted (selectable, not the default): the first-party hosted /stream service — streaming STT plus a server-side LLM cleanup pass, company-funded (metered under the flowvoice cost source at FLOWVOICE_USD_PER_MIN). It runs on a shared company credential, handed to every signed-in user, which is why it is no longer what a new user starts on.

Why Gemini is the default. The hosted engine gives every signed-in user the same company key with no per-user minute cap. The relay instead bills each dictation to the signed-in user against company caps, and it has an offline fallback, so dictation keeps working when the relay cannot be reached.

Only ONE side ever bills a dictation. On the relay path the gateway bills it, and the desktop deliberately writes no local stt-gemini row — billing it in both places would charge the same audio twice and throttle you against your own cap for spend you never made. Neither Gemini path touches the company flowvoice meter, which belongs to the hosted engine.

The shared hosted credential is fetched only while it is needed. The app asks the cloud for the hosted server credential (flowvoiceCredential) only when the selected engine is hosted (flowvoice/cloud-credential.ts), including at boot — every health check runs only on the hosted engine: the startup probe, and the settings card's Recheck (which also re-reads your OS permissions, and does that on any engine). Before that gate landed (2026-10-04) the boot health check resolved it unconditionally, so every signed-in user was handed the shared company credential whether or not they used the hosted engine. What stayed open until 2026-10-05 is the Recheck probe itself: on a Gemini install it still RAN the hosted probe on every mount of the card. By then the credential gate refused that fetch, so no credential was handed out — the cost was a needless probe, a "no server config (signed out or unconfigured)" log for a correctly configured install, and that refusal STORED as a verdict for a service this engine never uses. The engine gate closes those three, not a leak. Switching engines also clears the stored Server verdict, so the card never shows the previous engine's result for the one you just picked.

Worth knowing before a long session. Relayed dictation draws on the same company AI allowance as every other company-funded feature, so roughly half an hour of dictation in a day can exhaust a user's daily allowance and pause their other company-funded AI until midnight. Adding your own Gemini key removes that entirely. A single relayed dictation is also bounded at about four minutes — Cloud Run hangs up at its own request deadline, so the relay closes first with a reason rather than letting you be cut off mid-sentence.

Two engine-specific mechanics: (1) Audio format — Gemini consumes raw PCM16 16 kHz mono, so when it is selected the recorder-host captures via the shared AudioWorklet at a forced 16 kHz AudioContext (voice-audio-capture.ts) instead of the WebM/Opus MediaRecorder the hosted path uses. (2) Cleanup — fidelity is chosen at the ENGINE. off asks Gemini for mode: VERBATIM, so you get every word exactly as spoken; skipping only the local pass would not have been enough, because Gemini cleans upstream and off would still have returned edited text. Otherwise Gemini is asked for mode: SMART, so it returns an ALREADY-CLEANED transcript (Google removes fillers, stutters and false starts, resolves self-corrections, and punctuates). So light runs NO local pass on this engine — it would be a second disfluency pass over text with none left, which deleted real words ("Testing, testing, testing." came back "Testing."). strong still runs locally via the shared lightweight-LLM lane (llmProviderService.chat), but ADDITIVE-ONLY: it may add punctuation, casing and paragraph breaks and may never remove or reword anything, because grammar-fixing turned "Can you push the deploy?" into "We can push the deploy." Fail-soft to the raw text on any error (flowvoice/transcript-cleanup.ts). The hosted engine's server-side cleanup is unchanged, and delivery (paste/insert at the cursor) is identical for both engines.

What dictation costs, and where you see it

The FlowVoice card (Settings → Voice Control → Dictation) carries a Usage & spend block. It deliberately shows two figures that are not the same figure, side by side and labelled as such, because dictation is billed on two different ledgers depending on which engine ran it:

  • Today's metered dictation — your machine's own ledger. This is what the daily cap governs, and the cap is editable right in the block, next to the number it limits (a meter you can see but not adjust would be a dead end). Setting it to 0 turns the cap off; the block then says so rather than painting a full red bar at a user who simply disabled it. The figure is an estimate, and the block says so, showing the per-minute rate it meters at — nothing here is a bill.
  • Company-funded relay — our cost, not yours. Dictation that runs on Omniscio's own key (the relay path used when you have no Gemini key of your own) is billed to the company, recorded on the gateway, and shown here as covered by Omniscio. It never touches your bill and it is not what the daily cap meters. This one is a monthly figure, unlike the daily one above it — the block states each period rather than implying a single window.

An unreadable figure is never rendered as $0.00. If the gateway is unreachable, you are signed out, or the reply carries no figure for dictation, the line reads "Unknown — we could not read this from your account right now." Zero because nothing was used and "we could not find out" are different facts, and only one of them would be honest. The same rule governs the history list below: a dictation row shows a cost only where the local ledger genuinely metered one, and shows duration only otherwise — an unmetered row (any Gemini dictation, which bills elsewhere by design) carries no cost field at all, so nothing can mistake absence for a zero.

This is also where you can see the mechanism described above working: relayed dictation now draws on your monthly company AI allowance, exactly as the note below describes — until 2026-09-16 that record was written to the company-wide and daily counters only, so the allowance meter reported none of it. Company money was being spent and the number that reports it was invisible.

Cleanup runs while you speak, not after

On the Gemini path the cleanup pass used to be a single LLM call over the entire transcript, fired only once dictation stopped — the one post-speech step whose cost grew with how long you talked, so a five-minute dictation waited far longer than a thirty-second one before anything appeared ("it sends the whole block at once at the end and then does the translation").

Gemini finalises transcript segments continuously mid-speech, so those segments now feed a sentence-boundary chunker (flowvoice/transcript-chunker.ts) and each completed chunk is cleaned in the background while the user keeps talking; stopping only has to clean the small remaining tail. The post-speech wait is therefore roughly constant instead of proportional to dictation length.

Delivery is deliberately unchanged — still ONE complete cleaned block inserted at the cursor. This is a latency change, not incremental delivery: no partial text, and no text that visibly rewrites itself. Chunks split only between complete sentences (so a self-correction is never cut in half) with a word-boundary fallback for unpunctuated speech, run at most two at a time (the lightweight-LLM lane shares a circuit breaker app-wide), join by index rather than completion order, and fail soft per chunk to their raw text. The cleanup token budget scales with input length — it was a flat 2,048, which silently discarded the cleaned text and delivered the raw transcript on very long dictations.

HUD + composer mic button

A transparent always-on-top HUD pill (FlowVoiceHud.tsx, rendered in the flowvoice-hud.html overlay window) floats while FlowVoice is enabled. It follows your theme — surfaces, accent, body font and app text size all come from the same tokens as the rest of Omniscio, like the recording and meeting HUDs (until 2026-09-15 it kept FlowVoice's own warm-dark + amber palette and its own font, so it looked like a different product parked on top of the app). The pill shows: an idle pill (hover reveals Dictate / History / Settings; clicking it starts dictation), a live waveform with Cancel and Accept while recording, an accent-coloured sweep while processing, and a dot + message on error. The error pill clears itself after about 8 seconds and clicking it retries — it deliberately floats even for a hidden-pill user, because dictation is driven by a global hotkey from inside other apps, so an in-app surface would not reach you at the moment a dictation fails. (It used to have no way down at all: the pre-flight failures — signed out, no Gemini key, daily cap reached — stop before dictation starts, so "Dictation is unavailable — sign in to enable it." stayed on screen until some later dictation happened to succeed.) A fourth refusal joined them, covering both things the recorded reading can already tell us: when it says the engine is unreachable, or that the server refused our credential, the start re-probes once with a deliberately short ceiling (you are holding a key down) and refuses before the microphone opens only when that re-probe is evidence — the server answered the same way, or the connection itself was refused — so you are not left speaking a whole dictation into a session that was never going to work. A reading that says down refuses as an outage; a credential the server has turned away refuses with the sign-in copy and drops the cached credential, so the "try again" it asks for genuinely re-fetches rather than replaying the token that was just rejected. A server that has come back, a server that simply did not answer in time, and a fresh boot whose first read has not landed all still start. And when the server refuses the stream upgrade itself, that refusal's own HTTP status now picks the copy — a limit on your allowance says so and says it clears on its own (naming no time, because the server answers the same status both for a daily budget and for one that frees up in seconds), an expired dictation sign-in says so and tells you to try again to refresh it — instead of every refusal arriving as the same connection-error line. A refused sign-in drops the cached credential on every road a refusal can reach you by, including the replay behind the pill's try-again, so a retry genuinely re-fetches instead of looping on the token the server just rejected — and it says so on every one of those roads: a refusal that happens before any request is made (a dead sign-in, an engine known to be down, the daily cap, a signed-out token) now shows you its own reason instead of the generic "your recording is saved, click to try again", which used to invite exactly the retry that would fail the same way. The Show floating pill toggle (showIdleHud, default on) hides the idle pill; the HUD still appears during active dictation. Sound cues (playSoundFeedback, default on) plays a soft chime at the recording start and stop boundaries, synthesised in the recorder-host window. The stop chime sounds on the key-up itself — not when the mic actually closes — so the confirmation never lags half a second behind the release while the stop tail (above) runs. The start chime is the honest "you may speak now" signal, and it is worth waiting for: the mic cannot open instantly (opening the OS capture device measured 49-196 ms on a dev box, occasionally seconds), and the chime is only dispatched once capture is genuinely live. Words spoken before it are not clipped or buffered — there is no audio yet, so they never existed. After a dictation the microphone stays open for 30 seconds, so a second dictation in that window starts with no device-open wait at all (the system's microphone-in-use light stays on for those 30 seconds). Everything that does NOT need the microphone is already built ahead of the press (flowvoice.contract.md nothing-sits-ahead-of-the-microphone), so that wait is the device open and nothing else; startDictation logs the measured gap ([flowvoice] mic live Nms after start) on every dictation.

The session composer, the council panel, and the Team Chat composer each get a FlowVoice mic button in their left tool zone (FlowVoiceMicButton.tsx): click to start, click again to stop, a small pulsing red dot while recording; right-click opens the FlowVoice settings in a dialog over the composer — you tune dictation where you are, never bounced out of the session to the Settings page (flowvoice.contract.md settings-open-where-the-user-is) — and the tooltip shows both configured hotkeys. The dialog hosts the same settings card the Settings page shows, imported rather than forked, and it is a SIBLING of the button rather than nested inside it, so switching FlowVoice off from within the card hides the button without slamming the dialog shut. It focuses its own composer textarea on every press — start and stop — so the OS-level paste at stop has a guaranteed sink. Both presses need it: delivery is a paste, which lands wherever document focus is at that moment, so a user who pressed the mic without clicking the text box first had the transcript pasted into whatever happened to hold focus (or nowhere at all), and one who clicked elsewhere in the page mid-recording had the stop transcript go the same way. The button sits inside one composer and a press on it means "put the words here", so both presses take focus back there (within the app's focused window — a press cannot drag OS window focus back from another app, and the OS-wide hotkey path never routes through the button). The button self-hides when FlowVoice is disabled or on mobile (FlowVoice is desktop-only), and the Show mic button in the message box toggle (showComposerMicButton, default on) hides it for a user who drives dictation entirely from the hotkeys — dictation stays fully live, only the button goes. Together with showIdleHud that means every always-visible FlowVoice affordance can be dismissed WITHOUT turning the feature off (flowvoice.contract.md every-visible-affordance-has-its-own-opt-out); both are presentation-only and are read at the RENDER site, never on a start/stop path. Visibility reaches the button by being mirrored into FlowVoiceState and fed to the renderer's flowvoice-store by one app-level FLOWVOICE_STATE_CHANGED listener — there is one button per keep-alive SessionPanel, so none of them subscribe individually.

When the app you are dictating into is running as administrator

Windows refuses to let one program type into a window that belongs to a higher-privilege process (a protection called UIPI). A paste into an app you launched with Run as administrator is therefore refused by the operating system, and the app on the receiving end never sees a keystroke. Omniscio detects that before it presses paste and tells you plainly: the dictation is on your clipboard, so press Ctrl+V in that window to put it where you want it. Nothing is lost — the words are also in the dictation history — and an app Omniscio cannot measure is never blamed for it.

A dictation is delivered whole — or its audio is kept, never both lost

Two promises, added 2026-09-30 after a 53-second dictation delivered 285 characters and its audio was then deleted as a success:

  • What the engine heard is delivered, even when it never finalized it. Speech engines stream a dictation on two channels: the committed text, and the in-progress read of the sentence still being spoken. The in-progress channel is not a preview — measured against Google's Gemini Live model, a dictation's last sentences can exist ONLY there, with the committed text stopping mid-thought. FlowVoice now takes whatever the in-progress channel carries that the committed text does not, so the words arrive instead of being dropped. Nothing is ever repeated: text the committed channel already delivered is recognised and skipped.
  • A transcript that cannot match its audio keeps that audio. Once a dictation's text is safely in history the recording is normally deleted — it has done its job. But text far too short for the audio it came from is not a success worth being certain about, so the recording is kept instead and appears under Saved recordings as Transcription may be incomplete — retry for the rest, with the same Transcribe again action as any other kept recording. Ordinary dictation is unaffected, twice over: only dictations of twenty seconds or more are judged at all (a nine-second sentence is too brief for the ratio to mean anything), and every long dictation on record reads 15–21 characters per second of speech against an incident reading of 5.3.

If a dictation ever comes back looking short, the pill says so — "Some of that dictation may be missing — your recording is saved. Click to transcribe it again." — and clicking it re-transcribes that recording rather than starting a new one.

Your recording is kept — nothing about the dictation deletes it

Every dictation keeps its audio on your machine, not just the ones that went wrong. A transcript is what the speech engine heard, and engines drop words — so the recording, which is the only copy of what you actually said, outlives the dictation that produced it. Every dictation appears under Saved recordings with its audio, ready to re-transcribe or delete.

What a dictation's own outcome can never do is delete its audio: not a successful transcription, not an empty one, not a silent one, not a start that failed. A recording goes away only when you say so — ✕ Cancel on the pill mid-dictation, Delete or Clear all in the list, erasing your data, or signing out — or when your data-retention window passes (Settings → Data & Storage, 30 days by default). That window is the natural cleanup: it is far longer than noticing a dictation that came back short takes, and you can lengthen or shorten it whenever you like.

When transcription fails — your recording is saved

Every dictation is saved to disk while you speak (saved-recordings.ts, under userData/flowvoice-recordings/), so no failure can lose what you said:

  • The connection drops mid-dictation, or the result comes back broken or empty. On release, the app transcribes the saved audio again by itself: first a fresh session of the same engine (the audio replayed at 8× real time), then the built-in offline transcriber (free, no network, English only, for the Gemini engine's audio). The recovered text is pasted where you were, like a normal dictation. A transcript from a stream that errored part-way is never pasted as if it were complete.
  • Nothing could transcribe it right now (for example, you're offline and the offline model isn't available). The recording is kept. The pill says "Couldn't transcribe right now — your recording is saved. Click to try again.", and clicking it retries that recording and pastes at your cursor.
  • Nothing was transcribed at all — you spoke, and both transcribers answered silence. The recording is kept, and the pill now SAYS so instead of settling back to idle in silence (until 2026-10-06 it said nothing anywhere, so "the app ignored me" and "nothing was said" looked identical). A press too short to hold a spoken word is the one ending that stays quiet: it is kept, but an accidental tap of the hotkey is not announced. Which line you read comes from how loud the capture actually was: "No speech was heard — nothing was transcribed." when the room was silent, or, when the capture carried real but very quiet audio (something like 26 dB below an ordinary dictation), "No speech was heard — your microphone was very quiet. Check the input device." Neither carries a click-to-retry, because re-transcribing the same audio would fail the same way.
  • Saved recordings list. Settings → Voice Control → Dictation (FlowVoice) → History shows a Saved recordings group: by default only the recordings that need a look — failed, possibly incomplete, or recovered — each with Transcribe again (the text goes into history and is copied to your clipboard) and Delete. Every other dictation's audio is kept too, behind a Show all recordings toggle, so the list is not one row per dictation.
  • Quitting or crashing mid-dictation no longer throws the recording away: it appears in that list on the next start. Pressing ✕ Cancel still discards it.
  • The paste itself fails after a good transcript: the pill says the text is on your clipboard and in your history.
  • A window pops up while you dictate into Omniscio (Windows: a console window flashing up, for example): it never receives your words. Omniscio tells it apart from an app you switched to yourself by when its program started - after your dictation began means it appeared on its own - takes the focus back and pastes into Omniscio. If it cannot get the focus back, the pill says the text is on your clipboard. An app you deliberately switched to still gets the dictation at its cursor, unsent.
  • The offline transcriber narrates silence ("[BLANK_AUDIO]", "(music)"): those markers are stripped before anything is delivered, so a silent recording comes back as silence rather than as that text in your message box.
  • Privacy: a dictation's audio is kept on your machine until your data-retention window passes (Settings → Data & Storage, 30 days by default) — that window is the only automatic cleanup, and it is one you control. A dictation's own outcome never deletes it; nor does anything else, except the deletions you ask for: delete it, clear all, erase your data, or sign out.

History (inline in the settings card)

The standalone FlowVoice sidebar tile + FlowVoiceView panel were retired; FLOWVOICE_PROJECT_ID (__flowvoice__) survives in src/shared/virtual-project-ids.ts only so stored references still resolve. History now renders inline in the settings card: the 20 most recent dictations (the store caps at 50 in flowvoice-history.json under userData, history-store.ts), each showing target app, window title, relative time, and duration, plus the cleaned text, a raw-transcript disclosure when it differs, per-row Copy / Remove, and a Clear all action. The list live-updates on the flowvoice:result push.

History is age-purged as well as capped. The 50-record cap is only the first limit — what is kept is also deleted once it is older than your data-retention window (30 days by default, set at Settings → Data & Storage), so a dictation you made months ago is gone whether or not the 50 slots filled up. The purge covers everything a dictation record holds: the raw transcript, the cleaned text, and the app and window title you dictated into — that last one is why the window matters if you dictate into something sensitive. It runs on the app's periodic retention sweep rather than on a timer of its own, and the Clear all action and the per-row Remove clear a record immediately if you would rather not wait. There is a kill switch for the sweep (AMC_DISABLE_FLOWVOICE_HISTORY_PURGE=1).

For agents

Trigger from the CLI

Three POST routes on the CLI control server (127.0.0.1:19519, bearer-authed, 10/min mutation budget) drive dictation headlessly — the parity twins of the FLOWVOICE_START / _STOP / _CANCEL command channels, calling the same controller.ts entry points as the hotkey and the composer mic button (cli-server-flowvoice-routes.ts):

  • POST /flowvoice/start — begin a dictation (startDictation('manual')). Returns { state }; the returned phase reflects what happened (recording on success, idle/error otherwise).
  • POST /flowvoice/stop — finish, transcribe via the selected engine, deliver at the OS cursor, and return { result, state } so a script can read the transcript back. result is the just-completed history record, or null when the capture produced nothing (a freshness window keeps a stale earlier row from being returned).
  • POST /flowvoice/cancel — abort the in-flight capture without transcribing. Returns { state }.

Saved recordings (dictations whose transcription failed or was recovered) have their own routes:

  • GET /flowvoice/recordings — list them (plain read).
  • POST /flowvoice/recordings/:id/retry — transcribe one again through the same recovery chain; returns { result } ({ ok: true, text } or { ok: false, reason: 'failed' }), 404 for an unknown id, 409 while it is already being transcribed. A CLI retry never pastes at the OS cursor.
  • DELETE /flowvoice/recordings/:id — delete one.

GET /flowvoice/state is the read twin (plain bearer read, 60/min). Because delivery still types at the OS cursor, a CLI-triggered dictation lands in whatever app is focused, exactly like the hotkey.

Two settings routes make the whole feature drivable headlessly:

  • GET /flowvoice/settings — read the persisted settings (plain read). Behavioral config only; FlowVoice's store holds no credential.
  • PATCH /flowvoice/settings — update them. Partial, and .strict(): anything omitted keeps its current value, while an unknown or ill-typed key is a 400.

Turning dictation on headlessly is PATCH /flowvoice/settings with {"enabled":true} — startDictation no-ops while FlowVoice is off, so a start against a disabled install returns an idle phase rather than an error. That PATCH goes through the same applyFlowVoiceSettings path the settings card uses, so enabling also binds the global hotkey, records the F037 audio-egress consent, mirrors the runtime state the composer mic button reads, floats the HUD, and warms the recorder host. A store-only write would skip all five and leave dictation looking on but inert.

Related

The in-app microphone — talking to a session instead of typing at the OS cursor — is a different feature on the Voice and TTS page. Language and accent handling for both is on Language support, and the cost surfaces FlowVoice meters into are described on Cost control.

Last verified 2026-10-06