---
title: Voice L3 - Ask About a Thread (session:converse)
---

# Voice L3 - Ask About a Thread (session:converse)

## What it is

Voice L3 lets you ask hands-free questions about a specific agent session without switching to that session, typing anything, or interrupting whatever is running. You trigger it exactly the same way you trigger any other voice command -- say the wake word, or press Alt+V -- and then ask about the thread by name.

There are two phrasings:

- **Summary** -- "tell me about the auth refactor thread" -- Omniscio speaks the Omni briefing for that session. This is the same per-session summary Omni already knows how to produce (Ctrl+Shift+J on the session), now available completely hands-free and by name.
- **Question** -- "in the auth refactor thread, what is blocking it?" -- Omniscio pulls that thread's transcript and answers your specific question using xAI Grok via OpenRouter. The answer is drawn only from the content of that thread, spoken aloud.

After Omniscio speaks the summary or the answer, it asks "What would you like to know?" and listens for a spoken follow-up. Ask your next question out loud, Omniscio answers it from the same thread, then asks again -- a real back-and-forth. Each follow-up is answered with memory of the questions and answers already spoken in this exchange, so "what about the second one?" resolves against what was just said. The conversation runs until you stay silent, you reach the per-conversation turn limit, or the daily cost cap is hit.

The feature is **off by default**. Nothing changes in how Omniscio behaves until you turn it on.

## Where to find it

There is no panel of its own. The switch is in **Settings → Voice Control**, under **Ask about a thread (voice)**. You reach the feature by voice instead — the wake word or Alt+V — and its briefing, its answer and its **daily limit reached** message are spoken aloud, with a toast shown whenever a spoken session name matches no open session.

## How it behaves

### How to turn it on

Open **Settings &rarr; Voice Control** and find the **Ask about a thread (voice)** toggle (`voiceL3Enabled`). Flip it on.

That is the only switch. Once it is enabled, the `session:converse` intent is live and Omniscio will recognize the summary and question phrasings described above.

### How session name resolution works

When you name a session in your voice command ("the auth refactor thread"), Omniscio does a fuzzy match against the names of your open sessions. If it finds a match, that session's transcript is used.

If you do not name a session, Omniscio falls back to the currently active session (the one you are looking at). So "tell me what's happening" with no session name briefed produces a summary of whichever session is in focus.

If the name you spoke does **not** match any open session, Omniscio does **not** fall back - it **stops and tells you** ("I could not find a thread called auth refactor."), spoken aloud and shown as a toast, so it never summarizes the wrong thread. Say the name again more clearly, or open the thread you mean. (A deictic reference like "this conversation" is not a name - it always resolves to the session on screen.)

### Summary path -- Omni briefing reuse

When the intent resolves as a summary request, Omniscio calls the same Omni briefing logic that powers Ctrl+Shift+J. The result is a spoken summary of that session's recent activity -- what the agent is working on, what it last reported, whether it needs your attention. The Omni briefing returns a cached result when the session has not changed since the last briefing, or makes one briefing-model call (Haiku) when the cache is stale.

### Question path -- Grok via OpenRouter, read-only, transcript-only

When the intent resolves as a question (you phrased it as "in X, what..." or "ask X..."), Omniscio sends your question to xAI Grok via OpenRouter. The model setting is `voiceL3Model` (defaults to an `x-ai/grok-*` slug). The context given to Grok is drawn exclusively from that session's transcript -- it does not have access to other sessions, your filesystem, the web, or any live agent state. It reads what is already in the thread and answers from that.

The operation is strictly **read-only**. Nothing is sent to the live agent in that session, nothing is queued, nothing triggers an approval. The session continues running exactly as it was.

The answer is spoken aloud via your selected TTS provider, the same voice you hear for Voice Report Back and other TTS surfaces.

### Follow-up turns (bounded multi-turn)

Right after the summary or an answer finishes playing, Omniscio speaks a short cue ("What would you like to know?") and re-arms the microphone for your next follow-up. Your follow-up is captured as plain dictation -- it is not run through the command parser -- and is answered from the same thread by the same model. After it answers, Omniscio asks again, and keeps going as a real conversation.

Each follow-up carries the prior spoken questions and answers as conversation history, so a follow-up that refers back ("and the owner?", "what about the second one?") is answered in context. The current question always carries the freshest thread transcript; the prior exchange rides along as memory.

The loop ends as soon as any of these happens: you stay silent (or the listen times out), the exchange reaches its turn limit (`L3_MAX_FOLLOWUP_TURNS`, currently 6 follow-ups), an answer fails, or the daily cost cap is hit. The re-arm reuses the normal per-microphone-cycle capture, so your microphone setup and silence detection behave exactly as they do for ordinary dictation. The mic is only live during a follow-up listen, not while Omniscio is speaking, so you cannot interrupt a summary or an answer by talking over it (see the limitation below).

### When L3 cannot answer: offer to ask the live agent (L3 -> L4 handoff)

Sometimes the answer to your question simply is not in the thread's transcript. When that happens, instead of dead-ending on "I do not see that in this thread", Omniscio can offer to go ask the live agent directly.

This only happens when **all** of the following are true:

- You asked a **question** (not a summary). The briefing path never makes this offer.
- **Ask the live agent (voice)** is also turned on (`voiceL4Enabled`, off by default), so both L3 and L4 must be enabled.
- The thread you asked about is **still running** (it has not ended, errored, been paused, or archived). L4 never starts a thread, so there is no point offering to ask an agent that is not there.

When all three hold, Omniscio waits for its not-found answer to finish, then speaks a short offer that names the thread: "Want me to ask the Auth Refactor agent directly?" It listens once for your reply.

- Say **yes** (or "go ahead", "sure", and similar) and Omniscio hands your original question, unchanged, to the existing Ask-the-live-agent queue (L4). From there it behaves exactly like asking the agent directly: the question is delivered as a genuine operator turn when the thread next goes idle, and you are cued when the agent answers.
- Say **no**, stay silent, or say something unrelated, and nothing is queued and nothing is sent. The turn simply ends.

The offer replaces the usual "What would you like to know?" follow-up for that turn, so you never get two prompts stacked back to back. Nothing reaches any agent without your spoken yes, and the handed-off question is asked read-only (the agent is asked to answer, not to take action).

### Daily cost cap

Omniscio tracks the cumulative spend for Voice L3 question answers against a daily limit (`voiceL3DailyCapUSD`, default $1.00). The cap is checked before each Grok call goes out. If today's spend has already reached the limit, Omniscio speaks a "daily limit reached" message and stops -- no Grok call is made and no cost is incurred. In a multi-turn exchange the cap is enforced on every turn: the moment it is reached, Omniscio speaks the limit message once and ends the conversation rather than re-prompting for another follow-up. The cap resets at midnight (calendar day boundary). You can adjust the cap in Settings.

### Limitations -- what this version does not yet do

This page describes what is built and shipped in Voice L3, including the bounded multi-turn conversation. The items below are intentionally deferred.

- **Handoff only on a not-found answer.** When L3 can answer from the transcript, it just answers; it does not offer the live agent. The handoff (see above) fires only when L3 cannot answer, both L3 and L4 are on, and the thread is running. You can also ask the live agent directly at any time with the separate Ask-the-live-agent voice command.
- **No barge-in during the spoken answer.** Once the TTS playback starts, you cannot interrupt it by speaking. This matches the same limitation as Voice Report Back slice 1.
- **No mid-speech cancel.** Toggling `voiceL3Enabled` off while audio is already playing does not stop the current utterance.

## Related

- [voice-and-tts.md](voice-and-tts.md) -- the underlying Voice Control and TTS stack this feature depends on (provider selection, voice IDs, chunking, caching, Omni session briefings).
- [voice-report-back.md](voice-report-back.md) -- the companion feature that speaks LLM-layer prose answers back after ordinary voice commands.
