---
title: Cheap / utility LLM routing
---

# Cheap / utility LLM routing

## What it is

The small, utility LLM calls Omniscio makes on your behalf — AI-suggested reply chips, Omni per-session briefings, TTS summaries (the short spoken readouts), and voice-intent parsing (what turns "snooze this til tomorrow" into an action) — run on a cheap, fast model rather than your main Claude session.

**That cheap model is hardcoded to Groq's Llama 4 Scout 17B — there is no user setting for it.** A "Settings → AI Provider" panel used to let you pick it from a unified model dropdown; it was removed (2026-06-20) because the choice added surface area without real value — the default served virtually everyone, and every feature that genuinely needs a different model already pins its own in code (Omni → Anthropic, session titles → OpenRouter + Qwen, email pre-screen → OpenRouter's Gemini 2.5 Flash-Lite). The underlying `defaultLlmProvider` / `defaultLlmModel` / `lightweightLlmApiKey` settings still exist (kept for upgrade-migration safety, and `lightweightLlmApiKey` doubles as the multi-model council's Together-AI key slot) but are no longer user-editable.

The same **unified model picker** that powered that panel still drives four OTHER settings surfaces that DO let you choose a model — see "Configuring it" below.

If Groq fails in a way that clearly cost nothing (bad request, 5xx, auth error, or an empty answer), Omniscio silently retries the same prompt against Anthropic Haiku so the feature still works. The exception is a **timeout or a cancelled request**: the provider may have already run (and billed) it before the reply was lost, so Omniscio does NOT fire a second paid call for it — it surfaces the error instead, so you're never charged twice for one request. A run of repeated timeouts still trips the circuit breaker, which then routes everything to Anthropic. **This fallback is silent by design** — the old amber settings-panel chip and the "open settings" toast were removed with the panel.

**What is NOT routed through the hardcoded default (deliberately):**

1. **Session title generation** is hardcoded to OpenRouter + Qwen 3 32B (the same model [Plain Speak](plain-speak.md) uses) regardless of the default. Llama 3.1 8B Instant on Groq broke prompt adherence on short prompts ("What project is this?" produced an editorialized "(empty response — insufficient context)" instead of a real title). Pinning to a stronger model with the same prompt fixed it. There's no settings UI to change this — it's the same default for everyone. Multimodal titles (image attachments) still route to Anthropic directly per the carveout below.
2. **The inbound email pre-screen** (the classifier that inspects incoming email for prompt injection before any agent sees it) and **the AI conditions evaluator** inside Automation Rules. Both of those are security-critical paths that pin their model in code rather than honoring the default — a user-picked third-party provider could be tricked into a wrong verdict, and the cost of either is far higher than the savings. The AI conditions evaluator calls Anthropic Haiku directly. The inbound email pre-screen runs on OpenRouter's `google/gemini-2.5-flash-lite` (a fixed pin, auto-falling back to Claude Haiku 4.5 whenever OpenRouter fails or its breaker is open). Either way these are not user-configurable.

## Where to find it

### Configuring it

**There is nothing to configure for the utility calls — the cheap model is hardcoded.** To change it in a future model swap, edit `DEFAULT_MODELS.groq` (the per-provider default in [/src/shared/model-latest.ts](/src/shared/model-latest.ts)) or the `CHEAP_UTILITY_PROVIDER` constant in [/src/main/services/llm-provider-service.ts](/src/main/services/llm-provider-service.ts) — see "How it works". Provider API keys for the few keyless providers (e.g. Together AI for the council) live under **Settings → Accounts**, where the same save-time key check ("Checking…" / red-reject / **"Save anyway"** on a network blip or on a provider rejection) applies — mechanics in `secret-key-input-contract.md` I9.

The **unified model picker** is still wired into four settings surfaces that DO let you pick a model, all reading from the same catalog: **Voice Commands → Omni briefing model**, **Catch-Up Card settings → summary model**, **Email Summarizer settings → default model**, and **Quick Replies → AI suggestion model**. Each surface picks its own recommended fallback (e.g. Sonnet for catch-up summaries) but uses the same dropdown — every supported model (Anthropic, Groq, Together AI, DeepSeek, OpenAI, OpenRouter) grouped by provider, friendly names, a `(reasoning)` flag on reasoning models, and the same inline amber reasoning-model warning. These pickers are unaffected by the AI Provider panel's removal.

## How it behaves

### How it works

The unified picker is the React component [/src/renderer/src/components/ui/ModelPicker.tsx](/src/renderer/src/components/ui/ModelPicker.tsx). It reads from a single curated catalog at [/src/shared/model-catalog.ts](/src/shared/model-catalog.ts) — every catalog entry carries a `provider`, an optional `isReasoning` flag, and an optional `recommended: true` marker (one per provider). The picker renders the standardized `<Select>` abstraction (fed via its `options` prop with provider-grouped option lists, not native `<select>`/`<option>` children), derives the provider for any chosen model via `getProviderForModel(modelId)`, and emits the `recommended` per-provider default as the fallback when the user clears the value.

The four remaining picker surfaces each write their OWN settings field (e.g. `jarvisBriefingModel`, `continuousSummaryModel`) via `onChange` — there is no longer an `AiProviderSettings.tsx` writing the global `defaultLlmProvider` / `defaultLlmModel` fields. Those global fields are no longer written by any UI; `chat()` ignores them for the cheap utility default (it uses the hardcoded `CHEAP_UTILITY_PROVIDER`), and they persist only for upgrade-migration safety and the council's Together-AI key slot.

Every routed feature calls a single function, `llmProviderService.chat({ source, system, messages, accountId, ... })`, defined in [/src/main/services/llm-provider-service.ts](/src/main/services/llm-provider-service.ts). For the cheap utility default it uses the hardcoded `CHEAP_UTILITY_PROVIDER` (`'groq'`) unless the caller passes `opts.provider`; it still honors the optional `settings.lightweightLlmApiKey` override (for keyless providers like Together AI), calls the chosen provider via its OpenAI-compatible endpoint (Groq: `api.groq.com/openai/v1`, Together: `api.together.xyz/v1`, DeepSeek: `api.deepseek.com/v1`, OpenAI: `api.openai.com/v1`, OpenRouter: `openrouter.ai/api/v1`), and records per-call token cost via `trackApiCost`. Callers can pin a specific provider for one call by passing `provider: 'anthropic'` in the options — useful when a call can't tolerate the fallback pattern.

The model name itself is resolved by `resolveModel()` in the same file, which reads from a centralized `DEFAULT_MODELS` constant defined in [/src/shared/model-latest.ts](/src/shared/model-latest.ts) (re-exported through the `/src/shared/types.ts` barrel) — the single source of truth for per-provider defaults. Resolution priority: explicit `opts.model` → per-source override from `GROQ_MODELS_BY_SOURCE` (Groq only — e.g. voice-intent uses `llama-3.1-8b-instant`) → provider default from `DEFAULT_MODELS`. (The old user-picked `settings.defaultLlmModel` step was removed with the AI Provider panel.) To change a default model in a future model swap, edit `DEFAULT_MODELS` in [/src/shared/model-latest.ts](/src/shared/model-latest.ts); to change a per-source pin, edit `GROQ_MODELS_BY_SOURCE` in [/src/shared/types/claude-models.ts](/src/shared/types/claude-models.ts).

**Title-gen is the one source that opts out of the resolver entirely.** Inside [/src/main/services/ai/ai-suggestion-service.ts](/src/main/services/ai/ai-suggestion-service.ts) `callWithApiKey()`, the cheap-eligible branch detects `costSource === 'title-gen'` and passes `provider: 'openrouter'` + `model: 'qwen/qwen3-32b'` (`MODEL_QWEN3_32B_OPENROUTER`) through to `llmProviderService.chat()`. Those explicit values short-circuit both the user's `settings.defaultLlmProvider` and the per-source `GROQ_MODELS_BY_SOURCE` lookup. Qwen 3 is a **reasoning** model, and the `reasoning: { effort: 'low' }` that `chat()` sends when a caller states no budget does NOT bound it on OpenRouter (on/off only there, and its busiest host, DeepInfra, ignores the parameter). So the Qwen naming calls — session titles and SMS contact names — switch thinking off themselves: Qwen's `/no_think` soft switch in the message plus `effort: 'none'` (`QWEN_NO_THINK_SUFFIX` in [/src/main/services/ai/llm-egress-client.ts](/src/main/services/ai/llm-egress-client.ts)). Measured 2026-09-29: with thinking, a title took a median 10.5–15 s and ran past its 20 s deadline on about 1 in 6 real first messages, and the SMS-name call's 50-token budget ran out mid-thought on every real call (each then re-asked of Haiku); without it, a title takes ~1 s at half the cost and blind-judged no worse. The Anthropic Haiku fallback still applies — if OpenRouter returns a non-2xx, the circuit opens and subsequent title-gen calls go to Haiku exactly like the other cheap-eligible sources. That Haiku fallback runs on the user's OWN key, so a **keyless PAID** user gets a company-gateway Haiku backstop instead (`callTextTitleViaGateway`, fired when the `title-gen` circuit is OPEN **or** the cheap call hit a definitive non-timeout failure — a cheap TIMEOUT still waits, so no double-charge — under the `title-gen` daily-$ cap); and if every AI path is down the final give-up derives a clean, scaffolding-stripped LOCAL title (`deriveLocalTitle`) rather than echoing the raw first-60 chars — see `title-gen-recovery-contract.md` I5/I6. A sustained cheap-provider outage (which the fallback otherwise masks) also raises a self-clearing "AI features are using a backup provider" notice — see the gateway-routing contract's "unmask" §. Multimodal title-gen (when the first message has images) bypasses this routing entirely and goes straight to Anthropic Haiku 4.5 through the company gateway lane — never OpenRouter — see "Multimodal title carveout" below.

A circuit breaker guards the cheap provider: **5 consecutive failures open the circuit**, after which Omniscio fails fast to Anthropic Haiku for a **60-second cooldown** (transient errors) or a **10-minute cooldown** (permanent errors like HTTP 401/403, "invalid api key", or "model not found"). A single success in the half-open probe slot closes the circuit again. The circuit-break event is still broadcast on the `LLM_PROVIDER_FALLBACK` IPC channel as breaker instrumentation, but it now has NO renderer listener — the settings-panel chip and the permanent-error toast (`useAppLlmProviderAlerts.ts`) were removed with the AI Provider panel, so the cheap→Anthropic fallback is silent. The channel is baselined as emit-only in `main-renderer-push-emitter-coverage.test.ts`.

**Double-charge protection on a timeout-after-accept.** A metered call that _times out_ or is _cancelled_ is ambiguous — the provider may have already run and billed it before the reply was lost. Two guards keep Omniscio from paying twice for one logical call: (1) on a timeout/abort, `chat()` does NOT fall back to a second paid Anthropic call — it re-throws (only clearly-unbilled failures like 4xx / 5xx / empty-content fall back), routed through `isTimeoutOrAbortError()`; and (2) every metered `create()` carries a fresh per-request `Idempotency-Key` header so the SDKs' own internal retries dedupe into a single charge where the provider honors it — real on OpenAI, a forward-compatible no-op on Anthropic (no idempotency support today). The timeout still counts toward the circuit breaker. Full posture, the SDK-plumbing detail, and the cross-reference to the Google method-restriction rule live in the external-integration-retry contract, `a-metered-ai-call-never-double-charges`.

**Reasoning-model handling (Groq).** Reasoning models (`gpt-oss-*`, `qwen-qwq-*`, `deepseek-r1`, `deepseek-reasoner`, OpenAI `o1`/`o3`) split the `max_tokens` budget between an internal chain-of-thought and the visible content — once the budget is exhausted, the model returns empty content. Omniscio handles this in three layers so titles/suggestions still land:

1. **Cheap-provider request** (Groq only): `llmProviderService.callCheapProvider()` always sends `reasoning_effort: 'low'` to keep the reasoning step short. Non-reasoning models on Groq ignore the field. Together AI, DeepSeek, OpenAI, and OpenRouter do not currently use this field via the same path.
2. **Title-gen budget**: `aiSuggestionService.generateSessionTitle()` allocates `maxTokens: 1000` (was `50`) so even a chatty reasoning step still leaves comfortable headroom for the visible title.
3. **Empty-content tripwire**: if the cheap provider returns an empty `text` despite the above, `callCheapProvider()` throws a synthetic 502 so the existing Anthropic Haiku fallback kicks in instead of propagating an empty string up the stack.

The picker surfaces an inline amber note when the chosen model is flagged `isReasoning` in the catalog (or matches the legacy regex pattern for custom ids), so the user understands why responses on those models can feel slower than non-reasoning ones.

**Which providers need a key.** Groq and DeepSeek are the two cheap-utility providers a distribution can be pre-configured with, so on some builds the cheap utility calls work before you paste anything. Everything else is keyless until you supply one: Together AI, OpenAI and OpenRouter have no built-in key, so a call to them returns a 401 from `callCheapProvider()`, which then falls through to the Anthropic fallback described above. To use them, add the key under **Settings → Accounts**. Your own `lightweightLlmApiKey` (`Settings → Accounts`) always takes precedence over any built-in key.

The current routed call sites are [/src/main/services/ai/ai-suggestion-service.ts](/src/main/services/ai/ai-suggestion-service.ts) (suggestion chips — `suggestions`; titles also call `llmProviderService.chat` but with hardcoded `provider: 'openrouter'` + `model: 'qwen/qwen3-32b'` overrides as described above, so the user's setting does not apply to them), [/src/main/services/session/session-briefing-llm.ts](/src/main/services/session/session-briefing-llm.ts) (Omni), [/src/main/services/tts/tts-service.ts](/src/main/services/tts/tts-service.ts) (TTS readout summaries), and [/src/main/services/voice/voice-llm-providers.ts](/src/main/services/voice/voice-llm-providers.ts) (voice intent). The security carveouts live in [/src/main/services/email/email-prescreen.ts](/src/main/services/email/email-prescreen.ts) and [/src/main/services/automation/legacy/automation-condition-evaluator.ts](/src/main/services/automation/legacy/automation-condition-evaluator.ts). The conditions evaluator instantiates the Anthropic SDK directly. The email pre-screen calls `llmProviderService.chat()` with a hardcoded `provider: 'openrouter'` + `model: 'google/gemini-2.5-flash-lite'` (so the user's dropdown does not apply); `chat()`'s own fallback then routes a failed or circuit-open call to Claude Haiku 4.5.

**Image title routing — Haiku DIRECT from Anthropic via the gateway (NOT OpenRouter).** When a session's first operator message carries an image attachment (PNG / JPEG / GIF / WebP), the image title call routes through `llmProviderService.chat({ provider: 'anthropic', model: MODEL_HAIKU_LATEST, routeAnthropicThroughGateway: true, source: 'title-gen' })` with the image attached. The `routeAnthropicThroughGateway` opt-in sends it to the company gateway's `/internal/v1/anthropic` lane, where LiteLLM's `anthropic/*` route calls Anthropic's Messages API DIRECTLY on the company key (Haiku 4.5). **Text titles stay on Qwen via OpenRouter; image titles bypass OpenRouter entirely.** It is metered under the same `title-gen` gateway feature as the text titles. This replaced a bespoke gateway `fetch` (`callMultimodalTitleApi`) + a direct-Anthropic-vision fallback (`callAnthropicVisionTitle`) on 2026-07: the old fetch rode a SEPARATE `title-gen-multimodal` gateway feature that sat at a constant 402 "at capacity" cap (since the 2026-07-06 gateway routing change), so screenshot titles broke app-wide; routing to the Anthropic lane under `title-gen` dodges it (see [internal-ai-gateway-routing-contract](/.claude/memory/contracts/internal-ai-gateway-routing-contract.md)). Haiku 4.5 was picked over Qwen3-VL (a prior route) because Qwen tended to describe the most salient thing on screen rather than weight the operator's typed words as the primary intent; it costs ~$1/M in, ~$5/M out (~$0.002 per titled image). The first up to 3 images on the first message pass through as Anthropic-shape `{type:'image',source:{type:'base64',media_type,data}}` blocks (`chat()`'s `toOpenAiMessage` converts them to OpenAI `image_url` blocks for the gateway's OpenAI-compatible endpoint, which LiteLLM translates back to Anthropic); each is size-capped at 5 MB, unsupported MIME types (PDFs / Office formats) are silently skipped, and text-document attachments (`.md` / `.txt` / `.json` / etc.) are inlined as truncated text alongside. When the gateway is unavailable (signed out / toggle off), `chat()` falls back to a DIRECT Anthropic call WITH the image on the user's own key — so an image-only first message still gets a real title (instead of the empty-prompt "Untitled Session" sentinel), and it is Haiku direct from Anthropic either way. The company lane's model allow-list ([internal-models.ts](/gateway/gateway/internal-models.ts)) must list the Anthropic Haiku id or the lane 403s. On any miss the image call returns null and title-gen falls through to the text-only route; the orchestrator's retry schedule (`TITLE_RETRY_DELAYS = [5_000, 45_000]`) covers retries. A `title-gen` daily-$ backstop in `aiSuggestionService.callImageTitleViaAnthropic()` caps runaway image-regen spend (chat()'s per-source breaker handles transient backoff). Cost is tracked by `chat()` under `source: 'title-gen'` + provider `anthropic`. This lives in `aiSuggestionService.callImageTitleViaAnthropic()`, gated by `pickUsableImages()` in [/src/main/ipc/title-orchestrator.ts](/src/main/ipc/title-orchestrator.ts).

**No minimum character threshold.** In `ai_title` mode (the default), Omniscio generates an AI title for any non-empty first message regardless of length — a 3-word prompt like `"fix the bug"` still gets a meaningful AI-generated title rather than echoing the message text. Only the empty + zero-images case skips the AI call. The `first_prompt` mode opt-out (Settings → Sessions → Session naming) is the way to suppress AI title generation entirely; that mode uses the prompt verbatim as the title and never spends tokens.

## For agents

### Adding a new model

To extend the catalog with a new model:

1. Append an entry to `MODEL_CATALOG` in [/src/shared/model-catalog.ts](/src/shared/model-catalog.ts) with `modelId`, friendly `label`, short `description`, and `provider`. Set `isReasoning: true` if it splits its token budget on chain-of-thought. Mark exactly one entry per provider with `recommended: true` (the per-provider default).
2. Add a matching pricing row to `MODEL_PRICING` in [/src/shared/model-pricing.ts](/src/shared/model-pricing.ts) — the single source of truth for per-model prices (input + output in USD per million tokens; `api-cost-tracker` re-exports it, and every consumer reads it via `getModelPrice` / `getModelPricePerM`). NEVER hardcode a price literal anywhere else — the single-source guard [/tests/unit/lint/model-price-single-source.test.ts](/tests/unit/lint/model-price-single-source.test.ts) fails CI if you do. The catalog-vs-pricing parity test in [/tests/unit/services/api-cost-tracker.test.ts](/tests/unit/services/api-cost-tracker.test.ts) fails CI if you skip this step (otherwise the cost ledger silently logs the wrong figure for that model).
3. Run `node scripts/vitest.js run tests/unit/shared/model-catalog.test.ts tests/unit/services/api-cost-tracker.test.ts` and fix until green.

The model immediately appears in every picker surface (Voice, Catch-Up Card, Email Summarizer, Quick Replies) — no per-surface wiring needed.

## Related

- [voice-and-tts.md](voice-and-tts.md) — TTS and voice intent both route through this provider; the Groq API key is configured there
- [daily-digest.md](daily-digest.md) — Omni briefings route through this provider
- [use-quick-responses.md](use-quick-responses.md) — AI suggestion chips (Alt+1/2/3) route through this provider
- [email-inbound-prescreen.md](email-inbound-prescreen.md) — explicit non-user of this setting (pinned to Gemini 2.5 Flash-Lite via OpenRouter, Haiku 4.5 fallback)
- [automations-and-auto-replies.md](automations-and-auto-replies.md) — AI conditions are another explicit non-user (always Haiku)
