---
title: Semantic Search (AI-powered meaning-based search)
---

# Semantic Search (AI-powered meaning-based search)

## What it is

**Semantic Search** lets Global Search find sessions by _meaning_, not just keywords — so a query like `"why is the deploy slow"` can surface a session titled _"Fixing CI pipeline latency"_ even though none of the query words appear in that title. It's powered by a local 33 MB embedding model (`all-MiniLM-L6-v2` via Hugging Face Transformers.js) that Omniscio downloads the first time you enable the setting, then runs entirely on your machine — no data leaves your computer. Once built, each session's title plus first message become a 384-dimensional vector stored in SQLite, and every search query gets embedded the same way and compared against the index. Semantic hits are blended into the top of your Ctrl+K results invisibly — there's no visible "AI" badge, no semantic-vs-keyword toggle, no "Top N matches" wording. Omniscio just returns the most-likely results in order. It ships **on by default**.

## Where to find it

### How to use it

1. **Confirm it's on.** Open Settings → **Sessions** → scroll to **AI-Powered Search** and check the toggle. If it's off and you flip it on, Omniscio downloads the model (~33 MB, one-time) and backfills embeddings for all your existing sessions in batches of 50. The backfill is silent — you can keep working during it.
2. **Search normally.** Press **Ctrl+K** and type a meaning-based query: `"bug in auth flow"`, `"help me outline a blog post"`, `"that Stripe thing last week"`. You don't need special syntax — semantic matching runs automatically when the eligibility criteria are met (default Ctrl+K state matches them).
3. **No UI to flip.** As of late April 2026, the previous **Semantic / Local** toggle in the Ctrl+K palette is gone — semantic ranking blends silently with keyword results. The only user-visible control for semantic search is the master setting in Settings → Sessions.
4. **Want literal-only matching for one query?** Use FTS5 syntax to short-circuit semantic ranking: wrap the term in quotes for an exact phrase (`"refund flow"`), prefix with `-` to exclude (`refund -test`), or use `OR` to combine literals (`staging OR production`). Semantic still _runs_, but FTS5 keyword hits dominate when your query is precise. There's no in-UI escape hatch beyond the global setting.
5. **When it won't fire.** Semantic search only runs on _clean_ queries: sorted by **Most Relevant**, no source/date filters, `searchIn = all`, tool actions excluded, channel type `conversation`, and first page. The Ctrl+K palette's defaults match the eligibility criteria, so most searches work without thought; if you add filters or flip to **Newest/Oldest**, you get pure keyword results.
6. **Turn it off if you want.** Back in Settings → Sessions, toggle **AI-Powered Search** off. Existing embeddings stay on disk (harmless) but no new ones will be created and searches will be keyword-only. No indexing happens, no model loads, no CPU spent.

## How it behaves

### If it repeatedly crashes on your machine (automatic safety)

The embedding model runs through a native engine (onnxruntime) that, on a small number of machines, can crash hard — historically an intermittent crash right as a new session starts indexing its first message (see [crash-recovery.md](crash-recovery.md) and [slow-computer.md](slow-computer.md)). The engine now runs in its **own separate process**, isolated from the app: if it crashes, only that helper process dies — Omniscio notices, falls back to keyword-only search, and the app itself stays up. Your sessions and data are never affected — only whether meaning-based ranking is active. As a further backstop, Omniscio also **self-heals**: if it recovers from **3 of these embedding crashes within two weeks**, it automatically turns Semantic Search off on the next launch and raises a dismissible inbox notice with a one-click **Open Semantic Search settings** button. Turning it back on in Settings → Sessions clears the counter and gives it a fresh start.

Under the hood there are two layers. The **primary** one is process isolation: the embedding engine runs in an Electron `utilityProcess` child (the "ML service"), so a hard native abort — the `Cannot create a handle without a HandleScope` crash, CID 9bf615c2 — kills only that child; the host degrades to keyword search and can respawn it. The **backstop** is the semantic-search **circuit breaker**: for an onnx abort that still reaches the main process AND is attributed to embedding, the next-launch crash-sweep records it in a rolling 14-day window in `gpu-stability.json`, and early bootstrap's `applySemanticSearchCircuitBreaker` disables `semanticSearchEnabled` past the threshold — before the engine can spawn again. **Attribution matters now that more engines are isolated:** embedding, local speech-to-text, and screen-recorder captions each run in their own child, so a main-process onnx abort can only come from the one engine still running onnx inline — wake-word. The crash-sweep therefore reads a persisted *last-live-main-onnx-host* breadcrumb (`onnx-host-breadcrumb.ts`, written by the wake-word engine) and trips this breaker ONLY when the abort was embedding's — a wake-word crash never disables semantic search (the pre-breadcrumb misfire). It is a **confirmed-cause** trigger (safe, unlike the GPU auto-disable, because the keyword-search fallback has no crash vector of its own); with isolation shipped it is now a near-vestigial last resort. Full envelope: [ml-engine-reliability-contract.md](/.claude/memory/contracts/ml-engine-reliability-contract.md); ORT teardown mechanics: [embedding-worker-offload-contract.md](/.claude/memory/contracts/embedding-worker-offload-contract.md).

## For agents

### How it works

The toggle lives at [/src/renderer/src/features/settings/sections/session/SessionSettings.tsx](/src/renderer/src/features/settings/sections/session/SessionSettings.tsx) bound to `AppSettings.semanticSearchEnabled` (default `true` after the one-time migration in [/src/main/services/config-store/migrations/migrate-semantic-search-default.ts](/src/main/services/config-store/migrations/migrate-semantic-search-default.ts)). When enabled at startup, [/src/main/index.ts](/src/main/index.ts) imports [/src/main/services/embedding-model.ts](/src/main/services/embedding-model.ts) and calls `initEmbeddingModel()` → `backfillSessions()` from [/src/main/services/embedding-service.ts](/src/main/services/embedding-service.ts). The model is `Xenova/all-MiniLM-L6-v2` pulled via `@huggingface/transformers`; the 33 MB ONNX weights cache to the userData `models/` directory. Each session's source text = `sessionName + ' — ' + firstOperatorMessage`; the 384-dim float vector is stored as a 1536-byte BLOB in the `session_embeddings` table (migration v74, [incremental-migrations.ts](/src/main/db/incremental-migrations.ts)). Indexing is triggered two ways: the startup backfill, and on every `SESSION_RENAMED` event from [/src/main/ipc/title-orchestrator.ts](/src/main/ipc/title-orchestrator.ts) so new sessions get embedded as soon as they're auto-titled. At search time, the `IPC.SEARCH_ALL` handler in [/src/main/ipc/session-handlers.ts](/src/main/ipc/session-handlers.ts) gates semantic search behind 9 eligibility checks, then calls `semanticSearch(query, 10)` which cosine-compares against all stored vectors, drops anything below `SIMILARITY_THRESHOLD = 0.5`, dedupes against keyword hits, and `unshift`-inserts the survivors at the top of the response. The renderer no longer marks those rows specially — the previous AI badge and "Semantic match (NN% similar)" snippet wording were removed when the Ctrl+K UI was redesigned, so semantic and keyword matches now look identical to the user. Progress is still pushed on `IPC.SEMANTIC_SEARCH_PROGRESS` (no UI widget consumes it yet). The CLI control server exposes a read-only search surface via [/src/main/services/cli/cli-server-search.ts](/src/main/services/cli/cli-server-search.ts): `/search/semantic`, `/search/fts`, `/sessions`, and `/session/<id>/messages?all=true`, all supporting `startedAfter` / `startedBefore` timeframe filters (ISO 8601 or `7d`/`24h`/`90m` offsets). These endpoints power the [`search-sessions` AI skill](/.claude/skills/search-sessions/SKILL.md) — agents use an escalation ladder (semantic → FTS → timeframe list → full transcripts) to locate prior sessions by natural-language description.

## Related

- [global-search.md](global-search.md) — keyword-side of the same Ctrl+K palette; semantic results _augment_ keyword results, they don't replace them, and they're no longer visually distinguished
- [mempalace-memory.md](mempalace-memory.md) — MemPalace drawer search uses its own semantic pipeline (different embeddings, different surface)
