QW Triage Tool (mine the live DB for parser bugs)
The read-only developer tool that scans every agent message in your live Omniscio database, scores each one for likely QuestionWidget parser bugs, and writes a sorted JSONL review queue a human triages — the flags it takes, how the scores are read, and the re-baseline rule.
What it is
npm run triage:qw is a developer tool that scans every source='agent' message in your live Omniscio SQLite database, scores each one for likely QuestionWidget (QW) parser bugs, and writes a sorted JSONL review queue under tests/fixtures/question-widget/_triage/<UTC-timestamp>/. Two failure modes the tool surfaces:
- Silent miss — the parser returned
null(no widget rendered) but the content has strong QW signals (bold header, lettered options, question mark). The user sees plain text where they should have seen pills. - Over-fire — the parser DID render a widget but the content is log-line spam, raw code blocks, or other prose that shouldn't have triggered. The user sees a pill widget on a paragraph that should have been plain markdown.
The tool is read-only against the database, purely rule-based (no AI, no model calls, $0), and idempotent — re-running without code changes produces the same JSONL because it's a pure function of (DB content, parser SHA in the baseline, scorer regex tables). It writes nothing back to Omniscio; output stays under _triage/.
This is not a regression test — it's a mining tool that surfaces candidates a human reviews and (optionally) promotes to corpus fixtures. The QW corpus regression suite (npm run test:qw-corpus) is a separate gate that runs on every parser change; this tool is what you run when you want to expand the corpus or audit a long-tail of real production messages for bugs you missed.
See qw-parser-snapshots.md for the related frozen-snapshot diff workflow.
Where to find it
This is a developer tool with no product surface — there is no screen, menu or setting for it. You run it from a shell with npm run triage:qw, and its output lands in a dated folder beside the other QuestionWidget test fixtures rather than anywhere in the running app.
How it behaves
How to run
npm run triage:qw # full scan, emit JSONL
npm run triage:qw -- --dry-run # scan + print summary, no file write
npm run triage:qw -- --limit 100 # write only top 100 by score
npm run triage:qw -- --min-score 50 # raise the floor (default 30)
npm run triage:qw -- --dry-run --limit 50 # combine flags
Flags supported by scripts/qw-triage.mjs:
| Flag | Default | Meaning |
|---|---|---|
--limit N |
∞ | Keep only top N candidates by score after sort (descending). |
--dry-run |
off | Print summary + top 10 to stdout. Skips writing triage-report.jsonl. |
--min-score |
30 | Drop candidates scoring below this. Lower = noisier queue; higher = strictest. |
The runner scripts/electron-node-launch.cjs qw-triage.mjs spawns qw-triage.mjs under Electron-as-Node (ELECTRON_RUN_AS_NODE=1 + electron.exe) + tsx, because the wrapper dynamic-imports parseContentSegments from the renderer's TypeScript source via tsx, so any signals derived from the agent-markdown parser stay in sync with what the live UI sees — and that preload is an Electron-as-Node runtime.
Correction, 2026-09-30. This paragraph used to justify the wrapper with
better-sqlite3: "a native module compiled against Electron's V8 ABI (NODE_MODULE_VERSION 140); plain Node 22/24 (NMV 137) crashes withERR_DLOPEN_FAILED". That was true of the node-gyp era and is false of the version this repo pins (13.x), which is a Node-API add-on whose in-package prebuild is loaded before any node-gyp output — so no ABI is compared and plainnodeloads it. The wrapper is unchanged; only its stated reason was wrong.
The database path defaults to:
%APPDATA%\omniscio\mission-control.db # Windows
~/Library/Application Support/omniscio/... # macOS (via APPDATA shim)
Override via AMC_DB_PATH=/some/other.db if you want to triage a sandbox / e2e instance instead of your live DB.
The tool prints its DB path + baseline message count + fixture count BEFORE scanning so you can Ctrl-C if the wrong instance is targeted:
[qw-triage] DB=C:\Users\...\mission-control.db (18271 baseline msgs, 192 existing fixtures)
Where the output lands
tests/fixtures/question-widget/_triage/<UTC-timestamp>/
triage-report.jsonl
triage-summary.json
<UTC-timestamp> is the script's new Date().toISOString() with : and . replaced by - (e.g. 2026-05-09T00-10-56-953Z). Each batch lands in its own directory so concurrent runs don't clobber each other.
The whole _triage/ directory is gitignored except for .gitkeep. Triage batches are regeneratable from the DB + parser SHA, so committing a 4 MB JSONL would just be churn — it goes stale on every parser change.
To clean up old batches, just rm -rf tests/fixtures/question-widget/_triage/<batch>.
triage-report.jsonl — one candidate per line
Each line is a single JSON object, written sorted by descending score:
{
"messageId": "msg_abc123",
"timestamp": "2026-05-08T14:23:01.547Z",
"parserDecision": "null",
"parserBranches": [],
"segmentIndex": 0,
"signals": {
"boldHeader": true,
"letteredOptions": true,
"numberedOptions": false,
"bulletOptions": false,
"questionMark": true,
"choiceImperative": false,
"codeFence": false,
"logLine": false,
"optionCount": 4,
"scanTruncated": false,
"contentLength": 1342
},
"score": 90,
"suspicionReason": "silent-miss-strong",
"contentExcerpt": "**Should we proceed?**\n\nA. Yes, ship it now\nB. ...(first 500 chars)..."
}
Field guide:
parserDecision— what the parser did with this message according to the frozen baseline attests/fixtures/question-widget/_backtest/baseline.json. Values:'null'(no widget rendered),'preprocess-empty'(the parser preprocessor stripped everything),'renders'(a widget WAS rendered),'unknown'(no baseline entry — message arrived after the last re-baseline), or'fetch-error'(the parser threw — score forced to 100, see content excerpt for the exception text).parserBranches— for'renders'rows, the named cascade branches that fired (e.g.['WRAP_BOLD_COLON']). Empty for'null'and'preprocess-empty'.segmentIndex— agent messages are split into prose / tool segments before parsing; this is the index of the segment the scorer picked as the most suspicious. Multi-segment messages have a winner.signals— boolean / count flags extracted by the rule-based regex tables inscripts/qw-triage/signals.mjs. Drive the score directly.scanTruncated: truemeans the message was longer than 50,000 bytes and only the prefix was scanned.score— 0-100, computed byscripts/qw-triage/score.mjs. The bucket boundaries are documented below.suspicionReason— short stable label describing why the candidate ranked. One ofsilent-miss-strong,silent-miss-weak,preprocess-stripped-question,overfire-log-line, or the literal stringscore=<n>if no named pattern fits.contentExcerpt— first 500 characters of the winning segment's prose. Enough to triage manually without opening the full message; truncated for JSONL row size.
triage-summary.json — companion summary
{
"parserSha": "da2af34ae",
"triageScriptSha": null,
"batchTimestamp": "2026-05-09T00:10:56.953Z",
"dbPath": "C:\\Users\\...\\mission-control.db",
"counts": {
"scanned": 21738,
"skippedCovered": 81,
"skippedLowScore": 17530,
"emitted": 4127
},
"scoreBuckets": {
"90-100": 56,
"70-89": 13,
"50-69": 2412,
"30-49": 1646
}
}
parserSha— short hash of the QW parser source at the time the baseline was captured. If the parser has changed since, the JSONL'sparserDecisionvalues are stale and you should re-baseline (see "Re-baseline contract" below).counts.skippedCovered— message IDs already represented in the fixture corpus. Dedupe runs by reading everyexpected.jsonundertests/fixtures/question-widget/should-render/andshould-not-render/and skipping any message whoseidappears in the union.counts.skippedLowScore— messages where the winning segment scored below--min-score.counts.emitted— how many made it to the JSONL. Equalsmin(--limit, count above floor).scoreBuckets— distribution across the four scorer tiers, useful at-a-glance for whether the parser regressed or improved between batches.
How to interpret scores
Score = sum of weighted signals, clamped to [0, 100]. The weight table lives in scripts/qw-triage/score.mjs and depends on the parser's decision for that message:
parserDecision === 'null'(silent-miss bucket): bold header (+30), lettered options (+30), question mark (+15), choice imperative ("which" / "should we" / "want me to" / etc., +10), 2+ options (+10), 4+ options (+5).parserDecision === 'preprocess-empty'(parser preprocessor ate the question): any one signal (+50), bold + lettered together (+20).parserDecision === 'renders'(over-fire bucket): log-line content (+70), code fence without bold header (+30).
Negative pressure: logLine signal in a non-'renders' row subtracts 50 (it's evidence the candidate is log spam, not a real question). Truncated scans are capped at 40 (we didn't see the whole message, so we won't claim high confidence).
Score buckets
| Score range | Meaning |
|---|---|
| 90-100 | Strong silent-miss or preprocess-stripped — bold header AND lettered options AND question mark, all in a 'null' parser decision. Triage these first; the highest-confidence bugs live here. |
| 70-89 | Strong silent-miss without all three primary signals, OR a strong over-fire (log-line content with 'renders'). |
| 50-69 | Mid-confidence. Bulk of the queue. Many real bugs here, but also many borderline cases — choice imperative without lettered options, etc. Skim, don't chase every one. |
| 30-49 | Weak signal. Default floor. Most are noise; spot-check a handful. Lower the floor with --min-score 0 only if you're explicitly hunting long-tail edge cases. |
The scorer is monotonic — adding signals never decreases the score. There's a unit-test invariant (tests/unit/scripts/qw-triage-score.test.ts) that pins this: score("What?") <= score("What?\nA. yes") <= score("**Q:** What?\nA. yes\nB. no"). If you add a new signal that reverses this, the test fails. Don't bypass it — monotonicity is what makes the score buckets meaningful as triage tiers.
suspicionReason labels
silent-miss-strong— parser said'null', content has bold header AND lettered options. Highest-confidence bug in the queue.silent-miss-weak— parser said'null', content has lettered options but no bold header. Could be a real question phrased without bold emphasis, or could be a legitimate "Files changed: A. one B. two" listing.preprocess-stripped-question— parser said'preprocess-empty'. The content has signals but the parser's preprocessor (which strips tool output, code blocks, etc.) removed everything before the regex cascade ran.overfire-log-line— parser said'renders'and the content has a[2026-05-08]or[ERROR]log-line pattern. The widget is firing on log spam.score=<n>— fallback label when no named pattern matches. The numeric score still sorts correctly but you'll need to read the signals manually.
How to promote a candidate to a fixture
Manual flow for v1:
- Open
triage-report.jsonland scroll through the top of the queue (it's pre-sorted by score). Use any text editor that handles 4 MB files — VS Code, Notepad++,lessin WSL. - For each candidate worth investigating: copy the
messageIdand look it up in your Omniscio database, OR open the source session in Omniscio and find the message. ThecontentExcerptfield gives you 500 chars to triage without leaving the JSONL. - Decide which side of the corpus the fixture goes in:
- Real silent-miss → fixture under
tests/fixtures/question-widget/should-render/<descriptive-slug>/. - Real over-fire → fixture under
tests/fixtures/question-widget/should-not-render/<descriptive-slug>/. - Borderline / debatable → leave it. The queue is allowed to have noise.
- Real silent-miss → fixture under
- Author the fixture per the schema in
tests/fixtures/question-widget/README.md—content.md(the raw agent message text) +expected.json({ shouldRender, source: { messageId, sessionId, timestamp }, notes }). - Run the corpus suite to verify the fixture exercises the parser the way you expect:
Red on a should-render fixture means the parser is missing it (file the bug + write the failing assertion first per the Omniscionpm run test:qw-corpusQW parsing bugsrule); green means you've already covered it.
There is no automated promote-to-fixture command in v1. The fixture format is small enough that copy-paste is fine for the volume the triage tool produces, and a forced shape would prevent you from adding the human-readable notes field that makes the corpus useful as a regression document.
Decision gate — do we add AI re-rank?
The original plan (docs/plans/2026-05-08-qw-triage-tool-plan.md) had a conditional Phase 4 to add Anthropic Haiku re-ranking on top of the rule-based scorer. The gate to enable it was: after the first full run, spot-check the top 50 candidates. If ≥30% are real bugs (silent-miss or over-fire), ship rule-based-only and skip AI. If <30%, the rule-based signal is too noisy and AI is worth the cost.
Result of the 2026-05-08 first-run gate: ≥30% real bugs in top 50 → rule-based shipped, AI skipped. The current scorer is the production tool.
If a future parser change degrades the signal (e.g. a new branch lands that fires on log-lines, polluting the silent-miss bucket), re-run the spot-check. If the rate drops below 30%, the conditional Phase 4 wiring is described in the plan — it would add a cache-backed Haiku call against the top 20 candidates with --confirm-cost gating and a hard --max-cost 1.00 cap.
Re-baseline contract
The triage tool's parserDecision field comes from the frozen baseline at tests/fixtures/question-widget/_backtest/baseline.json. That baseline is captured by npm run backtest:qw against the parser at the time of capture, and stamped with the parser SHA.
You must re-baseline whenever the QW parser changes, otherwise:
- A real silent-miss the new parser would now catch still shows up in the queue (because the baseline says it returned
null). - A real over-fire the new parser no longer fires on still shows up (because the baseline says it returned
'renders').
Re-baseline command:
npm run backtest:qw -- --update-baseline
This rewrites _backtest/baseline.json with the current parser's decisions across every message in the database, and updates the embedded parserSha field. Commit the new baseline alongside the parser change so subsequent triage runs use the fresh decisions.
The triage-summary.json records parserSha per batch, so you can always look at an old batch and check whether it's stale relative to today's parser by comparing SHAs.
Concurrency and safety
- DB access is read-only:
better-sqlite3opens with{ readonly: true, fileMustExist: true }. The triage tool cannot mutate the live Omniscio database. - Atomic file writes: both
triage-report.jsonlandtriage-summary.jsonwrite to a.stagingsibling first, thenrenameSyncinto place. A killed run mid-write never corrupts a previous batch. - Concurrent runs: each batch directory is named after
new Date().toISOString(). Twotriage:qwruns would only collide on directory name if they fire in the same millisecond (the timestamp includes milliseconds) — accept the low probability. - Omniscio must not be holding a write lock: the tool opens the DB read-only, but if Omniscio is mid-WAL-checkpoint, the open may briefly block. Run with Omniscio closed if you want maximum safety; run with Omniscio open if you want a current snapshot.
For agents
Implementation pointers (for agents touching this code)
- Entry script:
scripts/qw-triage.mjs— reads DB, loops messages, calls scorer per segment, sorts and writes JSONL. - Native-binding wrapper:
scripts/electron-node-launch.cjs qw-triage.mjs—ELECTRON_RUN_AS_NODE=1+electron.exe+tsxsobetter-sqlite3and the renderer's TS source both load. - Signals:
scripts/qw-triage/signals.mjs— the rule-based regex tables (bold header, lettered options, log line, etc.). - Scorer:
scripts/qw-triage/score.mjs— weighted signal sum, monotonic, clamped 0-100. - Segment picker:
scripts/qw-triage/pick-segment.mjs— picks the highest-scoring prose segment when a message has multiple. - Dedupe:
scripts/qw-triage/dedupe.mjs— reads everyexpected.jsonunder the corpus root and returns aSet<messageId>of already-covered IDs. - Emit:
scripts/qw-triage/emit.mjs— atomic JSONL + summary writer. - Validation gate (regression test):
tests/integration/qw-triage-validation-gate.test.ts— fails CI if the scorer wrongly flags more than 5 should-render fixtures as silent-miss-strong, or if it would surface ANY should-not-render fixture as a false-positive over-fire candidate (when content has neither a log-line nor a code-fence). This is what stops "improvements" to the scorer from regressing precision against the curated corpus. - Score unit tests:
tests/unit/scripts/qw-triage-score.test.ts— pin the per-bucket scoring contract (silent-miss-strong ≥ 80, log-line over-fire ≥ 60, rhetorical questions < 30, monotonicity).
References
tests/fixtures/question-widget/README.md— fixture authoring schema for promoted candidates.docs/plans/2026-05-08-qw-triage-tool-plan.md— the build plan, including the conditional AI re-rank phase.
Related
question-widget.md covers what the widget is, the parser versions Omniscio ships, and the plumbing behind them. The other half of the parser-regression story is qw-parser-snapshots.md, the locked-snapshot diff workflow — a different tool with a different purpose, run offline against a frozen corpus rather than against your live database. And qw-miss-reporting.md is the in-app report a user files when the parser misses a question they expected to see as a widget, which is how a shape neither of these tools could predict reaches the parser team.
Last verified 2026-10-06