Omniscio documentation
Browse all documentation
  1. Getting Started13
  2. Sessions & Agents115
  3. Inbox & Notifications59
  4. Projects & Tasks95
  5. Automation & Scheduling75
  6. Knowledge & Memory26
  7. AI Features60
  8. Integrations100
  9. Plugins & Marketplace33
  10. Cloud & Teams56
  11. Settings & Customization58
  12. Account & Billing28
  13. Troubleshooting84
  14. CLI & API Reference22
  15. Legal & Policies4
  16. Uncategorised22

QW Triage Tool (mine the live DB for parser bugs)

The read-only developer tool that scans every agent message in your live Omniscio database, scores each one for likely QuestionWidget parser bugs, and writes a sorted JSONL review queue a human triages — the flags it takes, how the scores are read, and the re-baseline rule.

What it is

npm run triage:qw is a developer tool that scans every source='agent' message in your live Omniscio SQLite database, scores each one for likely QuestionWidget (QW) parser bugs, and writes a sorted JSONL review queue under tests/fixtures/question-widget/_triage/<UTC-timestamp>/. Two failure modes the tool surfaces:

  1. Silent miss — the parser returned null (no widget rendered) but the content has strong QW signals (bold header, lettered options, question mark). The user sees plain text where they should have seen pills.
  2. Over-fire — the parser DID render a widget but the content is log-line spam, raw code blocks, or other prose that shouldn't have triggered. The user sees a pill widget on a paragraph that should have been plain markdown.

The tool is read-only against the database, purely rule-based (no AI, no model calls, $0), and idempotent — re-running without code changes produces the same JSONL because it's a pure function of (DB content, parser SHA in the baseline, scorer regex tables). It writes nothing back to Omniscio; output stays under _triage/.

This is not a regression test — it's a mining tool that surfaces candidates a human reviews and (optionally) promotes to corpus fixtures. The QW corpus regression suite (npm run test:qw-corpus) is a separate gate that runs on every parser change; this tool is what you run when you want to expand the corpus or audit a long-tail of real production messages for bugs you missed.

See qw-parser-snapshots.md for the related frozen-snapshot diff workflow.

Where to find it

This is a developer tool with no product surface — there is no screen, menu or setting for it. You run it from a shell with npm run triage:qw, and its output lands in a dated folder beside the other QuestionWidget test fixtures rather than anywhere in the running app.

How it behaves

How to run

npm run triage:qw                          # full scan, emit JSONL
npm run triage:qw -- --dry-run             # scan + print summary, no file write
npm run triage:qw -- --limit 100           # write only top 100 by score
npm run triage:qw -- --min-score 50        # raise the floor (default 30)
npm run triage:qw -- --dry-run --limit 50  # combine flags

Flags supported by scripts/qw-triage.mjs:

Flag Default Meaning
--limit N ∞ Keep only top N candidates by score after sort (descending).
--dry-run off Print summary + top 10 to stdout. Skips writing triage-report.jsonl.
--min-score 30 Drop candidates scoring below this. Lower = noisier queue; higher = strictest.

The runner scripts/electron-node-launch.cjs qw-triage.mjs spawns qw-triage.mjs under Electron-as-Node (ELECTRON_RUN_AS_NODE=1 + electron.exe) + tsx, because the wrapper dynamic-imports parseContentSegments from the renderer's TypeScript source via tsx, so any signals derived from the agent-markdown parser stay in sync with what the live UI sees — and that preload is an Electron-as-Node runtime.

Correction, 2026-09-30. This paragraph used to justify the wrapper with better-sqlite3: "a native module compiled against Electron's V8 ABI (NODE_MODULE_VERSION 140); plain Node 22/24 (NMV 137) crashes with ERR_DLOPEN_FAILED". That was true of the node-gyp era and is false of the version this repo pins (13.x), which is a Node-API add-on whose in-package prebuild is loaded before any node-gyp output — so no ABI is compared and plain node loads it. The wrapper is unchanged; only its stated reason was wrong.

The database path defaults to:

%APPDATA%\omniscio\mission-control.db        # Windows
~/Library/Application Support/omniscio/...   # macOS (via APPDATA shim)

Override via AMC_DB_PATH=/some/other.db if you want to triage a sandbox / e2e instance instead of your live DB.

The tool prints its DB path + baseline message count + fixture count BEFORE scanning so you can Ctrl-C if the wrong instance is targeted:

[qw-triage] DB=C:\Users\...\mission-control.db (18271 baseline msgs, 192 existing fixtures)

Where the output lands

tests/fixtures/question-widget/_triage/<UTC-timestamp>/
  triage-report.jsonl
  triage-summary.json

<UTC-timestamp> is the script's new Date().toISOString() with : and . replaced by - (e.g. 2026-05-09T00-10-56-953Z). Each batch lands in its own directory so concurrent runs don't clobber each other.

The whole _triage/ directory is gitignored except for .gitkeep. Triage batches are regeneratable from the DB + parser SHA, so committing a 4 MB JSONL would just be churn — it goes stale on every parser change.

To clean up old batches, just rm -rf tests/fixtures/question-widget/_triage/<batch>.

triage-report.jsonl — one candidate per line

Each line is a single JSON object, written sorted by descending score:

{
  "messageId": "msg_abc123",
  "timestamp": "2026-05-08T14:23:01.547Z",
  "parserDecision": "null",
  "parserBranches": [],
  "segmentIndex": 0,
  "signals": {
    "boldHeader": true,
    "letteredOptions": true,
    "numberedOptions": false,
    "bulletOptions": false,
    "questionMark": true,
    "choiceImperative": false,
    "codeFence": false,
    "logLine": false,
    "optionCount": 4,
    "scanTruncated": false,
    "contentLength": 1342
  },
  "score": 90,
  "suspicionReason": "silent-miss-strong",
  "contentExcerpt": "**Should we proceed?**\n\nA. Yes, ship it now\nB. ...(first 500 chars)..."
}

Field guide:

  • parserDecision — what the parser did with this message according to the frozen baseline at tests/fixtures/question-widget/_backtest/baseline.json. Values: 'null' (no widget rendered), 'preprocess-empty' (the parser preprocessor stripped everything), 'renders' (a widget WAS rendered), 'unknown' (no baseline entry — message arrived after the last re-baseline), or 'fetch-error' (the parser threw — score forced to 100, see content excerpt for the exception text).
  • parserBranches — for 'renders' rows, the named cascade branches that fired (e.g. ['WRAP_BOLD_COLON']). Empty for 'null' and 'preprocess-empty'.
  • segmentIndex — agent messages are split into prose / tool segments before parsing; this is the index of the segment the scorer picked as the most suspicious. Multi-segment messages have a winner.
  • signals — boolean / count flags extracted by the rule-based regex tables in scripts/qw-triage/signals.mjs. Drive the score directly. scanTruncated: true means the message was longer than 50,000 bytes and only the prefix was scanned.
  • score — 0-100, computed by scripts/qw-triage/score.mjs. The bucket boundaries are documented below.
  • suspicionReason — short stable label describing why the candidate ranked. One of silent-miss-strong, silent-miss-weak, preprocess-stripped-question, overfire-log-line, or the literal string score=<n> if no named pattern fits.
  • contentExcerpt — first 500 characters of the winning segment's prose. Enough to triage manually without opening the full message; truncated for JSONL row size.

triage-summary.json — companion summary

{
  "parserSha": "da2af34ae",
  "triageScriptSha": null,
  "batchTimestamp": "2026-05-09T00:10:56.953Z",
  "dbPath": "C:\\Users\\...\\mission-control.db",
  "counts": {
    "scanned": 21738,
    "skippedCovered": 81,
    "skippedLowScore": 17530,
    "emitted": 4127
  },
  "scoreBuckets": {
    "90-100": 56,
    "70-89": 13,
    "50-69": 2412,
    "30-49": 1646
  }
}
  • parserSha — short hash of the QW parser source at the time the baseline was captured. If the parser has changed since, the JSONL's parserDecision values are stale and you should re-baseline (see "Re-baseline contract" below).
  • counts.skippedCovered — message IDs already represented in the fixture corpus. Dedupe runs by reading every expected.json under tests/fixtures/question-widget/should-render/ and should-not-render/ and skipping any message whose id appears in the union.
  • counts.skippedLowScore — messages where the winning segment scored below --min-score.
  • counts.emitted — how many made it to the JSONL. Equals min(--limit, count above floor).
  • scoreBuckets — distribution across the four scorer tiers, useful at-a-glance for whether the parser regressed or improved between batches.

How to interpret scores

Score = sum of weighted signals, clamped to [0, 100]. The weight table lives in scripts/qw-triage/score.mjs and depends on the parser's decision for that message:

  • parserDecision === 'null' (silent-miss bucket): bold header (+30), lettered options (+30), question mark (+15), choice imperative ("which" / "should we" / "want me to" / etc., +10), 2+ options (+10), 4+ options (+5).
  • parserDecision === 'preprocess-empty' (parser preprocessor ate the question): any one signal (+50), bold + lettered together (+20).
  • parserDecision === 'renders' (over-fire bucket): log-line content (+70), code fence without bold header (+30).

Negative pressure: logLine signal in a non-'renders' row subtracts 50 (it's evidence the candidate is log spam, not a real question). Truncated scans are capped at 40 (we didn't see the whole message, so we won't claim high confidence).

Score buckets

Score range Meaning
90-100 Strong silent-miss or preprocess-stripped — bold header AND lettered options AND question mark, all in a 'null' parser decision. Triage these first; the highest-confidence bugs live here.
70-89 Strong silent-miss without all three primary signals, OR a strong over-fire (log-line content with 'renders').
50-69 Mid-confidence. Bulk of the queue. Many real bugs here, but also many borderline cases — choice imperative without lettered options, etc. Skim, don't chase every one.
30-49 Weak signal. Default floor. Most are noise; spot-check a handful. Lower the floor with --min-score 0 only if you're explicitly hunting long-tail edge cases.

The scorer is monotonic — adding signals never decreases the score. There's a unit-test invariant (tests/unit/scripts/qw-triage-score.test.ts) that pins this: score("What?") <= score("What?\nA. yes") <= score("**Q:** What?\nA. yes\nB. no"). If you add a new signal that reverses this, the test fails. Don't bypass it — monotonicity is what makes the score buckets meaningful as triage tiers.

suspicionReason labels

  • silent-miss-strong — parser said 'null', content has bold header AND lettered options. Highest-confidence bug in the queue.
  • silent-miss-weak — parser said 'null', content has lettered options but no bold header. Could be a real question phrased without bold emphasis, or could be a legitimate "Files changed: A. one B. two" listing.
  • preprocess-stripped-question — parser said 'preprocess-empty'. The content has signals but the parser's preprocessor (which strips tool output, code blocks, etc.) removed everything before the regex cascade ran.
  • overfire-log-line — parser said 'renders' and the content has a [2026-05-08] or [ERROR] log-line pattern. The widget is firing on log spam.
  • score=<n> — fallback label when no named pattern matches. The numeric score still sorts correctly but you'll need to read the signals manually.

How to promote a candidate to a fixture

Manual flow for v1:

  1. Open triage-report.jsonl and scroll through the top of the queue (it's pre-sorted by score). Use any text editor that handles 4 MB files — VS Code, Notepad++, less in WSL.
  2. For each candidate worth investigating: copy the messageId and look it up in your Omniscio database, OR open the source session in Omniscio and find the message. The contentExcerpt field gives you 500 chars to triage without leaving the JSONL.
  3. Decide which side of the corpus the fixture goes in:
    • Real silent-miss → fixture under tests/fixtures/question-widget/should-render/<descriptive-slug>/.
    • Real over-fire → fixture under tests/fixtures/question-widget/should-not-render/<descriptive-slug>/.
    • Borderline / debatable → leave it. The queue is allowed to have noise.
  4. Author the fixture per the schema in tests/fixtures/question-widget/README.md — content.md (the raw agent message text) + expected.json ({ shouldRender, source: { messageId, sessionId, timestamp }, notes }).
  5. Run the corpus suite to verify the fixture exercises the parser the way you expect:
    npm run test:qw-corpus
    
    Red on a should-render fixture means the parser is missing it (file the bug + write the failing assertion first per the Omniscio QW parsing bugs rule); green means you've already covered it.

There is no automated promote-to-fixture command in v1. The fixture format is small enough that copy-paste is fine for the volume the triage tool produces, and a forced shape would prevent you from adding the human-readable notes field that makes the corpus useful as a regression document.

Decision gate — do we add AI re-rank?

The original plan (docs/plans/2026-05-08-qw-triage-tool-plan.md) had a conditional Phase 4 to add Anthropic Haiku re-ranking on top of the rule-based scorer. The gate to enable it was: after the first full run, spot-check the top 50 candidates. If ≥30% are real bugs (silent-miss or over-fire), ship rule-based-only and skip AI. If <30%, the rule-based signal is too noisy and AI is worth the cost.

Result of the 2026-05-08 first-run gate: ≥30% real bugs in top 50 → rule-based shipped, AI skipped. The current scorer is the production tool.

If a future parser change degrades the signal (e.g. a new branch lands that fires on log-lines, polluting the silent-miss bucket), re-run the spot-check. If the rate drops below 30%, the conditional Phase 4 wiring is described in the plan — it would add a cache-backed Haiku call against the top 20 candidates with --confirm-cost gating and a hard --max-cost 1.00 cap.

Re-baseline contract

The triage tool's parserDecision field comes from the frozen baseline at tests/fixtures/question-widget/_backtest/baseline.json. That baseline is captured by npm run backtest:qw against the parser at the time of capture, and stamped with the parser SHA.

You must re-baseline whenever the QW parser changes, otherwise:

  • A real silent-miss the new parser would now catch still shows up in the queue (because the baseline says it returned null).
  • A real over-fire the new parser no longer fires on still shows up (because the baseline says it returned 'renders').

Re-baseline command:

npm run backtest:qw -- --update-baseline

This rewrites _backtest/baseline.json with the current parser's decisions across every message in the database, and updates the embedded parserSha field. Commit the new baseline alongside the parser change so subsequent triage runs use the fresh decisions.

The triage-summary.json records parserSha per batch, so you can always look at an old batch and check whether it's stale relative to today's parser by comparing SHAs.

Concurrency and safety

  • DB access is read-only: better-sqlite3 opens with { readonly: true, fileMustExist: true }. The triage tool cannot mutate the live Omniscio database.
  • Atomic file writes: both triage-report.jsonl and triage-summary.json write to a .staging sibling first, then renameSync into place. A killed run mid-write never corrupts a previous batch.
  • Concurrent runs: each batch directory is named after new Date().toISOString(). Two triage:qw runs would only collide on directory name if they fire in the same millisecond (the timestamp includes milliseconds) — accept the low probability.
  • Omniscio must not be holding a write lock: the tool opens the DB read-only, but if Omniscio is mid-WAL-checkpoint, the open may briefly block. Run with Omniscio closed if you want maximum safety; run with Omniscio open if you want a current snapshot.

For agents

Implementation pointers (for agents touching this code)

  • Entry script: scripts/qw-triage.mjs — reads DB, loops messages, calls scorer per segment, sorts and writes JSONL.
  • Native-binding wrapper: scripts/electron-node-launch.cjs qw-triage.mjs — ELECTRON_RUN_AS_NODE=1 + electron.exe + tsx so better-sqlite3 and the renderer's TS source both load.
  • Signals: scripts/qw-triage/signals.mjs — the rule-based regex tables (bold header, lettered options, log line, etc.).
  • Scorer: scripts/qw-triage/score.mjs — weighted signal sum, monotonic, clamped 0-100.
  • Segment picker: scripts/qw-triage/pick-segment.mjs — picks the highest-scoring prose segment when a message has multiple.
  • Dedupe: scripts/qw-triage/dedupe.mjs — reads every expected.json under the corpus root and returns a Set<messageId> of already-covered IDs.
  • Emit: scripts/qw-triage/emit.mjs — atomic JSONL + summary writer.
  • Validation gate (regression test): tests/integration/qw-triage-validation-gate.test.ts — fails CI if the scorer wrongly flags more than 5 should-render fixtures as silent-miss-strong, or if it would surface ANY should-not-render fixture as a false-positive over-fire candidate (when content has neither a log-line nor a code-fence). This is what stops "improvements" to the scorer from regressing precision against the curated corpus.
  • Score unit tests: tests/unit/scripts/qw-triage-score.test.ts — pin the per-bucket scoring contract (silent-miss-strong ≥ 80, log-line over-fire ≥ 60, rhetorical questions < 30, monotonicity).

References

  • tests/fixtures/question-widget/README.md — fixture authoring schema for promoted candidates.
  • docs/plans/2026-05-08-qw-triage-tool-plan.md — the build plan, including the conditional AI re-rank phase.

Related

question-widget.md covers what the widget is, the parser versions Omniscio ships, and the plumbing behind them. The other half of the parser-regression story is qw-parser-snapshots.md, the locked-snapshot diff workflow — a different tool with a different purpose, run offline against a frozen corpus rather than against your live database. And qw-miss-reporting.md is the in-app report a user files when the parser misses a question they expected to see as a widget, which is how a shape neither of these tools could predict reaches the parser team.

Last verified 2026-10-06