---
title: QW Triage Tool (mine the live DB for parser bugs)
---

# QW Triage Tool (mine the live DB for parser bugs)

## What it is

`npm run triage:qw` is a developer tool that scans every `source='agent'` message in your live Omniscio SQLite database, scores each one for likely QuestionWidget (QW) parser bugs, and writes a sorted JSONL review queue under `tests/fixtures/question-widget/_triage/<UTC-timestamp>/`. Two failure modes the tool surfaces:

1. **Silent miss** — the parser returned `null` (no widget rendered) but the content has strong QW signals (bold header, lettered options, question mark). The user sees plain text where they should have seen pills.
2. **Over-fire** — the parser DID render a widget but the content is log-line spam, raw code blocks, or other prose that shouldn't have triggered. The user sees a pill widget on a paragraph that should have been plain markdown.

The tool is **read-only against the database**, **purely rule-based** (no AI, no model calls, $0), and **idempotent** — re-running without code changes produces the same JSONL because it's a pure function of (DB content, parser SHA in the baseline, scorer regex tables). It writes nothing back to Omniscio; output stays under `_triage/`.

This is not a regression test — it's a **mining tool** that surfaces candidates a human reviews and (optionally) promotes to corpus fixtures. The QW corpus regression suite (`npm run test:qw-corpus`) is a separate gate that runs on every parser change; this tool is what you run when you want to expand the corpus or audit a long-tail of real production messages for bugs you missed.

See [qw-parser-snapshots.md](qw-parser-snapshots.md) for the related frozen-snapshot diff workflow.

## Where to find it

This is a developer tool with no product surface — there is no screen, menu or setting for it. You run it from a shell with `npm run triage:qw`, and its output lands in a dated folder beside the other QuestionWidget test fixtures rather than anywhere in the running app.

## How it behaves

### How to run

```bash
npm run triage:qw                          # full scan, emit JSONL
npm run triage:qw -- --dry-run             # scan + print summary, no file write
npm run triage:qw -- --limit 100           # write only top 100 by score
npm run triage:qw -- --min-score 50        # raise the floor (default 30)
npm run triage:qw -- --dry-run --limit 50  # combine flags
```

Flags supported by `scripts/qw-triage.mjs`:

| Flag          | Default | Meaning                                                                        |
| ------------- | ------- | ------------------------------------------------------------------------------ |
| `--limit N`   | ∞       | Keep only top N candidates by score after sort (descending).                   |
| `--dry-run`   | off     | Print summary + top 10 to stdout. Skips writing `triage-report.jsonl`.         |
| `--min-score` | 30      | Drop candidates scoring below this. Lower = noisier queue; higher = strictest. |

The runner script `scripts/qw-triage-runner.cjs` spawns `qw-triage.mjs` under Electron-as-Node (`ELECTRON_RUN_AS_NODE=1` + electron.exe) + `tsx`, because the wrapper dynamic-imports `parseContentSegments` from the renderer's TypeScript source via `tsx`, so any signals derived from the agent-markdown parser stay in sync with what the live UI sees — and that preload is an Electron-as-Node runtime.

> **Correction, 2026-09-30.** This paragraph used to justify the wrapper with `better-sqlite3`: "a native module compiled against Electron's V8 ABI (NODE_MODULE_VERSION 140); plain Node 22/24 (NMV 137) crashes with `ERR_DLOPEN_FAILED`". That was true of the node-gyp era and is false of the version this repo pins (13.x), which is a Node-API add-on whose in-package prebuild is loaded before any node-gyp output — so no ABI is compared and plain `node` loads it. The wrapper is unchanged; only its stated reason was wrong.

The database path defaults to:

```
%APPDATA%\omniscio\mission-control.db        # Windows
~/Library/Application Support/omniscio/...   # macOS (via APPDATA shim)
```

Override via `AMC_DB_PATH=/some/other.db` if you want to triage a sandbox / e2e instance instead of your live DB.

The tool prints its DB path + baseline message count + fixture count BEFORE scanning so you can Ctrl-C if the wrong instance is targeted:

```
[qw-triage] DB=C:\Users\...\mission-control.db (18271 baseline msgs, 192 existing fixtures)
```

### Where the output lands

```
tests/fixtures/question-widget/_triage/<UTC-timestamp>/
  triage-report.jsonl
  triage-summary.json
```

`<UTC-timestamp>` is the script's `new Date().toISOString()` with `:` and `.` replaced by `-` (e.g. `2026-05-09T00-10-56-953Z`). Each batch lands in its own directory so concurrent runs don't clobber each other.

The whole `_triage/` directory is **gitignored** except for `.gitkeep`. Triage batches are regeneratable from the DB + parser SHA, so committing a 4 MB JSONL would just be churn — it goes stale on every parser change.

To clean up old batches, just `rm -rf tests/fixtures/question-widget/_triage/<batch>`.

#### `triage-report.jsonl` — one candidate per line

Each line is a single JSON object, written sorted by descending score:

```json
{
  "messageId": "msg_abc123",
  "timestamp": "2026-05-08T14:23:01.547Z",
  "parserDecision": "null",
  "parserBranches": [],
  "segmentIndex": 0,
  "signals": {
    "boldHeader": true,
    "letteredOptions": true,
    "numberedOptions": false,
    "bulletOptions": false,
    "questionMark": true,
    "choiceImperative": false,
    "codeFence": false,
    "logLine": false,
    "optionCount": 4,
    "scanTruncated": false,
    "contentLength": 1342
  },
  "score": 90,
  "suspicionReason": "silent-miss-strong",
  "contentExcerpt": "**Should we proceed?**\n\nA. Yes, ship it now\nB. ...(first 500 chars)..."
}
```

Field guide:

- **`parserDecision`** — what the parser did with this message according to the frozen baseline at `tests/fixtures/question-widget/_backtest/baseline.json`. Values: `'null'` (no widget rendered), `'preprocess-empty'` (the parser preprocessor stripped everything), `'renders'` (a widget WAS rendered), `'unknown'` (no baseline entry — message arrived after the last re-baseline), or `'fetch-error'` (the parser threw — score forced to 100, see content excerpt for the exception text).
- **`parserBranches`** — for `'renders'` rows, the named cascade branches that fired (e.g. `['WRAP_BOLD_COLON']`). Empty for `'null'` and `'preprocess-empty'`.
- **`segmentIndex`** — agent messages are split into prose / tool segments before parsing; this is the index of the segment the scorer picked as the most suspicious. Multi-segment messages have a winner.
- **`signals`** — boolean / count flags extracted by the rule-based regex tables in `scripts/qw-triage/signals.mjs`. Drive the score directly. `scanTruncated: true` means the message was longer than 50,000 bytes and only the prefix was scanned.
- **`score`** — 0-100, computed by `scripts/qw-triage/score.mjs`. The bucket boundaries are documented below.
- **`suspicionReason`** — short stable label describing why the candidate ranked. One of `silent-miss-strong`, `silent-miss-weak`, `preprocess-stripped-question`, `overfire-log-line`, or the literal string `score=<n>` if no named pattern fits.
- **`contentExcerpt`** — first 500 characters of the winning segment's prose. Enough to triage manually without opening the full message; truncated for JSONL row size.

#### `triage-summary.json` — companion summary

```json
{
  "parserSha": "da2af34ae",
  "triageScriptSha": null,
  "batchTimestamp": "2026-05-09T00:10:56.953Z",
  "dbPath": "C:\\Users\\...\\mission-control.db",
  "counts": {
    "scanned": 21738,
    "skippedCovered": 81,
    "skippedLowScore": 17530,
    "emitted": 4127
  },
  "scoreBuckets": {
    "90-100": 56,
    "70-89": 13,
    "50-69": 2412,
    "30-49": 1646
  }
}
```

- **`parserSha`** — short hash of the QW parser source at the time the baseline was captured. If the parser has changed since, the JSONL's `parserDecision` values are stale and you should re-baseline (see "Re-baseline contract" below).
- **`counts.skippedCovered`** — message IDs already represented in the fixture corpus. Dedupe runs by reading every `expected.json` under `tests/fixtures/question-widget/should-render/` and `should-not-render/` and skipping any message whose `id` appears in the union.
- **`counts.skippedLowScore`** — messages where the winning segment scored below `--min-score`.
- **`counts.emitted`** — how many made it to the JSONL. Equals `min(--limit, count above floor)`.
- **`scoreBuckets`** — distribution across the four scorer tiers, useful at-a-glance for whether the parser regressed or improved between batches.

### How to interpret scores

Score = sum of weighted signals, clamped to `[0, 100]`. The weight table lives in `scripts/qw-triage/score.mjs` and depends on the parser's decision for that message:

- **`parserDecision === 'null'` (silent-miss bucket)**: bold header (+30), lettered options (+30), question mark (+15), choice imperative ("which" / "should we" / "want me to" / etc., +10), 2+ options (+10), 4+ options (+5).
- **`parserDecision === 'preprocess-empty'` (parser preprocessor ate the question)**: any one signal (+50), bold + lettered together (+20).
- **`parserDecision === 'renders'` (over-fire bucket)**: log-line content (+70), code fence without bold header (+30).

Negative pressure: `logLine` signal in a non-`'renders'` row subtracts 50 (it's evidence the candidate is log spam, not a real question). Truncated scans are capped at 40 (we didn't see the whole message, so we won't claim high confidence).

#### Score buckets

| Score range | Meaning                                                                                                                                                                                       |
| ----------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **90-100**  | Strong silent-miss or preprocess-stripped — bold header AND lettered options AND question mark, all in a `'null'` parser decision. Triage these first; the highest-confidence bugs live here. |
| **70-89**   | Strong silent-miss without all three primary signals, OR a strong over-fire (log-line content with `'renders'`).                                                                              |
| **50-69**   | Mid-confidence. Bulk of the queue. Many real bugs here, but also many borderline cases — choice imperative without lettered options, etc. Skim, don't chase every one.                        |
| **30-49**   | Weak signal. Default floor. Most are noise; spot-check a handful. Lower the floor with `--min-score 0` only if you're explicitly hunting long-tail edge cases.                                |

The scorer is **monotonic** — adding signals never decreases the score. There's a unit-test invariant (`tests/unit/scripts/qw-triage-score.test.ts`) that pins this: `score("What?") <= score("What?\nA. yes") <= score("**Q:** What?\nA. yes\nB. no")`. If you add a new signal that reverses this, the test fails. Don't bypass it — monotonicity is what makes the score buckets meaningful as triage tiers.

#### `suspicionReason` labels

- **`silent-miss-strong`** — parser said `'null'`, content has bold header AND lettered options. Highest-confidence bug in the queue.
- **`silent-miss-weak`** — parser said `'null'`, content has lettered options but no bold header. Could be a real question phrased without bold emphasis, or could be a legitimate "Files changed: A. one B. two" listing.
- **`preprocess-stripped-question`** — parser said `'preprocess-empty'`. The content has signals but the parser's preprocessor (which strips tool output, code blocks, etc.) removed everything before the regex cascade ran.
- **`overfire-log-line`** — parser said `'renders'` and the content has a `[2026-05-08]` or `[ERROR]` log-line pattern. The widget is firing on log spam.
- **`score=<n>`** — fallback label when no named pattern matches. The numeric score still sorts correctly but you'll need to read the signals manually.

### How to promote a candidate to a fixture

Manual flow for v1:

1. **Open `triage-report.jsonl` and scroll through the top of the queue** (it's pre-sorted by score). Use any text editor that handles 4 MB files — VS Code, Notepad++, `less` in WSL.
2. **For each candidate worth investigating**: copy the `messageId` and look it up in your Omniscio database, OR open the source session in Omniscio and find the message. The `contentExcerpt` field gives you 500 chars to triage without leaving the JSONL.
3. **Decide which side of the corpus the fixture goes in**:
   - Real silent-miss → fixture under `tests/fixtures/question-widget/should-render/<descriptive-slug>/`.
   - Real over-fire → fixture under `tests/fixtures/question-widget/should-not-render/<descriptive-slug>/`.
   - Borderline / debatable → leave it. The queue is allowed to have noise.
4. **Author the fixture per the schema in `tests/fixtures/question-widget/README.md`** — `content.md` (the raw agent message text) + `expected.json` (`{ shouldRender, source: { messageId, sessionId, timestamp }, notes }`).
5. **Run the corpus suite to verify the fixture exercises the parser the way you expect**:
   ```bash
   npm run test:qw-corpus
   ```
   Red on a should-render fixture means the parser is missing it (file the bug + write the failing assertion first per the Omniscio `QW parsing bugs` rule); green means you've already covered it.

There is no automated promote-to-fixture command in v1. The fixture format is small enough that copy-paste is fine for the volume the triage tool produces, and a forced shape would prevent you from adding the human-readable `notes` field that makes the corpus useful as a regression document.

### Decision gate — do we add AI re-rank?

The original plan (`docs/plans/2026-05-08-qw-triage-tool-plan.md`) had a conditional Phase 4 to add Anthropic Haiku re-ranking on top of the rule-based scorer. The gate to enable it was: **after the first full run, spot-check the top 50 candidates. If ≥30% are real bugs (silent-miss or over-fire), ship rule-based-only and skip AI. If <30%, the rule-based signal is too noisy and AI is worth the cost.**

**Result of the 2026-05-08 first-run gate**: ≥30% real bugs in top 50 → rule-based shipped, AI skipped. The current scorer is the production tool.

If a future parser change degrades the signal (e.g. a new branch lands that fires on log-lines, polluting the silent-miss bucket), re-run the spot-check. If the rate drops below 30%, the conditional Phase 4 wiring is described in the plan — it would add a cache-backed Haiku call against the top 20 candidates with `--confirm-cost` gating and a hard `--max-cost 1.00` cap.

### Re-baseline contract

The triage tool's `parserDecision` field comes from the frozen baseline at `tests/fixtures/question-widget/_backtest/baseline.json`. That baseline is captured by `npm run backtest:qw` against the parser at the time of capture, and stamped with the parser SHA.

**You must re-baseline whenever the QW parser changes**, otherwise:

- A real silent-miss the new parser would now catch still shows up in the queue (because the baseline says it returned `null`).
- A real over-fire the new parser no longer fires on still shows up (because the baseline says it returned `'renders'`).

Re-baseline command:

```bash
npm run backtest:qw -- --update-baseline
```

This rewrites `_backtest/baseline.json` with the current parser's decisions across every message in the database, and updates the embedded `parserSha` field. Commit the new baseline alongside the parser change so subsequent triage runs use the fresh decisions.

The `triage-summary.json` records `parserSha` per batch, so you can always look at an old batch and check whether it's stale relative to today's parser by comparing SHAs.

### Concurrency and safety

- **DB access is read-only**: `better-sqlite3` opens with `{ readonly: true, fileMustExist: true }`. The triage tool cannot mutate the live Omniscio database.
- **Atomic file writes**: both `triage-report.jsonl` and `triage-summary.json` write to a `.staging` sibling first, then `renameSync` into place. A killed run mid-write never corrupts a previous batch.
- **Concurrent runs**: each batch directory is named after `new Date().toISOString()`. Two `triage:qw` runs would only collide on directory name if they fire in the same millisecond (the timestamp includes milliseconds) — accept the low probability.
- **Omniscio must not be holding a write lock**: the tool opens the DB read-only, but if Omniscio is mid-WAL-checkpoint, the open may briefly block. Run with Omniscio closed if you want maximum safety; run with Omniscio open if you want a current snapshot.

## For agents

### Implementation pointers (for agents touching this code)

- Entry script: `scripts/qw-triage.mjs` — reads DB, loops messages, calls scorer per segment, sorts and writes JSONL.
- Native-binding wrapper: `scripts/qw-triage-runner.cjs` — `ELECTRON_RUN_AS_NODE=1` + `electron.exe` + `tsx` so `better-sqlite3` and the renderer's TS source both load.
- Signals: `scripts/qw-triage/signals.mjs` — the rule-based regex tables (bold header, lettered options, log line, etc.).
- Scorer: `scripts/qw-triage/score.mjs` — weighted signal sum, monotonic, clamped 0-100.
- Segment picker: `scripts/qw-triage/pick-segment.mjs` — picks the highest-scoring prose segment when a message has multiple.
- Dedupe: `scripts/qw-triage/dedupe.mjs` — reads every `expected.json` under the corpus root and returns a `Set<messageId>` of already-covered IDs.
- Emit: `scripts/qw-triage/emit.mjs` — atomic JSONL + summary writer.
- Validation gate (regression test): `tests/integration/qw-triage-validation-gate.test.ts` — fails CI if the scorer wrongly flags more than 5 should-render fixtures as silent-miss-strong, or if it would surface ANY should-not-render fixture as a false-positive over-fire candidate (when content has neither a log-line nor a code-fence). This is what stops "improvements" to the scorer from regressing precision against the curated corpus.
- Score unit tests: `tests/unit/scripts/qw-triage-score.test.ts` — pin the per-bucket scoring contract (silent-miss-strong ≥ 80, log-line over-fire ≥ 60, rhetorical questions < 30, monotonicity).

### References

- `tests/fixtures/question-widget/README.md` — fixture authoring schema for promoted candidates.
- `docs/plans/2026-05-08-qw-triage-tool-plan.md` — the build plan, including the conditional AI re-rank phase.

## Related

[question-widget.md](question-widget.md) covers what the widget is, the parser versions Omniscio ships, and the plumbing behind them. The other half of the parser-regression story is [qw-parser-snapshots.md](qw-parser-snapshots.md), the locked-snapshot diff workflow — a different tool with a different purpose, run offline against a frozen corpus rather than against your live database. And [qw-miss-reporting.md](qw-miss-reporting.md) is the in-app report a user files when the parser misses a question they expected to see as a widget, which is how a shape neither of these tools could predict reaches the parser team.
