---
title: QW Parser Snapshots (regression-diff baseline)
---

# QW Parser Snapshots (regression-diff baseline)

## What it is

**HISTORICAL (v1).** This snapshot methodology — the locked 18,271-row corpus and the `onBranchFire` / `QW_PARSER_BRANCHES` branch-fire instrumentation it relies on — was **v1-specific, and v1 was removed (audit F019).** The page is kept for history. Current QW parser work goes through the [QW parser contract](/.claude/memory/contracts/qw-parser-contract.md) and the v3 corpus — [`question-widget-corpus.test.ts`](/tests/unit/components/question-widget-corpus.test.ts) and [`qw-product-path-characterization.test.ts`](/tests/unit/components/qw-product-path-characterization.test.ts).

The v1 QuestionWidget (QW) parser (the now-deleted `src/shared/question-widget-parser.ts`) decided whether an agent message contains a multiple-choice question and, if so, where the question text starts and which options are attached. Every regex / structural-guard tweak in that parser risked two failure modes:

1. **Over-fire** — prose that wasn't a question gets parsed as one. The user sees a pill widget on a paragraph that should have been plain text.
2. **Miss** — a real question stops being recognized. The user has to type the answer instead of clicking a pill.

To catch both, we run the parser against a **frozen sample of 18,271 real agent messages** and diff the output before and after the change. The frozen sample is called the **production snapshot**.

This page describes what the snapshot is, where it lives, how to re-capture it, and why we deliberately keep it stale.

## Where to find it

This is a developer-only workflow with no product surface — there is no screen, menu or setting for it. Everything lives in a scratch directory under your system temp folder, called `qw-mining`, alongside the other QW regression artifacts; the exact paths and file formats are listed under **For agents** below.

## How it behaves

### Why we keep it locked

The snapshot is **deliberately stale**. We do not auto-refresh it.

The reason is regression-diff legibility. When you change the parser and re-run the cycle, the orchestrator reports two sets:

- **appeared** — message IDs that fire NOW but didn't before. If they're real questions the parser was missing, great. If they're prose, that's a new over-fire.
- **disappeared** — message IDs that fired before but don't now. If they were over-fires the change was meant to remove, great. If they were real questions, that's a regression.

If the snapshot is allowed to drift between runs, **a brand-new message that arrived since the last capture also shows up in `appeared`** — indistinguishable from a real over-fire introduced by the parser change. The triage gets ambiguous. So we freeze the snapshot once and compare every parser change against the same 18,271 rows. The signal stays clean.

**Re-capture policy**: explicit approval only. Default cadence is quarterly or before a major architectural change to the parser. After a re-capture, expect a one-time spike in `appeared` rows that need fresh triage — bake that triage into the re-capture decision.

### How the snapshot is used

Three QW tools read the snapshot in different ways:

| Tool                         | Reads snapshot via                | Output                                                                        |
| ---------------------------- | --------------------------------- | ----------------------------------------------------------------------------- |
| `tools/qw-mine-overfires.ts` | `--snapshot <path>`               | `candidates.jsonl` — every message that fires the parser, ranked by suspicion |
| `tools/qw-sample-render.ts`  | `--snapshot <path>`               | `sample.json` — deterministic 100-row render at `--seed 42` for visual diffs  |
| `tools/qw-cycle.sh`          | passes `--snapshot` to both above | `cycle-summary.md` — verdict (`OK` / `HALT: layer N`) + four layer reports    |

`qw-cycle.sh` is the one-button orchestrator: it runs the four regression layers in order (corpus test → mine-diff → sample-render → sample-diff) against a per-cycle output directory so two cycles can run side-by-side without clobbering each other. Pass `--snapshot <path>` so all layers see the same frozen rows.

### What to do when over-fires evolve

Over-fire patterns aren't static. New agent prose styles emerge, new model versions phrase things differently, and intentional parser changes will deliberately remove some over-fires while adding others.

The honest workflow is:

1. **Before the change** — run `qw-cycle.sh` and save the four baselines (`baseline-corpus.txt`, `baseline-mine.jsonl`, `baseline-sample.json`, `baseline-commit.txt`).
2. **Apply the change.**
3. **After the change** — run `qw-cycle.sh` again, then diff `candidates.jsonl` against `baseline-mine.jsonl` and `sample.json` against `baseline-sample.json`.
4. **Triage**:
   - Rows in **disappeared** that match the change's intent → expected, ship it.
   - Rows in **appeared** → inspect every one. Each is either (a) a real new over-fire (regression — fix or revert) or (b) a real question the parser was missing that is now correctly fired (good — note in the commit message).
5. **If the diff is too large to triage manually** — narrow the parser change. The whole point of the cycle is to keep diffs small enough that a human can read every row.

Per-branch instrumentation in production (Task 8, 2026-05-03) feeds the same signal from a different angle: every accept point in the parser fires `feature_events: { feature: 'qw_parser_branch_fire', metadata: { branch: 'WRAP_BOLD_COLON' | … } }`. The Stats virtual project can show which branches dominate live traffic, so a parser change that drains one branch's volume on the locked snapshot can be cross-checked against the live event stream.

## For agents

### Where the snapshot lives

```
${os.tmpdir()}/qw-mining/messages.jsonl
```

On this Windows box that resolves to:

```
C:/Users/<you>/AppData/Local/Temp/qw-mining/messages.jsonl
```

- Format: one JSON object per line — `{ id, session_id, timestamp, content }`.
- Source rows: every `conversation_messages` row where `source = 'agent'` (the renderer never sees this file; it's a developer artifact).
- Size: ~104 MB.
- Lines: 18,271 (as of the 2026-05-03 capture).

The directory name `qw-mining` is also the working directory for every other QW regression artifact — `candidates.jsonl` (mine output), `sample.json` (deterministic 100-row render), per-cycle baselines (`baseline-corpus.txt`, `baseline-mine.jsonl`, `baseline-sample.json`, `baseline-commit.txt`).

### How to re-capture

The capture tool reads Omniscio's live `mission-control.db` and writes every `source='agent'` message to JSONL.

```bash
cd "C:/Software Projects/Agent Orchestrator"   # or your worktree
SNAP="$(node -e 'console.log(require("path").join(require("os").tmpdir(), "qw-mining"))')"
mkdir -p "$SNAP"
ELECTRON_RUN_AS_NODE=1 \
  ./node_modules/electron/dist/electron.exe \
  tools/qw-mine-overfires.ts --capture-snapshot "$SNAP/messages.jsonl"
```

Why the `ELECTRON_RUN_AS_NODE=1` + Electron-binary dance? The original reason was `better-sqlite3` — "a native module compiled against Electron's V8 ABI (NODE_MODULE_VERSION 140), so plain `node` (ABI 137) fails with `ERR_DLOPEN_FAILED`". **That reason is retired as of 2026-09-30**: the version this repo pins (13.x) is a Node-API add-on whose in-package prebuild is loaded before any node-gyp output, so no ABI is compared and plain `node` loads it. The command above is unchanged — running the tool on the same runtime the app and the test launcher use is still what it is for.

The write is **atomic**: the tool writes to `<path>.tmp` and then renames into place. A killed capture mid-write never corrupts the existing snapshot.

Sanity-check the new snapshot before relying on it:

```bash
wc -l "$SNAP/messages.jsonl"            # expect tens of thousands
head -1 "$SNAP/messages.jsonl" | jq .   # expect { id, session_id, timestamp, content }
```

### References

- Plan: `docs/plans/2026-05-03-qw-parser-corpus-convergence-plan.md` (25-task convergence roadmap).
- Active corpus tests: `tests/unit/components/question-widget-corpus.test.ts` (192 active fixtures + 41 pending — pending is the convergence target).
- Branch-fire instrumentation (v1, removed with the v1 parser): was the now-deleted `src/shared/question-widget-parser.ts` `onBranchFire` callback + its `src/main/process/ndjson-utils.ts` subscriber. Neither the file nor the `onBranchFire` symbol exists today.

## Related

[question-widget.md](question-widget.md) covers what the QuestionWidget is, its parser versions and the plumbing behind it. The same two failure modes this snapshot diff catches offline are what [qw-triage.md](qw-triage.md) hunts for online, with its read-only rule-based scorer mining the live database for likely silent-miss and over-fire bugs. And when a real user spots a miss the parser team never predicted, [qw-miss-reporting.md](qw-miss-reporting.md) is the in-app one-click report that captures those fresh shapes the frozen snapshot can never see.
