Weekly Session Analysis (agent-friction dashboard + weekly deep analysis)
A read-only developer panel that mines your past agent sessions for friction and waste. A free, always-on dashboard tallies timeouts, hook blocks, tool errors and redundant cd prefixes, plus spend concentration; a paid, opt-in deep run spawns one silent session that writes a ranked, lever-tagged report.
What it is
A read-only developer-tools panel inside Omniscio that mines your past agent sessions for friction and waste, so you can tune the things that make your fleet slower or costlier. It has two layers:
- A free, always-on deterministic friction dashboard. The friction signatures — command timeouts, hook blocks, tool errors, redundant
cd "<worktree>" &&prefixes (and bash-command volume as the denominator) — are tallied per session as its output streams in (at NDJSON ingestion, while the raw tool results are still in hand; the result text is inspected for the signatures then discarded, so nothing bulky is stored). On a bounded schedule the dashboard just SUMs those per-session counters over the recent-long window, plus spend concentration (total window spend, how much of it lands in the priciest 10% of sessions, and the top-cost tail). No AI, no spend, and no transcript files — so it covers every session, not just the ones whose transcripts still happen to be on disk. - An opt-in deep agent-run analysis. On demand it spawns a single silent Claude session that runs a forensic methodology over your recent sessions and writes a ranked, lever-tagged report (each finding maps to a lever: HOOK / SKILL / CLAUDE.md / MEMORY / CONTRACT / TOOL / CONFIG / WORKFLOW / CONTEXT / OTHER). This is paid work — it is cost-capped, runs one at a time, and over the CLI it is approval-gated.
It lives in the Developer Tools sidebar group, and it is a non-spawnable virtual project — you don't start Claude sessions "inside" it; the deep-run session spawns under a real project. The panel is visible; the scan behind it is not running by default, for the reason in the next section.
Where to find it
Turning it on — and the gate with no switch
Two gates, and they are not the same one.
The panel is shipped. The registry entry is status: 'shipped', and the visibility gate returns true unconditionally for that state — so the surface is visible on a normal install with nothing to reveal and no Lab row, because listInDevelopmentFeatures() only lists in-development and labs entries and this one is neither.
The scan is the gate that decides everything else. sessionForensicsEnabled must be exactly true, and it defaults to false. Two paths honour it: the deep-run IPC handler (src/main/ipc/session-forensics-handlers.ts) and the periodic scanner (scanNow() in src/main/services/session-forensics/session-forensics-service.ts, which also stands down while the service is paused). The "Scan now" path is deliberately exempt — the manual trigger POST /session-forensics/scan calls scanOnce() → runScan() directly, bypassing both checks, so a scan the user explicitly asked for always runs. (The same is true of the other CLI run routes.) The gate is aimed at the automatic work, not at a person pressing the button.
And nothing in the interface writes it: no file under src/renderer names sessionForensicsEnabled at all, so there is no toggle to find — the Lab row the registry's settingKey implies is never rendered, precisely because the entry says shipped. On a default install the dashboard therefore has nothing to show and a deep run is refused.
Until a switch ships, the only way to open the gate is to set sessionForensicsEnabled to true in the config directly. The registry entry's AMC_SHOW_SESSION_FORENSICS escape hatch is inert here — an env var only reveals a surface, and this surface is already unconditionally visible because the entry is shipped, so there is nothing left for it to reveal and it has no bearing on the scan gate either way.
The companion settings ride the same flag: the scan interval (sessionForensicsScanIntervalHours, 1–168h, default 6), the deep-run daily spend cap (sessionForensicsDailyCostCapUsd, $0.05–$20, default $2), and the weekly auto-run toggle (sessionForensicsAutoWeeklyEnabled, default off).
How it behaves
The dashboard
When the panel has at least one scan, the layout reads top-to-bottom as a triage view:
- A health headline (the focal point) — names your single worst friction signal in that signal's severity colour, or shows a calm "Running clean" verdict when every signal is healthy. The worst is ranked by severity tier (bad > notable > healthy), tie-broken by a fixed priority — never by raw percentage, because the signals use different denominators.
- Friction signals — four severity-coloured cards: Timeout sessions, Hook-block sessions, Tool-error sessions, and Redundant cd (redundant
cdprefixes as a share of all bash calls). Each shows its percentage, a thin severity bar, and the raw count. Severity reads three ways at once — colour, bar length, and the number — so it never rides on colour alone: a calm surface tone when healthy, amber (status-needs-you) when notable, red (status-error) when bad. Thresholds: the session-percent signals are healthy < 10 / notable 10–33 / bad ≥ 33; redundant-cd (a lower-stakes, differently-scaled signal) is healthy < 5 / notable 5–20 / bad ≥ 20. - Agent workflow insights (F016) — three informational cards projected from the same deterministic cross-session dialogue aggregates the paid deep run uses, surfaced back to you for free (reused, never recomputed, no schema change): Correction load (the share of your human messages that re-steer the agent, with a few PII-scrubbed sample quotes), Repeated launch templates (recurring session kick-offs with no matching skill — automation candidates, shown as an "automate these" list), and Reactive vs proactive (the firefighting-vs-building work mix, as a two-segment bar). Neutral-toned (these aren't "wrong", so no severity colour); each card shows a calm empty state until a scan has data.
- Spend & cost — a separate, informational section (money isn't "wrong", so it is never severity-coloured): the Spend (window) total, the Top-10% spend share concentration, and the Priciest sessions list (rank · name · source · hours · turns · cost, each with a relative-cost bar).
A Scan now button runs an immediate deterministic pass (re-SUMs the DB counters). Everything reflects only the recent window (the friction counters are aggregated only over sessions that ran in the window), so it is labelled by window, never implied as all-time. Before the first scan the empty state previews the metrics a scan will surface (a labelled skeleton grid — no invented numbers) so the panel shows its value instead of a bare button.
The deep analysis
Run deep analysis starts a paid forensic run. It is taxonomy-driven and token-light: the heavy "aggregate across every session" work is done deterministically FIRST (no LLM), so the paid agent reasons over a few-KB digest instead of ~150 raw transcripts. It:
- pre-flight checks the daily cost cap (
SUM(cost_micro_usd)of today's runs vs the cap) and refuses if one is already running; - deterministically pre-computes a cross-session digest (
dialogue-aggregates.ts— correction-language rate + top clusters, repeated launch-templates w/ skill-exists flag, idle-ratio outliers, model-mix, reactive:proactive ratio, spend tail) and embeds it inline in the prompt — NO temp file (a red-team simplification; removes a path/IO failure mode); - spawns a silent session with
source: 'session-forensics'under a real project (the last active project, else the first one) on Sonnet (DEEP_RUN_MODEL = 'claude-sonnet-4-6'— a mid-tier model is ample for applying a fixed taxonomy to a digest, and keeps per-run cost low), injecting a bundled forensic prompt (version-controlled, not a runtime read of a user file) that carries the digest + an 8-family pattern taxonomy (F1 Repetition/Automation … F8 Agent/meta) + 7 finding-quality rules (aggregate-across-sessions · evidence-lock · concrete-change · quantify-impact · tag-owner · reinforce-what-works · don't-over-count-healthy-signal); - the agent writes
findings.json+report.mdto<userData>/session-forensics/<runId>/(a containment-validated path passed in the prompt); - a watcher (mtime-gated + a max-runtime watchdog) parses the findings into a
session_forensics_runsrow and records the spawned session's cost.
Each finding is richer than a bare lever tag — patternClass (taxonomy id), frequency (the cross-session count = the value multiplier), quote (verbatim evidence), concreteChange (the actual rule/skill/setting to make), impact (quantified payoff), and polarity (problem vs reinforce — a working pattern worth codifying). All the new fields are OPTIONAL + defensively parsed/clamped, so older run rows still render.
Runs are triageable in the panel — expand to read the ranked findings + report, mark read, or dismiss (reversible). Orphaned running rows left by a crash are reconciled to failed on startup.
Weekly auto-run (opt-in). With Run the deep analysis automatically every week on (sessionForensicsAutoWeeklyEnabled, default off), the deep analysis also fires on its own once a week (Monday morning, local time): at most one run per week, never overlapping an in-flight run, and within the same daily spend cap. It is doubly gated (both sessionForensicsEnabled and the weekly toggle must be on). The scheduler is an hourly check in session-forensics-service.ts calling the pure isWeeklyDeepRunDue in session-forensics-weekly.ts.
For agents
- Data model. Two tables (migration
20260717023721):session_forensics_scans(one deterministic snapshot per scan) andsession_forensics_runs(one deep-run report; carries thecost_micro_usdtwin for the daily cap). Queries:src/main/db/queries-session-forensics.ts. - Friction capture (at ingestion). The five friction signals are tallied per turn in
processNdjsonEvent(src/main/process/ndjson-event-handlers.ts) via the shared pure detectoraccumulateFrictionFromEntry(transcript-scan.ts), accumulated onsession.turnFriction, and flushed to the session row at the result event alongsideupdateSessionCost. Best-effort (try/catch): a friction write must NEVER throw into the streaming/cost path (freeze-safety — the user-experience-inviolable directive). Only integer counts are stored; the tool-result text is inspected then discarded. Per-session columnsfriction_bash_commands/_redundant_cd/_tool_errors/_timeouts/_hook_blocks(migration20260727012313, additive; integer baseline stays 261). Persist + aggregate helpers:incrementSessionFriction/getFrictionAggregateinqueries-session-forensics.ts. - Scanner.
src/main/services/session-forensics/session-forensics-service.ts—registerService+createPeriodicTask;runScanInnerSUMs the per-session friction counters over the recent-long window (getFrictionAggregate) — no transcript files, no 40-longest selection bias (the old path re-read~/.claude/projectsat scan time and prioritized the longest worktree-orchestrator sessions, whose transcripts had already rolled off disk → it analyzed ~1 of 40). Spend stats still read session metadata directly (bounded top-active set). Gated onsessionForensicsEnabled+ the service-registry pause.scanOnce()powers Scan-now;buildSessionForensicsOverview()powers the panel payload. No historical backfill — pre-existing sessions show 0 friction (their result text is already gone); the dashboard fills forward.transcriptsAvailableon the scan snapshot is kept for shape-stability and now equalssessionsScanned. - Deep run.
src/main/services/session-forensics/deep-run.ts—startDeepRun()(computes the digest →buildForensicPrompt({ aggregatesDigest })→ spawns onDEEP_RUN_MODEL = 'claude-sonnet-4-6') + a fire-and-forgetwatchRun. The canonical spawn is injected inindex.tsviasetForensicsSpawnDependencies(...)→createSessionWithPrompt(the spawn seam gained an optionalmodel; never/project/new).parseForensicFindings(raw)is the exported, pure, unit-tested normalizer — validates the lever/polarity enums, clamps field lengths, tolerates garbage →[]. Never invokestartDeepRunfrom tests/agents live — it spawns a real, billable session; validate via the built prompt string + afindings.jsonfixture instead. - Dialogue aggregates (the token-light core).
src/main/services/session-forensics/dialogue-aggregates.ts— purecomputeDialogueAggregates(rows)+renderAggregatesDigest(agg)(unit-tested intests/unit/services/session-forensics/dialogue-aggregates.test.ts) + a boundedgetDialogueAggregates()DB wrapper (windowed, capped, keyed on indexed columns, empty/archived-safe, never throws → degrades to an empty digest). Runs once per opt-in deep run AND once per free 6-hourly scan (F016, below) — never on a hot loop, so it stays off the freeze path. The prompt lives inforensic-prompt.ts(buildForensicPrompt+ the exportedFINDINGS_JSON_CONTRACT), unit-tested (forensic-prompt-and-parse.test.ts) for the 8-family taxonomy, the 7 rules, digest-embed, and a ≤9KB size bound. Two UI labels stay English pending a translation pass (gated feature) — exempted insrc/shared/i18n/untranslated-exemptions.ts(sessionForensicsSupercharge2026). - IPC. Channels
session-forensics:get-overview | scan-now | run-deep | mark-read | dismiss-run+ the pushsession-forensics:updated. Handlers insrc/main/ipc/session-forensics-handlers.ts(auto-discovered). - CLI.
GET /session-forensics,POST /session-forensics/scan,POST /session-forensics/runs/:id/read,POST /session-forensics/runs/:id/dismiss(apply-immediately), andPOST /session-forensics/run— approval-gated (session_forensics.run_deep, NON_TOGGLEABLE; thesession-forensics-runaction handler runsstartDeepRun()at approve time). - Free workflow insights (F016). The free dashboard surfaces a slim projection of the dialogue aggregates. On each free scan
runScanInnerrefreshes an in-memorylastDialogueInsightsfrom the already-boundedgetDialogueAggregates()read (reused, never recomputed, NO DB column — mirrors thelastScanErrorin-memory pattern), set BEFORE thesession-forensics:updatedpush so the panel reload sees fresh cards;buildSessionForensicsOverview()surfaces it asdialogueInsightson the overview payload and stays synchronous. The pure projectiontoDialogueInsights()(dialogue-insights.ts, unit-testeddialogue-insights.test.ts) re-shapes the fat aggregate onto the slimDialogueInsightsshared type — only the three cards' fields, somodelMix/topCost/idleRatioOutliersnever leak into the free payload. Rendered byWorkflowInsightsSection.tsx(display-only; all message text is plain-escaped React text and already PII-scrubbed at the read boundary).nulluntil the session's first scan; wiring locked bysession-forensics-dialogue-insights.test.ts. New labels are fully translated (all 23 locales), not exempted like the gated deep-run's two supercharge labels. - Severity model (presentation only).
src/renderer/src/features/session-forensics/forensics-severity.ts— a pure, unit-tested (tests/unit/renderer/forensics-severity.test.ts) helper holding the friction thresholds,severityOf(), andworstFrictionSignal()(tier-ranked, fixed-priority tie-break, NaN-safe). The panel paints what it returns through the canonicalstatus-*tokens; the scan payload and backend are unchanged. - Gating. Registered as an
amc-builtinunderdeveloper-tools-group(src/shared/integrations/session-forensics.ts); sentinel__session_forensics__; unreleased-feature idsession-forensics; UI panelsrc/renderer/src/features/session-forensics/SessionForensicsPanel.tsx(panelOwnsLayout: true).
Related
Forensics reads a session’s past; these are the raw records it reads.
- Session event log — the per-session record of what happened, in order.
- Startup trace — the same step-by-step trace, for startup.
- Logs and debugging — where the logs live and what to look for in them.
Last verified 2026-10-01