---
title: Main-process heartbeat (silent-crash forensic trail)
---

# Main-process heartbeat (silent-crash forensic trail)

## What it is

A dedicated log file — `heartbeat.log` — that captures a continuous 5-second-resolution tape of main-process health while Omniscio is running: event-loop lag P50/P99, RSS/heap memory, active-resource counts, and the last IPC channel that ran before each tick. It exists for the case where the Omniscio window simply **disappears with no crash dialog, no Sentry error, no Windows Event Log entry, and no Crashpad dump** — silent main-process death. With a heartbeat tape on disk, the next-launch post-mortem can answer:

1. **Was main healthy in the 0–5 s before death?** (capacity question — was lag climbing? was RSS bloating?)
2. **What was main DOING when it died?** (causation question — the `lastOp` slot stamps the most recent IPC channel name)
3. **Was this an internal hang or an external kill?** (presence vs. absence of `[CLEAN-EXIT]` / `[GOODBYE]` on shutdown; correlation with `[POST-MORTEM]` markers on next launch)

A sibling file — `heartbeat-final.log` — captures the last 10 ticks if an uncaught exception or unhandled rejection fires inside main, in case rotation dropped the live `heartbeat.log` before you could read it.

The file lives next to the other diagnostics logs:

- **Windows**: `%APPDATA%/omniscio/logs/heartbeat.log`
- **macOS**: `~/Library/Application Support/omniscio/logs/heartbeat.log`
- **Linux**: `~/.config/omniscio/logs/heartbeat.log`

Like `interaction.log` (and unlike `startup.log`), `heartbeat.log` is a **continuous tape across launches** — the same file grows every time you open the app, capped at **1 MB** with rotation to `.1` → … → `.5`. This is intentional: a silent crash that only repeats every few days needs a tape that spans those days.

## Where to find it

`heartbeat.log` sits in the Omniscio logs folder alongside the other diagnostic logs — on Windows under `%APPDATA%/omniscio/logs/`. There is no in-app viewer; it is read after the fact, most often as the evidence attached to a crash report.

## How it behaves

### How to use it

### After a silent app death

1. **Don't reopen Omniscio immediately.** The post-mortem analyser runs on the NEXT launch and reads the on-disk tape — that's good — but if you have other diagnostics to capture first (Task Manager, dxdiag, Reliability Monitor), do those before relaunching, since relaunch is also what writes the `[POST-MORTEM]` markers.
2. Relaunch Omniscio normally. During init, the heartbeat:
   - Walks back through `heartbeat.log` looking for the prior launch's `[HEARTBEAT-START]` banner.
   - Checks whether `[CLEAN-EXIT]` or `[GOODBYE]` appears in the window between that banner and the new one.
   - If neither appears, writes `[POST-MORTEM]` lines into the new launch's section flagging the silent death, the last known tick (with its `lastOp`), and any Windows Event Log entries from ±60 s around that tick.
3. Open **Settings → Diagnostics → Main-Process Heartbeat → Reveal Heartbeat Log**. The file manager opens with `heartbeat.log` highlighted.
4. Read in this order:
   - First, **`heartbeat-final.log`** in the same folder, if it exists. That's the emergency ring-buffer dump from an uncaught exception OR a render/child-process-gone — fastest path to a stack trace or to the surrounding ticks for a silent renderer death.
   - Next, the tail of **`heartbeat.log`**: scan up from the bottom for the most recent `[POST-MORTEM]` lines. They sit between the prior and current launch banners. If a `renderGone=N` or `childGone=N` count appears, the analyser also surfaces the last `[RENDER-GONE]` / `[CHILD-GONE]` line — that names the suspect (renderer vs. GPU/utility).
   - Then the last few `[tick]` / `[STALL]` / `[BOOT-STALL]` lines BEFORE the silent boundary — that's the 5–25 s window before death.
   - Finally, the **Windows Event Log scrape** lines (only on Windows): the analyser shells out to PowerShell `Get-WinEvent` for Application-Error / Application-Hang / Windows-Error-Reporting events near the last known tick. If something shows up here, the kill came from outside Omniscio (Windows OOM, security software, driver crash, etc.).
5. Attach `heartbeat.log` (and `heartbeat-final.log` if present) to the bug report.

### What you'll see in the file

```
═══════════════════════════════════════════════════════════════
 Omniscio MAIN-PROCESS HEARTBEAT
 …self-describing header (purpose, Version, Electron, Mode, PID,
  tick interval, stall threshold)…
═══════════════════════════════════════════════════════════════

── Launch resumed at 2026-05-22T18:00:00.123Z (PID 24680) ──
2026-05-22T18:00:00.124Z [HEARTBEAT-START] pid=24680 version=1.2.3 mode=development
2026-05-22T18:00:05.130Z [tick] uptime=5s lagP50=2ms lagP99=4ms rss=180MB heap=72MB lastOp="<boot>" resourceCount=18
2026-05-22T18:00:10.131Z [tick] uptime=10s lagP50=2ms lagP99=3ms rss=185MB heap=74MB lastOp="session:list" resourceCount=22
2026-05-22T18:00:15.137Z [STALL] uptime=15s lagP50=180ms lagP99=1240ms rss=195MB heap=78MB lastOp="recipe:run" resourceCount=24 resources={Timeout:12,TCPWRAP:6,FSReqCallback:3,Immediate:1,TTYWRAP:2}
… more ticks …
2026-05-22T18:02:01.211Z [CLEAN-EXIT] pid=24680 ticksWritten=24 uptime=121s
2026-05-22T18:02:01.213Z [GOODBYE] pid=24680 code=0 ticksWritten=24
```

Line legend:

- `[HEARTBEAT-START]` — banner stamped at every init. Marks a new launch.
- `[tick]` — normal 5-second sample. P99 lag below the stall threshold (1000 ms).
- `[STALL]` — same shape as `[tick]` but P99 ≥ 1000 ms AFTER the first IPC arrived. Suggests a CPU-bound operation in main (heavy synchronous SQLite, JSON parse on a huge buffer, NDJSON parse spike).
- `[BOOT-STALL]` — same as STALL but before any IPC ran. Suggests an early-startup hang (migration, native module load, sync I/O during bootstrap).
- `[STALL-CAUGHT]` — emitted by a separate 200 ms fast-tick watchdog whenever the gap between two consecutive watchdog fires is ≥ 2 000 ms (the loop was just unblocked after a multi-second pause). Line shape: `durationMs=<actual gap> lastOp="<channel>" inFlightCount=<n> inFlight=[…]`. The `inFlight` array lists every IPC that was mid-await at the moment of unblock, sorted DESC by `ageMs` (oldest first — most likely the culprit) and capped at 10 entries. Companion to `[STALL]`: the 5 s `[STALL]` says "lagP99 was high somewhere in this window," `[STALL-CAUGHT]` says "the loop was pinned for exactly N ms and these IPCs were waiting." A 2–5 s stall that ends mid-window can be missed by `[STALL]` (dilutes below the p99 threshold) but always shows up here. See `postmortem`.
- `[stall-stack]` — emitted right AFTER a `[STALL-CAUGHT]` when the V8 CPU profiler was armed and captured samples. **Pair the two by `stallId=<epochMs>`** — the freeze's unblock timestamp, carried by every record of one freeze (`[STALL-CAUGHT]`, `[stall-stack]`, `[STALL-GC]`, and the `[GC-PAUSE]` warn that lands in `main.log` instead of `heartbeat.log`). Pairing by `durationMs` was the old join and it does not work: a storm puts dozens of stalls in the same 2–4 s band, and `[GC-PAUSE]` never carried it at all. Line shape: `stallId=<epochMs> durationMs=<gap> top: <fn> (<file>:<line>) <n>% | … ← caller: <fn> (<file>:<line>)`. **`top:`** is the top-5 frames by SELF time — on a multi-second synchronous block the blocking frame dominates every other sample, so the first entry IS the blocking op (frequently a native leaf like `existsSync (native) 81%`). **`← caller:`** is the answer to "which of OUR code asked for that": the profiler walks UP the call tree from the hottest leaf to the first APP frame. **Node built-in frames are deliberately SKIPPED** — they carry a source url (`node:fs`, `node:internal/…`) but are only Node's own wrapper around the native leaf, so returning one would merely restate what `top:` already said. Before that skip existed, 22 of 26 attributed callers on a live box read `existsSync (node:fs:278)` / `statSync (node:fs:1733)` and named nothing useful. A `node:` frame still appears as a last-resort fallback when the path holds no app frame at all; `← caller:` is omitted entirely for a pure native / off-CPU wait, and the line reads `top: (no samples — likely native/off-CPU wait)` when the window had no samples. **NOTE — this is a DIFFERENT field from `blockedOp=` on the `[STALL-CAUGHT]` line**, which is fed by `trackMainOp` breadcrumbs — the DB layer, plus (since 2026-08-29) EVERY periodic background service, named `service:<label>`, because `createPeriodicTask` wraps each tick's synchronous slice. So a freeze inside a service tick now reads e.g. `blockedOp=service:agent-status-board` rather than `(none-instrumented)`; that fallback still appears for a synchronous block in code no breadcrumb covers, and for an ASYNC tick, which correctly registers no block at all. The profiler's answer cannot be folded into this field because the capture is asynchronous and resolves after that line is already written. Read the two lines together. Kill switch `AMC_DISABLE_MAIN_STALL_PROFILER=1`; armed only under box load. **Arm/disarm/rotate leave their own `[stall-profiler]` transition lines on this tape** (`armed intervalUs=… maxWindowMs=…` / `disarmed windowMs=…` / `rotated windowMs=… samplesDiscarded=…` / `arm-failed reason=…`), so "was the instrument running during the 13:20Z storm?" is answerable after the fact. **A `[stall-profiler] arm-failed` line means the profiler is DARK, and it can be dark while `armed` reads true to nobody but itself** — V8 refuses `setSamplingInterval` while a profile is live, so if a command's reply is lost under load the profiler keeps sampling while the module believes it is off. The arm path is therefore its own breaker: three consecutive failures go quiet for 60 s (logging once per episode, never once per 200 ms tick, and the tape keeps the token every time), and the next arm stops the suspect profiler before it re-enables one. This is the same trap the renderer twin hit on 2026-09-25 (bug 9acddc97), where the retry loop itself measured ~250 log lines a minute and 46.5% of a core. **And starting the profiler is itself expensive on a large process**: `Profiler.start` runs V8's `StartProfiling` on the main thread, which at a ~370 MB heap measured **2–4 s** (Help Desk 6069a960, 2026-09-29). Because that exceeds the 2 s stall threshold the watchdog recorded the instrument's own start-up as an app freeze → the freeze asked for a capture → the capture stopped the profile and re-armed → arming paid it again; `[stall-profiler] armed` went from ~20/hour to 500–980/hour while the UI lagged ~2 s per click. So the profiler now times its own start and **disarms itself for the rest of the run** when it exceeds `AMC_STALL_PROFILER_MAX_ARM_COST_MS` (500 ms — deliberately below the stall threshold, so the profiler's own work can never be counted as a stall). The tape says `[stall-profiler] self-defeated armCostMs=…`, and a later `[stall-stack]` from that process reads `unavailable=self-defeated`. **The kill switch must be set BEFORE the app launches**: setting `AMC_DISABLE_MAIN_STALL_PROFILER=1` and then using the in-app restart does NOT pick it up — `app.relaunch` inherits the old process environment, so it takes a full quit and relaunch. See `src/main/services/diagnostics/main-stall-profiler.ts`.
- `[CLEAN-EXIT]` — written at the top of `gracefulShutdown`. Disk is healthy, shutdown reached the heartbeat-stop path.
- `[GOODBYE]` — written from a `process.on('exit')` hook. Confirms the runtime actually unwound, not just that gracefulShutdown started.
- `[FATAL]` — written from `uncaughtException` / `unhandledRejection` handlers. Includes the lastOp slot and a 6-line stack snippet.
- `[RENDER-GONE]` — written from the `render-process-gone` handler in `src/main/index.ts` when `details.reason !== 'clean-exit'`. Stamps the current `lastOp` plus `reason` and `exitCode`, then dumps the ring buffer into `heartbeat-final.log` so the surrounding ticks are preserved even if rotation later trims the live tape. Tells you whether a silent death was the renderer (the BrowserWindow's web content) dying out from under main.
- `[CHILD-GONE]` — same mechanism as `[RENDER-GONE]` but from the app-level `child-process-gone` handler. Fires for GPU process / utility process / sandbox helper exits. Includes `type`, `reason`, `exitCode`, and (when present) `name` / `service`. A GPU-process death is a common cause of "the window disappeared" with no main-process exception.
- `[POST-MORTEM]` — written on the NEXT launch when the analyser sees a silent death pattern. Lines include `previousLaunchEndedSilently=true`, `lastKnownTick=…`, `os-event=…`. When `[RENDER-GONE]` or `[CHILD-GONE]` lines fell inside the prior-launch window, the analyser also surfaces `renderGone=N` / `childGone=N` counts and the last line of each, so the post-mortem names a suspect even after the live tape is gone.

`resources={Timeout:12,TCPWRAP:6,…}` shows the libuv active-resource histogram (counts of pending timers, sockets, fs requests, etc.) — present on STALL/BOOT-STALL lines because the typical silent-death pattern is "resource count climbing without bound."

### Knobs

- **Master kill switch (env)**: `AMC_DISABLE_MAIN_HEARTBEAT=1` skips init entirely. Use for perf measurements where you want zero heartbeat overhead. (Steady-state cost is ~1 ms per 5 s = ~0.02 % CPU — practically free — but the switch exists for the principled case.)
- **Master kill switch (config)**: `config.json` field `mainHeartbeatEnabled: false`. Read synchronously at init, before the config-store module is loaded. Defaults to `true`.
- **Reveal button**: Settings → Diagnostics → Main-Process Heartbeat → Reveal Heartbeat Log.

### Why it's its own file (not main.log)

`main.log` is a rolling firehose — every `log.info` from every service. After a silent crash, opening `main.log` means reading thousands of lines hoping the last few survived the buffer flush. `heartbeat.log` is focused: just the periodic samples and the lifecycle markers. The self-describing header at the top means an LLM reading a pasted copy with no other context understands every line.

## For agents

### Lifecycle wiring

- Initialised in `src/main/index.ts` immediately after the interaction-trace init (after `app.whenReady()`), via `startMainHeartbeat()`.
- Stopped at the top of `gracefulShutdown()` via `stopMainHeartbeat()` — writes `[CLEAN-EXIT]` while the disk is still healthy.
- `recordOp(channel)` called from `wrapHandler()` in `src/main/ipc/handler-wrapper.ts` so every IPC invocation updates the lastOp slot. Cost: ~100 ns per call (single string write to a module-level slot, plus a boolean flip on the first call).
- `markIpcInFlight(requestId, label)` / `unmarkIpcInFlight(requestId)` paired in the same `wrapHandler()` — mark before the handler body runs, unmark in the `finally` so a thrown exception still clears the slot. The fast-tick watchdog reads this Map on every `[STALL-CAUGHT]` capture.
- `recordProcessGone(kind, details)` called from the existing `render-process-gone` and `child-process-gone` handlers in `src/main/index.ts` when `details.reason !== 'clean-exit'`. Writes the `[RENDER-GONE]` or `[CHILD-GONE]` line and dumps the ring buffer into `heartbeat-final.log`.
- Pre-exit hooks (`uncaughtException`, `unhandledRejection`, `exit`) installed once during `startMainHeartbeat()`.

Implementation: `src/main/services/diagnostics/main-heartbeat.ts`.

## Related

The other diagnostic tapes Omniscio writes — the performance-sensitive interaction timeline and the general runtime log — are described on the [Logs and debugging](logs-and-debugging.md) page, and what happens after a crash on [Crash recovery](crash-recovery.md).
