Main-process heartbeat (silent-crash forensic trail)
A log file recording a continuous five-second tape of Omniscio's main-process health — event-loop lag, memory, live resource counts and the last operation that ran — so that a window which simply disappears with no crash dialog and no error report can still be explained after the fact.
What it is
A dedicated log file — heartbeat.log — that captures a continuous 5-second-resolution tape of main-process health while Omniscio is running: event-loop lag P50/P99, RSS/heap memory, active-resource counts, and the last IPC channel that ran before each tick. It exists for the case where the Omniscio window simply disappears with no crash dialog, no Sentry error, no Windows Event Log entry, and no Crashpad dump — silent main-process death. With a heartbeat tape on disk, the next-launch post-mortem can answer:
- Was main healthy in the 0–5 s before death? (capacity question — was lag climbing? was RSS bloating?)
- What was main DOING when it died? (causation question — the
lastOpslot stamps the most recent IPC channel name) - Was this an internal hang or an external kill? (presence vs. absence of
[CLEAN-EXIT]/[GOODBYE]on shutdown; correlation with[POST-MORTEM]markers on next launch)
A sibling file — heartbeat-final.log — captures the last 10 ticks if an uncaught exception or unhandled rejection fires inside main, in case rotation dropped the live heartbeat.log before you could read it.
The file lives next to the other diagnostics logs:
- Windows:
%APPDATA%/omniscio/logs/heartbeat.log - macOS:
~/Library/Application Support/omniscio/logs/heartbeat.log - Linux:
~/.config/omniscio/logs/heartbeat.log
Like interaction.log (and unlike startup.log), heartbeat.log is a continuous tape across launches — the same file grows every time you open the app, capped at 1 MB with rotation to .1 → … → .5. This is intentional: a silent crash that only repeats every few days needs a tape that spans those days.
Where to find it
heartbeat.log sits in the Omniscio logs folder alongside the other diagnostic logs — on Windows under %APPDATA%/omniscio/logs/. There is no in-app viewer; it is read after the fact, most often as the evidence attached to a crash report.
How it behaves
How to use it
After a silent app death
- Don't reopen Omniscio immediately. The post-mortem analyser runs on the NEXT launch and reads the on-disk tape — that's good — but if you have other diagnostics to capture first (Task Manager, dxdiag, Reliability Monitor), do those before relaunching, since relaunch is also what writes the
[POST-MORTEM]markers. - Relaunch Omniscio normally. During init, the heartbeat:
- Walks back through
heartbeat.loglooking for the prior launch's[HEARTBEAT-START]banner. - Checks whether
[CLEAN-EXIT]or[GOODBYE]appears in the window between that banner and the new one. - If neither appears, writes
[POST-MORTEM]lines into the new launch's section flagging the silent death, the last known tick (with itslastOp), and any Windows Event Log entries from ±60 s around that tick.
- Walks back through
- Open Settings → Diagnostics → Main-Process Heartbeat → Reveal Heartbeat Log. The file manager opens with
heartbeat.loghighlighted. - Read in this order:
- First,
heartbeat-final.login the same folder, if it exists. That's the emergency ring-buffer dump from an uncaught exception OR a render/child-process-gone — fastest path to a stack trace or to the surrounding ticks for a silent renderer death. - Next, the tail of
heartbeat.log: scan up from the bottom for the most recent[POST-MORTEM]lines. They sit between the prior and current launch banners. If arenderGone=NorchildGone=Ncount appears, the analyser also surfaces the last[RENDER-GONE]/[CHILD-GONE]line — that names the suspect (renderer vs. GPU/utility). - Then the last few
[tick]/[STALL]/[BOOT-STALL]lines BEFORE the silent boundary — that's the 5–25 s window before death. - Finally, the Windows Event Log scrape lines (only on Windows): the analyser shells out to PowerShell
Get-WinEventfor Application-Error / Application-Hang / Windows-Error-Reporting events near the last known tick. If something shows up here, the kill came from outside Omniscio (Windows OOM, security software, driver crash, etc.).
- First,
- Attach
heartbeat.log(andheartbeat-final.logif present) to the bug report.
What you'll see in the file
═══════════════════════════════════════════════════════════════
Omniscio MAIN-PROCESS HEARTBEAT
…self-describing header (purpose, Version, Electron, Mode, PID,
tick interval, stall threshold)…
═══════════════════════════════════════════════════════════════
── Launch resumed at 2026-05-22T18:00:00.123Z (PID 24680) ──
2026-05-22T18:00:00.124Z [HEARTBEAT-START] pid=24680 version=1.2.3 mode=development
2026-05-22T18:00:05.130Z [tick] uptime=5s lagP50=2ms lagP99=4ms rss=180MB heap=72MB lastOp="<boot>" resourceCount=18
2026-05-22T18:00:10.131Z [tick] uptime=10s lagP50=2ms lagP99=3ms rss=185MB heap=74MB lastOp="session:list" resourceCount=22
2026-05-22T18:00:15.137Z [STALL] uptime=15s lagP50=180ms lagP99=1240ms rss=195MB heap=78MB lastOp="recipe:run" resourceCount=24 resources={Timeout:12,TCPWRAP:6,FSReqCallback:3,Immediate:1,TTYWRAP:2}
… more ticks …
2026-05-22T18:02:01.211Z [CLEAN-EXIT] pid=24680 ticksWritten=24 uptime=121s
2026-05-22T18:02:01.213Z [GOODBYE] pid=24680 code=0 ticksWritten=24
Line legend:
[HEARTBEAT-START]— banner stamped at every init. Marks a new launch.[tick]— normal 5-second sample. P99 lag below the stall threshold (1000 ms).[STALL]— same shape as[tick]but P99 ≥ 1000 ms AFTER the first IPC arrived. Suggests a CPU-bound operation in main (heavy synchronous SQLite, JSON parse on a huge buffer, NDJSON parse spike).[BOOT-STALL]— same as STALL but before any IPC ran. Suggests an early-startup hang (migration, native module load, sync I/O during bootstrap).[STALL-CAUGHT]— emitted by a separate 200 ms fast-tick watchdog whenever the gap between two consecutive watchdog fires is ≥ 2 000 ms (the loop was just unblocked after a multi-second pause). Line shape:durationMs=<actual gap> lastOp="<channel>" inFlightCount=<n> inFlight=[…]. TheinFlightarray lists every IPC that was mid-await at the moment of unblock, sorted DESC byageMs(oldest first — most likely the culprit) and capped at 10 entries. Companion to[STALL]: the 5 s[STALL]says "lagP99 was high somewhere in this window,"[STALL-CAUGHT]says "the loop was pinned for exactly N ms and these IPCs were waiting." A 2–5 s stall that ends mid-window can be missed by[STALL](dilutes below the p99 threshold) but always shows up here. Seepostmortem.[stall-stack]— emitted right AFTER a[STALL-CAUGHT]when the V8 CPU profiler was armed and captured samples. Pair the two bystallId=<epochMs>— the freeze's unblock timestamp, carried by every record of one freeze ([STALL-CAUGHT],[stall-stack],[STALL-GC], and the[GC-PAUSE]warn that lands inmain.loginstead ofheartbeat.log). Pairing bydurationMswas the old join and it does not work: a storm puts dozens of stalls in the same 2–4 s band, and[GC-PAUSE]never carried it at all. Line shape:stallId=<epochMs> durationMs=<gap> top: <fn> (<file>:<line>) <n>% | … ← caller: <fn> (<file>:<line>).top:is the top-5 frames by SELF time — on a multi-second synchronous block the blocking frame dominates every other sample, so the first entry IS the blocking op (frequently a native leaf likeexistsSync (native) 81%).← caller:is the answer to "which of OUR code asked for that": the profiler walks UP the call tree from the hottest leaf to the first APP frame. Node built-in frames are deliberately SKIPPED — they carry a source url (node:fs,node:internal/…) but are only Node's own wrapper around the native leaf, so returning one would merely restate whattop:already said. Before that skip existed, 22 of 26 attributed callers on a live box readexistsSync (node:fs:278)/statSync (node:fs:1733)and named nothing useful. Anode:frame still appears as a last-resort fallback when the path holds no app frame at all;← caller:is omitted entirely for a pure native / off-CPU wait, and the line readstop: (no samples — likely native/off-CPU wait)when the window had no samples. NOTE — this is a DIFFERENT field fromblockedOp=on the[STALL-CAUGHT]line, which is fed bytrackMainOpbreadcrumbs — the DB layer, plus (since 2026-08-29) EVERY periodic background service, namedservice:<label>, becausecreatePeriodicTaskwraps each tick's synchronous slice. So a freeze inside a service tick now reads e.g.blockedOp=service:agent-status-boardrather than(none-instrumented); that fallback still appears for a synchronous block in code no breadcrumb covers, and for an ASYNC tick, which correctly registers no block at all. The profiler's answer cannot be folded into this field because the capture is asynchronous and resolves after that line is already written. Read the two lines together. Kill switchAMC_DISABLE_MAIN_STALL_PROFILER=1; armed only under box load. Arm/disarm/rotate leave their own[stall-profiler]transition lines on this tape (armed intervalUs=… maxWindowMs=…/disarmed windowMs=…/rotated windowMs=… samplesDiscarded=…/arm-failed reason=…), so "was the instrument running during the 13:20Z storm?" is answerable after the fact. A[stall-profiler] arm-failedline means the profiler is DARK, and it can be dark whilearmedreads true to nobody but itself — V8 refusessetSamplingIntervalwhile a profile is live, so if a command's reply is lost under load the profiler keeps sampling while the module believes it is off. The arm path is therefore its own breaker: three consecutive failures go quiet for 60 s (logging once per episode, never once per 200 ms tick, and the tape keeps the token every time), and the next arm stops the suspect profiler before it re-enables one. This is the same trap the renderer twin hit on 2026-09-25 (bug 9acddc97), where the retry loop itself measured ~250 log lines a minute and 46.5% of a core. And starting the profiler is itself expensive on a large process:Profiler.startruns V8'sStartProfilingon the main thread, which at a ~370 MB heap measured 2–4 s (Help Desk 6069a960, 2026-09-29). Because that exceeds the 2 s stall threshold the watchdog recorded the instrument's own start-up as an app freeze → the freeze asked for a capture → the capture stopped the profile and re-armed → arming paid it again;[stall-profiler] armedwent from ~20/hour to 500–980/hour while the UI lagged ~2 s per click. So the profiler now times its own start and disarms itself for the rest of the run when it exceedsAMC_STALL_PROFILER_MAX_ARM_COST_MS(1000 ms — set ABOVE the measured start cost, whose 28 recorded defeats run 500–807 ms, so a healthy run does not latch the instrument off; and set well below the 2 s stall threshold — not exactly 2× under it, because the watchdog's gap also carries the arm's own pre-start round trips — so the profiler's own work cannot be counted as a stall. It was 500 ms until 2026-10-04, which sat below the very costs it was bounding and took the instrument out on essentially every run; re-derive it from thearmCostMs=now printed on thearmedline). The tape says[stall-profiler] self-defeated armCostMs=…, and a later[stall-stack]from that process readsunavailable=self-defeated. The kill switch must be set BEFORE the app launches: settingAMC_DISABLE_MAIN_STALL_PROFILER=1and then using the in-app restart does NOT pick it up —app.relaunchinherits the old process environment, so it takes a full quit and relaunch. Seesrc/main/services/diagnostics/main-stall-profiler.ts.[CLEAN-EXIT]— written at the top ofgracefulShutdown. Disk is healthy, shutdown reached the heartbeat-stop path.[GOODBYE]— written from aprocess.on('exit')hook. Confirms the runtime actually unwound, not just that gracefulShutdown started.[FATAL]— written fromuncaughtException/unhandledRejectionhandlers. Includes the lastOp slot and a 6-line stack snippet.[RENDER-GONE]— written from therender-process-gonehandler insrc/main/index.tswhendetails.reason !== 'clean-exit'. Stamps the currentlastOpplusreasonandexitCode, then dumps the ring buffer intoheartbeat-final.logso the surrounding ticks are preserved even if rotation later trims the live tape. Tells you whether a silent death was the renderer (the BrowserWindow's web content) dying out from under main.[CHILD-GONE]— same mechanism as[RENDER-GONE]but from the app-levelchild-process-gonehandler. Fires for GPU process / utility process / sandbox helper exits. Includestype,reason,exitCode, and (when present)name/service. A GPU-process death is a common cause of "the window disappeared" with no main-process exception.[POST-MORTEM]— written on the NEXT launch when the analyser sees a silent death pattern. Lines includepreviousLaunchEndedSilently=true,lastKnownTick=…,os-event=…. When[RENDER-GONE]or[CHILD-GONE]lines fell inside the prior-launch window, the analyser also surfacesrenderGone=N/childGone=Ncounts and the last line of each, so the post-mortem names a suspect even after the live tape is gone.
resources={Timeout:12,TCPWRAP:6,…} shows the libuv active-resource histogram (counts of pending timers, sockets, fs requests, etc.) — present on STALL/BOOT-STALL lines because the typical silent-death pattern is "resource count climbing without bound."
Knobs
- Master kill switch (env):
AMC_DISABLE_MAIN_HEARTBEAT=1skips init entirely. Use for perf measurements where you want zero heartbeat overhead. (Steady-state cost is ~1 ms per 5 s = ~0.02 % CPU — practically free — but the switch exists for the principled case.) - Master kill switch (config):
config.jsonfieldmainHeartbeatEnabled: false. Read synchronously at init, before the config-store module is loaded. Defaults totrue. - Reveal button: Settings → Diagnostics → Main-Process Heartbeat → Reveal Heartbeat Log.
Why it's its own file (not main.log)
main.log is a rolling firehose — every log.info from every service. After a silent crash, opening main.log means reading thousands of lines hoping the last few survived the buffer flush. heartbeat.log is focused: just the periodic samples and the lifecycle markers. The self-describing header at the top means an LLM reading a pasted copy with no other context understands every line.
For agents
Lifecycle wiring
- Initialised in
src/main/index.tsimmediately after the interaction-trace init (afterapp.whenReady()), viastartMainHeartbeat(). - Stopped at the top of
gracefulShutdown()viastopMainHeartbeat()— writes[CLEAN-EXIT]while the disk is still healthy. recordOp(channel)called fromwrapHandler()insrc/main/ipc/handler-wrapper.tsso every IPC invocation updates the lastOp slot. Cost: ~100 ns per call (single string write to a module-level slot, plus a boolean flip on the first call).markIpcInFlight(requestId, label)/unmarkIpcInFlight(requestId)paired in the samewrapHandler()— mark before the handler body runs, unmark in thefinallyso a thrown exception still clears the slot. The fast-tick watchdog reads this Map on every[STALL-CAUGHT]capture.recordProcessGone(kind, details)called from the existingrender-process-goneandchild-process-gonehandlers insrc/main/index.tswhendetails.reason !== 'clean-exit'. Writes the[RENDER-GONE]or[CHILD-GONE]line and dumps the ring buffer intoheartbeat-final.log.- Pre-exit hooks (
uncaughtException,unhandledRejection,exit) installed once duringstartMainHeartbeat().
Implementation: src/main/services/diagnostics/main-heartbeat.ts.
Related
The other diagnostic tapes Omniscio writes — the performance-sensitive interaction timeline and the general runtime log — are described on the Logs and debugging page, and what happens after a crash on Crash recovery.
Last verified 2026-10-04