Omniscio documentation
Browse all documentation
  1. Getting Started13
  2. Sessions & Agents115
  3. Inbox & Notifications59
  4. Projects & Tasks95
  5. Automation & Scheduling75
  6. Knowledge & Memory26
  7. AI Features60
  8. Integrations100
  9. Plugins & Marketplace33
  10. Cloud & Teams56
  11. Settings & Customization58
  12. Account & Billing28
  13. Troubleshooting84
  14. CLI & API Reference22
  15. Legal & Policies4
  16. Uncategorised22

Main-process heartbeat (silent-crash forensic trail)

A log file recording a continuous five-second tape of Omniscio's main-process health — event-loop lag, memory, live resource counts and the last operation that ran — so that a window which simply disappears with no crash dialog and no error report can still be explained after the fact.

What it is

A dedicated log file — heartbeat.log — that captures a continuous 5-second-resolution tape of main-process health while Omniscio is running: event-loop lag P50/P99, RSS/heap memory, active-resource counts, and the last IPC channel that ran before each tick. It exists for the case where the Omniscio window simply disappears with no crash dialog, no Sentry error, no Windows Event Log entry, and no Crashpad dump — silent main-process death. With a heartbeat tape on disk, the next-launch post-mortem can answer:

  1. Was main healthy in the 0–5 s before death? (capacity question — was lag climbing? was RSS bloating?)
  2. What was main DOING when it died? (causation question — the lastOp slot stamps the most recent IPC channel name)
  3. Was this an internal hang or an external kill? (presence vs. absence of [CLEAN-EXIT] / [GOODBYE] on shutdown; correlation with [POST-MORTEM] markers on next launch)

A sibling file — heartbeat-final.log — captures the last 10 ticks if an uncaught exception or unhandled rejection fires inside main, in case rotation dropped the live heartbeat.log before you could read it.

The file lives next to the other diagnostics logs:

  • Windows: %APPDATA%/omniscio/logs/heartbeat.log
  • macOS: ~/Library/Application Support/omniscio/logs/heartbeat.log
  • Linux: ~/.config/omniscio/logs/heartbeat.log

Like interaction.log (and unlike startup.log), heartbeat.log is a continuous tape across launches — the same file grows every time you open the app, capped at 1 MB with rotation to .1 → … → .5. This is intentional: a silent crash that only repeats every few days needs a tape that spans those days.

Where to find it

heartbeat.log sits in the Omniscio logs folder alongside the other diagnostic logs — on Windows under %APPDATA%/omniscio/logs/. There is no in-app viewer; it is read after the fact, most often as the evidence attached to a crash report.

How it behaves

How to use it

After a silent app death

  1. Don't reopen Omniscio immediately. The post-mortem analyser runs on the NEXT launch and reads the on-disk tape — that's good — but if you have other diagnostics to capture first (Task Manager, dxdiag, Reliability Monitor), do those before relaunching, since relaunch is also what writes the [POST-MORTEM] markers.
  2. Relaunch Omniscio normally. During init, the heartbeat:
    • Walks back through heartbeat.log looking for the prior launch's [HEARTBEAT-START] banner.
    • Checks whether [CLEAN-EXIT] or [GOODBYE] appears in the window between that banner and the new one.
    • If neither appears, writes [POST-MORTEM] lines into the new launch's section flagging the silent death, the last known tick (with its lastOp), and any Windows Event Log entries from ±60 s around that tick.
  3. Open Settings → Diagnostics → Main-Process Heartbeat → Reveal Heartbeat Log. The file manager opens with heartbeat.log highlighted.
  4. Read in this order:
    • First, heartbeat-final.log in the same folder, if it exists. That's the emergency ring-buffer dump from an uncaught exception OR a render/child-process-gone — fastest path to a stack trace or to the surrounding ticks for a silent renderer death.
    • Next, the tail of heartbeat.log: scan up from the bottom for the most recent [POST-MORTEM] lines. They sit between the prior and current launch banners. If a renderGone=N or childGone=N count appears, the analyser also surfaces the last [RENDER-GONE] / [CHILD-GONE] line — that names the suspect (renderer vs. GPU/utility).
    • Then the last few [tick] / [STALL] / [BOOT-STALL] lines BEFORE the silent boundary — that's the 5–25 s window before death.
    • Finally, the Windows Event Log scrape lines (only on Windows): the analyser shells out to PowerShell Get-WinEvent for Application-Error / Application-Hang / Windows-Error-Reporting events near the last known tick. If something shows up here, the kill came from outside Omniscio (Windows OOM, security software, driver crash, etc.).
  5. Attach heartbeat.log (and heartbeat-final.log if present) to the bug report.

What you'll see in the file

═══════════════════════════════════════════════════════════════
 Omniscio MAIN-PROCESS HEARTBEAT
 …self-describing header (purpose, Version, Electron, Mode, PID,
  tick interval, stall threshold)…
═══════════════════════════════════════════════════════════════

── Launch resumed at 2026-05-22T18:00:00.123Z (PID 24680) ──
2026-05-22T18:00:00.124Z [HEARTBEAT-START] pid=24680 version=1.2.3 mode=development
2026-05-22T18:00:05.130Z [tick] uptime=5s lagP50=2ms lagP99=4ms rss=180MB heap=72MB lastOp="<boot>" resourceCount=18
2026-05-22T18:00:10.131Z [tick] uptime=10s lagP50=2ms lagP99=3ms rss=185MB heap=74MB lastOp="session:list" resourceCount=22
2026-05-22T18:00:15.137Z [STALL] uptime=15s lagP50=180ms lagP99=1240ms rss=195MB heap=78MB lastOp="recipe:run" resourceCount=24 resources={Timeout:12,TCPWRAP:6,FSReqCallback:3,Immediate:1,TTYWRAP:2}
… more ticks …
2026-05-22T18:02:01.211Z [CLEAN-EXIT] pid=24680 ticksWritten=24 uptime=121s
2026-05-22T18:02:01.213Z [GOODBYE] pid=24680 code=0 ticksWritten=24

Line legend:

  • [HEARTBEAT-START] — banner stamped at every init. Marks a new launch.
  • [tick] — normal 5-second sample. P99 lag below the stall threshold (1000 ms).
  • [STALL] — same shape as [tick] but P99 ≥ 1000 ms AFTER the first IPC arrived. Suggests a CPU-bound operation in main (heavy synchronous SQLite, JSON parse on a huge buffer, NDJSON parse spike).
  • [BOOT-STALL] — same as STALL but before any IPC ran. Suggests an early-startup hang (migration, native module load, sync I/O during bootstrap).
  • [STALL-CAUGHT] — emitted by a separate 200 ms fast-tick watchdog whenever the gap between two consecutive watchdog fires is ≥ 2 000 ms (the loop was just unblocked after a multi-second pause). Line shape: durationMs=<actual gap> lastOp="<channel>" inFlightCount=<n> inFlight=[…]. The inFlight array lists every IPC that was mid-await at the moment of unblock, sorted DESC by ageMs (oldest first — most likely the culprit) and capped at 10 entries. Companion to [STALL]: the 5 s [STALL] says "lagP99 was high somewhere in this window," [STALL-CAUGHT] says "the loop was pinned for exactly N ms and these IPCs were waiting." A 2–5 s stall that ends mid-window can be missed by [STALL] (dilutes below the p99 threshold) but always shows up here. See postmortem.
  • [stall-stack] — emitted right AFTER a [STALL-CAUGHT] when the V8 CPU profiler was armed and captured samples. Pair the two by stallId=<epochMs> — the freeze's unblock timestamp, carried by every record of one freeze ([STALL-CAUGHT], [stall-stack], [STALL-GC], and the [GC-PAUSE] warn that lands in main.log instead of heartbeat.log). Pairing by durationMs was the old join and it does not work: a storm puts dozens of stalls in the same 2–4 s band, and [GC-PAUSE] never carried it at all. Line shape: stallId=<epochMs> durationMs=<gap> top: <fn> (<file>:<line>) <n>% | … ← caller: <fn> (<file>:<line>). top: is the top-5 frames by SELF time — on a multi-second synchronous block the blocking frame dominates every other sample, so the first entry IS the blocking op (frequently a native leaf like existsSync (native) 81%). ← caller: is the answer to "which of OUR code asked for that": the profiler walks UP the call tree from the hottest leaf to the first APP frame. Node built-in frames are deliberately SKIPPED — they carry a source url (node:fs, node:internal/…) but are only Node's own wrapper around the native leaf, so returning one would merely restate what top: already said. Before that skip existed, 22 of 26 attributed callers on a live box read existsSync (node:fs:278) / statSync (node:fs:1733) and named nothing useful. A node: frame still appears as a last-resort fallback when the path holds no app frame at all; ← caller: is omitted entirely for a pure native / off-CPU wait, and the line reads top: (no samples — likely native/off-CPU wait) when the window had no samples. NOTE — this is a DIFFERENT field from blockedOp= on the [STALL-CAUGHT] line, which is fed by trackMainOp breadcrumbs — the DB layer, plus (since 2026-08-29) EVERY periodic background service, named service:<label>, because createPeriodicTask wraps each tick's synchronous slice. So a freeze inside a service tick now reads e.g. blockedOp=service:agent-status-board rather than (none-instrumented); that fallback still appears for a synchronous block in code no breadcrumb covers, and for an ASYNC tick, which correctly registers no block at all. The profiler's answer cannot be folded into this field because the capture is asynchronous and resolves after that line is already written. Read the two lines together. Kill switch AMC_DISABLE_MAIN_STALL_PROFILER=1; armed only under box load. Arm/disarm/rotate leave their own [stall-profiler] transition lines on this tape (armed intervalUs=… maxWindowMs=… / disarmed windowMs=… / rotated windowMs=… samplesDiscarded=… / arm-failed reason=…), so "was the instrument running during the 13:20Z storm?" is answerable after the fact. A [stall-profiler] arm-failed line means the profiler is DARK, and it can be dark while armed reads true to nobody but itself — V8 refuses setSamplingInterval while a profile is live, so if a command's reply is lost under load the profiler keeps sampling while the module believes it is off. The arm path is therefore its own breaker: three consecutive failures go quiet for 60 s (logging once per episode, never once per 200 ms tick, and the tape keeps the token every time), and the next arm stops the suspect profiler before it re-enables one. This is the same trap the renderer twin hit on 2026-09-25 (bug 9acddc97), where the retry loop itself measured ~250 log lines a minute and 46.5% of a core. And starting the profiler is itself expensive on a large process: Profiler.start runs V8's StartProfiling on the main thread, which at a ~370 MB heap measured 2–4 s (Help Desk 6069a960, 2026-09-29). Because that exceeds the 2 s stall threshold the watchdog recorded the instrument's own start-up as an app freeze → the freeze asked for a capture → the capture stopped the profile and re-armed → arming paid it again; [stall-profiler] armed went from ~20/hour to 500–980/hour while the UI lagged ~2 s per click. So the profiler now times its own start and disarms itself for the rest of the run when it exceeds AMC_STALL_PROFILER_MAX_ARM_COST_MS (1000 ms — set ABOVE the measured start cost, whose 28 recorded defeats run 500–807 ms, so a healthy run does not latch the instrument off; and set well below the 2 s stall threshold — not exactly 2× under it, because the watchdog's gap also carries the arm's own pre-start round trips — so the profiler's own work cannot be counted as a stall. It was 500 ms until 2026-10-04, which sat below the very costs it was bounding and took the instrument out on essentially every run; re-derive it from the armCostMs= now printed on the armed line). The tape says [stall-profiler] self-defeated armCostMs=…, and a later [stall-stack] from that process reads unavailable=self-defeated. The kill switch must be set BEFORE the app launches: setting AMC_DISABLE_MAIN_STALL_PROFILER=1 and then using the in-app restart does NOT pick it up — app.relaunch inherits the old process environment, so it takes a full quit and relaunch. See src/main/services/diagnostics/main-stall-profiler.ts.
  • [CLEAN-EXIT] — written at the top of gracefulShutdown. Disk is healthy, shutdown reached the heartbeat-stop path.
  • [GOODBYE] — written from a process.on('exit') hook. Confirms the runtime actually unwound, not just that gracefulShutdown started.
  • [FATAL] — written from uncaughtException / unhandledRejection handlers. Includes the lastOp slot and a 6-line stack snippet.
  • [RENDER-GONE] — written from the render-process-gone handler in src/main/index.ts when details.reason !== 'clean-exit'. Stamps the current lastOp plus reason and exitCode, then dumps the ring buffer into heartbeat-final.log so the surrounding ticks are preserved even if rotation later trims the live tape. Tells you whether a silent death was the renderer (the BrowserWindow's web content) dying out from under main.
  • [CHILD-GONE] — same mechanism as [RENDER-GONE] but from the app-level child-process-gone handler. Fires for GPU process / utility process / sandbox helper exits. Includes type, reason, exitCode, and (when present) name / service. A GPU-process death is a common cause of "the window disappeared" with no main-process exception.
  • [POST-MORTEM] — written on the NEXT launch when the analyser sees a silent death pattern. Lines include previousLaunchEndedSilently=true, lastKnownTick=…, os-event=…. When [RENDER-GONE] or [CHILD-GONE] lines fell inside the prior-launch window, the analyser also surfaces renderGone=N / childGone=N counts and the last line of each, so the post-mortem names a suspect even after the live tape is gone.

resources={Timeout:12,TCPWRAP:6,…} shows the libuv active-resource histogram (counts of pending timers, sockets, fs requests, etc.) — present on STALL/BOOT-STALL lines because the typical silent-death pattern is "resource count climbing without bound."

Knobs

  • Master kill switch (env): AMC_DISABLE_MAIN_HEARTBEAT=1 skips init entirely. Use for perf measurements where you want zero heartbeat overhead. (Steady-state cost is ~1 ms per 5 s = ~0.02 % CPU — practically free — but the switch exists for the principled case.)
  • Master kill switch (config): config.json field mainHeartbeatEnabled: false. Read synchronously at init, before the config-store module is loaded. Defaults to true.
  • Reveal button: Settings → Diagnostics → Main-Process Heartbeat → Reveal Heartbeat Log.

Why it's its own file (not main.log)

main.log is a rolling firehose — every log.info from every service. After a silent crash, opening main.log means reading thousands of lines hoping the last few survived the buffer flush. heartbeat.log is focused: just the periodic samples and the lifecycle markers. The self-describing header at the top means an LLM reading a pasted copy with no other context understands every line.

For agents

Lifecycle wiring

  • Initialised in src/main/index.ts immediately after the interaction-trace init (after app.whenReady()), via startMainHeartbeat().
  • Stopped at the top of gracefulShutdown() via stopMainHeartbeat() — writes [CLEAN-EXIT] while the disk is still healthy.
  • recordOp(channel) called from wrapHandler() in src/main/ipc/handler-wrapper.ts so every IPC invocation updates the lastOp slot. Cost: ~100 ns per call (single string write to a module-level slot, plus a boolean flip on the first call).
  • markIpcInFlight(requestId, label) / unmarkIpcInFlight(requestId) paired in the same wrapHandler() — mark before the handler body runs, unmark in the finally so a thrown exception still clears the slot. The fast-tick watchdog reads this Map on every [STALL-CAUGHT] capture.
  • recordProcessGone(kind, details) called from the existing render-process-gone and child-process-gone handlers in src/main/index.ts when details.reason !== 'clean-exit'. Writes the [RENDER-GONE] or [CHILD-GONE] line and dumps the ring buffer into heartbeat-final.log.
  • Pre-exit hooks (uncaughtException, unhandledRejection, exit) installed once during startMainHeartbeat().

Implementation: src/main/services/diagnostics/main-heartbeat.ts.

Related

The other diagnostic tapes Omniscio writes — the performance-sensitive interaction timeline and the general runtime log — are described on the Logs and debugging page, and what happens after a crash on Crash recovery.

Last verified 2026-10-04