Cron failure alerts
What happens when a scheduled job fails for good: the Windows toast, the persistent inbox card that stays until you acknowledge, snooze, or the job recovers, the detail pane's Fix with AI / Pause / View actions, and the settings and silence rules that govern them.
What it is
Omniscio's cron scheduler runs scripts and recipes on a fixed schedule. When one of those jobs fails for the last time — meaning every retry has been exhausted — Omniscio raises a Windows toast and drops a persistent card into the unified inbox so you cannot miss the failure. The card stays in the inbox until you acknowledge it, snooze it (for a duration you pick), or the next scheduled run of that same job succeeds (in which case the card auto-dismisses without you doing anything).
This is a separate feature from cron self-healing. Self-healing is opt-in per job and spawns a Claude session to fix the failure. Failure alerts fire on every job, every time, without opting in — they're the visible breadcrumb that says "this thing broke, you should look at it." If you have self-healing turned on for a job AND that job's heal pipeline is mid-flight, the failure alert is suppressed for the same window so you don't see two cards racing for the same failure (see "Interaction with self-healing" below).
The feature is on by default (cronFailureAlertsEnabled = true). Turn it off entirely at Settings → Notifications → Cron failure alerts, or keep it on but route the OS toast around Focus Mode batching with the nested Always fire cron failure alerts toggle.
Where to find it
The two settings that control this feature live at Settings → Notifications. Everything else about it is something you receive rather than open: a failure reaches you as a Windows toast in the corner of your screen and as a persistent card in Omniscio's unified inbox, filed under CLI Pending Actions. The card is the one you act on — clicking the toast only brings the Omniscio window to the front.
How it behaves
What you see
The Windows toast
When a cron job exhausts its retries and lands in failed, Omniscio fires a system notification using the error sound profile:
- Title:
Cron Job Failed - Body:
<job name> failed after N attempt(s)(or just<job name> failedfor a job withretryCount = 0)
Click the toast and the Omniscio window comes to the foreground — that's the same focus-the-window behavior every other Omniscio notification uses.
The inbox card
A cron.failure_alert row lands in the unified inbox under CLI Pending Actions. Layout:
- A red status dot on the left (no other CLI Pending kind uses red — this is the visual cue that it's a failure card).
- The job name on top, truncated if long.
- The relative timestamp of the most recent failure on the right.
- Below the name:
Failed N time(s)— the cumulative count of failures captured by this card. - An optional amber Self-heal pending… chip when a heal is mid-flight for this job.
- A monospace one-line preview of the most recent stderr / error output.
The whole row is the click target — there is no inline button. Clicking opens the inline detail pane described below (same surface that arrow-key navigation, swipe activation, and archive auto-advance route to). The failure-count display caps at 50+ once it crosses 50 so a runaway job that's failed 4,000 times doesn't blow out the row width.
Card identity: one card per job
The producer is insert-or-bump, not insert-per-failure. The first failure of <job X> inserts a new pending row. The second failure of the same job updates the existing row's failureCount, refreshes lastErrorText and lastFailureAt, and leaves created_at and firstFailureAt untouched — so the inbox shows one card per failing job, not one per failure event. The total failureCount is the cumulative number of failures the card represents.
Auto-dismiss on success
When the job's next scheduled run succeeds, Omniscio marks the open card approved automatically (per-kind status semantics: approved = acknowledged) and emits a push so the inbox UI removes it without any user click. Auto-dismiss only fires once per row — if the row was already non-pending (you'd already acknowledged or dismissed it), the success branch skips the push. A row you snoozed is still pending (just hidden), so a later success still auto-dismisses it.
Auto-resolve on job deletion
Deleting the job also closes its open failure card. The card carries the job id in target_id, but deleteCronJob is a HARD DELETE FROM cron_jobs WHERE id=? with no FK cascade to cli_pending_actions — so a deleted job used to leave its card behind, and every action on it then resolved a job that no longer existed: Fix it (startManualHeal → getCronJob → null) bounced with "This job no longer exists.", Pause job failed the same way, and only Dismiss worked (with nothing telling you that). An AFTER DELETE ON cron_jobs trigger (trg_cron_jobs_failure_alert_resolve_delete) now marks the deleted job's PENDING cron.failure_alert / cron.failure_heal rows approved on every delete path — the manual delete, the plugin-bridge upsert/unschedule, and Nighty-Tidy unsubscribe — and the CRON_JOB_DELETE handler emits CLI_PENDING_CHANGED so the inbox drops the closed card live. A one-time backfill in the same migration cleared any cards already orphaned before the trigger shipped. Only pending rows are touched, so an already-acknowledged/dismissed/snoozed card is left as-is. Full invariant: the cron-failures inbox contract § I9.
The detail pane
Clicking the inbox card opens an inline detail pane (ApprovalPaneShell) showing the full failure context — the same shell every other CLI Pending approval kind uses. The activation path is the standard cli-pending one: setActiveApproval({ kind: 'cli-pending', id }), dispatched to CronFailureAlertPane by CliPendingApprovalModal based on actionKind === 'cron.failure_alert'. This means click, arrow-key nav (J/K), swipe activation, and archive auto-advance all converge on the same surface.
Footer action buttons, each with its keyboard shortcut advertised on hover — a key-cap in the button's tooltip (from its title + aria-keyshortcuts), NOT an always-visible footer hint line: N Fix · P Pause · O View (plus H Snooze on the header clock, and S Auto-fix whenever the Auto-fix row is shown — the job has loaded and no fix is mid-flight). The bare keys are routed through the global approval-keyboard intercept via a per-pane registry (approval-pane-hotkeys.ts) so they act on the open card — H snoozes this card, not a background session, and the pane's N beats the global New-Session N because the approval intercept runs first (its dispatcher sits ahead of the general binding lookup and consumes the key). Full invariants in the frontend-inbox-row contract § "Specialized cron-pane hotkeys".
The three footer actions are built around actually resolving the failure, not just hiding it (the older set — Snooze / Run Now / View job / Dismiss — only hid it, re-failed it, or navigated away):
- Fix with AI (
N, secondary, with the new-sessionMessageSquarePlusicon — the same icon Omniscio uses elsewhere for "start a session") — the primary useful action: starts a Claude session to debug and fix the failing job. It keys offN, Omniscio's standard "new session" shortcut, and opens the app's normal Start-session dialog — the same sharedStartSessionDialogthe digest / scheduled-message cards use — pre-filled with an editable opening message (the job name + its captured error) and defaulted to the job's own repo, so you can add to the message before launching. The dialog IS the confirm — a stray tap can't spawn a session; you still hit "Start session". On a successful launch the pane pauses the job (toggleJob(jobId, false)) so it can't keep re-failing while you work, toasts, and resolves the alert so the inbox advances. Disabled while a heal is already running (payload.healPending) and for job types a heal can't repair (session / one-off). (This replaced the earlier bare confirm-dialog +healNowserver-heal spawn — the escalation card still uses that manual-heal path, see cron-self-healing.md § "Manual Fix it". The launched session is a normal session, so its prompt does NOT carry the cron/run·/togglecontrol-server tools; you resume the paused schedule from the app once the fix is confirmed.) - Pause job (
P, secondary) —useCronStore.toggleJob(jobId, false): stops the job firing (and re-failing / re-alerting) until you deal with it. ToastsPaused "<job>"on success; leaves the card open so you can still Fix / View / Archive. - View job (
O, secondary) — switches the active project to the Cron Jobs virtual project (CRON_PROJECT_ID) and selects the failing job so you can inspect its schedule, env vars, command, and full run history. Closes the pane, leaves the pending row for later resolution. (Renamed from "View details" — it just opens the job.)
There is no footer Dismiss button — the top-right Archive button (below) already removes the card (a side-effect-free reject-based dismiss), so a separate acknowledge-style Dismiss was redundant.
Snooze is no longer a footer button — the shell header's clock icon already snoozes this row, so the duplicate footer Snooze was removed. H still opens Omniscio's universal snooze picker (snooze-entity CustomEvent → SnoozePalette → inbox_snoozes, the same picker the header clock and every other inbox item use). Because the universal snooze only HIDES the row (it does not resolve the cli_pending action), the pane subscribes to the snooze store and auto-closes once isInboxItemSnoozed('cli-pending-approval', row.id) is true.
Archive — the shell header's top-right corner carries the shared Archive button (the same one every inbox item has; on mobile it's pinned at the fixed top-right corner via MobileInboxArchivePin, exactly where you'd tap the X). It removes the card from the inbox and advances to the next item. Because a cron-failure alert is acknowledge-only (not a consent request), the shared button's default click resolves it through the standard inbox-dismiss chokepoint — a side-effect-free reject (Dismissed from inbox); reject never runs the dispatch, so nothing is spawned or re-fired. This is why an acknowledge-only card gets an Archive button while a true approve/reject card must not (archiving one would silently reject a real request). The header's close ✗ still just closes the pane without removing the row. The sibling cron self-heal escalation card and the nighty-tidy summary card carry the same Archive button for the same reason; the escalation card's Legacy (consent) variant deliberately does not. See mobile-archive-placement-contract.md.
Auto-fix (S) — a first-class row inside the details block (see the pane body below), NOT a footer button and no longer a loose inline "Turn on" link. The row is an Auto-fix label + an honest one-line state + a trailing <ToggleSwitch> bound to the job's healingEnabled. The toggle works both ways — flipping it calls useCronStore.updateJob(jobId, { healingEnabled: next }) (→ IPC.CRON_JOB_UPDATE) in place with an Auto-fix on/off for "<job>" toast, and is disabled while the update is in flight so a rapid double-tap can't double-fire. The S shortcut flips it too, and is registered whenever the row is shown (the job has loaded and no automatic fix is mid-flight, i.e. !payload.healPending; while a fix runs the amber heal-pending banner covers it and the row is hidden).
The pane body leads with a shared ApprovalDetailsBlock — the same bordered, hairline-divided icon + UPPERCASE-label + value block today's other CLI approval cards use — so this card reads consistently with them. The heal-pending banner sits above the block; Auto-fix is now a ROW inside the block (with a trailing on/off toggle — the one interactive control allowed inside the card, per the approval-standard-ui-contract I3 trailing-slot carve-out); the error output + recent runs stay below it. The body shows:
- Heal-pending banner (only when
payload.healPending === true, above the block) — amber callout readingA fix is already running automatically. - Job (the block's first row) — the full job name via
humanizeCronJobName(payload.jobName): a machine slug (enable-show-all-final-blocks) shows as "Enable Show All Final Blocks" with the raw slug ontitle=hover, and a name already written with spaces/capitals renders verbatim. This is the ONE place the COMPLETE name always lives — theApprovalPaneShellheader title istruncated, so on a narrow (mobile) viewport it clips toenable-sh…. The same humanized label is reused in the pane title, the Fix-it confirm, the toasts, and the snooze label so a raw slug never shows in one place while the clean name shows in another (mirrors the cron-approval card's Job row — cron-approval-display-contract I7). - Failures (the block's amber emphasis row) — a truthful summary derived from the job's REAL recent-run history via
summarizeRecentFailures(recentRuns, payload.failureCount):Failed the last N runs in a row · last succeeded <date>,Failed the latest run, orFailed N of the last M runs. It falls back toFailed N time(s)(the card's bump count) only when no run history is available — the bump count alone is misleading because it resets every time the card is dismissed, so a job that failed every week showed "Failed 1 time". - Last failed (a block row, only when
payload.lastFailureAtis set) — the most-recent-failure time viaformatInboxDateTime(fullJuly 11, 1:48 PM). Replaces the old inline· Last failure <date>append the loose summary text used to carry. - Auto-fix (a block row with a trailing toggle, shown when the job has loaded and
!healPending) — anAuto-fixlabel, a trailing on/offToggleSwitchreflectinghealingEnabled, and an HONEST one-line state as the value: OFF →Off. Turn on to auto-fix future failures.; ON →On, but it hasn't fixed this failure yet.The ON copy deliberately does NOT use a green✓ handledcheck: healing being enabled is not the same as it having fixed anything (heal attempts can be zero, or it can't fix this failure type), and a green "handled" check on a still-failing job reads as a lie. A fix genuinely in progress surfaces via the amber heal-pending banner instead (and this row is hidden then). This one row replaces BOTH the old separate "self-healing status" note and the loose inline "Turn on" control. - Last error output — the job's real stderr/error, in a monospace, scroll-on-overflow
<pre>block (capped at 1024 chars at the producer side). When the run genuinely captured no output, an honest empty-state (No error output was captured for this run.) replaces the blank box. Historically this box was ALWAYS empty for every card:notifyJobResultbuilt the alert from the pre-updaterunsnapshot (whoseerrorOutputwas still null), not the committed DB row. It now refetches the run (getCronJobRun) before building the alert, so every card carries the real error. - Recent runs — the last 5 runs of this job pulled from the database via
IPC.CRON_JOB_RUNS, rendered by the sharedCronRecentRunsList(cron-run-status.tsx): each row is a status pill (a colored dot + the Title-Case status label,Failed/Success) plus a cleanMon D, h:mm AMrun time — the same status style the job's own run-history screen (RunHistoryList) uses. Both this card and the sibling escalation card share the one strip; the old hand-rolled bare-lowercase-word + rawtoLocaleString()(seconds-precision) strip, which had been duplicated verbatim across the two panes, is gone. TheFailuressummary's "last succeeded" date and the run times now route through the shareddate-utilsformatters (formatInboxDate/ a clean short date+time), so the card no longer shows three different date styles at once.
If the payload JSON is malformed (should be impossible — the producer validates against cronFailureAlertPayloadSchema at insert time), the pane falls back to a degraded body: Unable to parse this alert. The pending row may be corrupt.
Why a pane and not a centered dialog?
The pane replaces an earlier centered DialogShell modal that bypassed the inbox approval producer→consumer dispatcher (it called CLI_PENDING_APPROVE directly), which meant Acknowledge did not auto-advance the inbox to the next approval and broke under push races. It also split activation across two surfaces — click opened the dialog, arrow-key nav opened the generic JSON-payload pane — so the experience drifted between input methods. The pane fixes both: one surface, dispatcher routing, auto-advance.
Skip conditions
The producer in cron-engine-service.ts walks the following gates in order. The first matching condition short-circuits — no card, no toast — and the order matters because the heal-pending and per-kind-cap checks happen inside the same db.transaction() that does the insert:
AMC_INSTANCE_IDis set — e2e and Claude-sandbox isolation. Both the toast (innotification-service.fireSystemNotification) and the card insert (ininsertOrBumpFailureAlert) bail out instantly when the env var is non-empty. This is why you never see cron-failure alerts in the sandbox or innpm run test:e2e:prod.cronFailureAlertsEnabled === false— the user opt-out at Settings → Notifications. Suppresses both the toast and the card insert.run.systemMarkedFailure === true— the startup reconciler flips this column on any unfinished runs after a crash. Those aren't real failures — they're the app falling over mid-flight — so we never alert on them.- An open
cron_heal_attemptsrow exists for this job (statuspendingorspawned) — the self-healing pipeline owns this failure for now, and surfacing both a heal card and an alert card for the same failure would be noisy. Read inside the same transaction as the insert so a heal that lands a microsecond before this insert is still seen. - Per-kind cap reached —
countPendingByKind('cron.failure_alert') >= 5. Five concurrent failed jobs in the queue is already a "your environment is broken" signal; capping at 5 leaves the rest of the global queue free for other approval kinds (cron approvals, automations, recipe runs, etc.). When the cap fires, the producer logscron.failure_alert per-kind cap (5) reached; skipping insert for job <id>to electron-log and skips the insert. - Daily idempotency — the
clientRequestIdis<jobId>-failure-<YYYY-MM-DD>. Once you've cleared a card today — acknowledged it (→approved) or dismissed it with the X key (→rejected) — a same-day re-failure re-enters the producer becausefindOpenAlertByTargetonly seespendingrows, but that tombstone still occupies the daily key in the(client_request_id, action_kind)UNIQUE index. The producer probesfindIdempotentfirst and skips re-insert with an info log:cron.failure_alert daily key already used for job <id> (status=<status>); skipping re-insert. So a cleared card won't re-fire until tomorrow's date. Snoozing (the header clock icon or H) is NOT a clear — it leaves the rowpendingand hidden viainbox_snoozesfor the duration you pick, so same-day re-failures bump the still-hidden row instead of being idempotency-skipped.
The Windows toast respects two additional gates inside notificationService.fireSystemNotification:
- Focus Mode — the toast goes through
focusModeService.evaluatelike every other notification-eligible item. If Focus Mode would suppress (rule not met), the toast is suppressed too. The user can override with the Always fire cron failure alerts setting (cronFailureAlertsAlwaysFire = true), which setsbypassFocusMode: trueon the notification call and pierces the gate. - Global silence and behavior settings —
notificationService.shouldNotify()covers the standard "Silence until" window, "Silence forever" toggle, and the cross-platform muted/idle gates. Same surface every other Omniscio notification uses; nothing special for cron-failure alerts.
The inbox card is independent of Focus Mode and silence — those settings only suppress the toast, not the persistent inbox card. The card always lands as long as the producer-side gates (1–6 above) all pass.
When multiple suppressors apply (e.g. cronFailureAlertsEnabled = false AND a heal is also pending), the earliest in the order wins; later checks never run. So if you flip the master setting off, no toast and no card, regardless of heal state.
A deferred run is not a failure — and never reaches here
A scheduled job can finish a run having done none of its work with nothing broken at all: the box was momentarily busy and refused it, so the job backed off and left the work for its next fire. That run is a deferral, and you get no toast and no card for it.
A script says so by exiting 75 (EX_TEMPFAIL). The engine records the run as skipped — not
failed — with the script's own reason kept on the run row, and a skip never reaches the producer
above: no card, no toast, no error chime, and no paid self-heal session spent debugging a script
that behaved exactly as designed. An in-app executor says the same thing in-band instead of through
an exit code.
75 is the only exit code that means this. Every other non-zero exit, and every deadline kill, is
still a real failure and still alerts you — including a 503 from the spawn route that carries no
retry-later code, which is a broken route rather than a busy moment. So if you expected a card for a
job and did not get one, open its run history: a skipped row names the reason, and anything
reading failed went through the normal gates above.
See deferred-run-exit.mjs for the contract, and cron-script-executor.ts for where the exit code is read.
A failed WAKE is carded only once the outage is sustained
A wake schedule — a session job that wakes an existing session rather than starting a fresh one —
is the one failure that does not card on its first miss. A wake whose fire was refused records its
run failed like any other, and that run row is what the outage is measured from, but the card
waits for BOTH three consecutive failed fires AND thirty minutes dark.
Either half alone would still page you for a blip: three misses fit inside one bad minute on a five-minute heartbeat, and a single miss of an hourly schedule sits dark for a full hour. Together they mean consecutive misses spread over real time.
Below that bar the wake is transient — no toast, no card, no chime; the failure is still on the run row and in the log, so nothing is lost, only the interruption is withheld. At or above it the ordinary cron-failure card appears once, under the same identity as any other job, so a recovery still auto-dismisses it and later fires of the same outage simply bump that one card without re-chiming. Once the wake is sustained, the crew that owns the woken session is told as well — its lead, or the crew's own board when the lead is the session that went dark — and told again only every six hours while the outage lasts, so a long outage does not become a chorus.
A non-wake job is unaffected: it still cards on its very first failure. And a wake whose run history cannot be read stays silent — a missed card is the cheaper mistake.
See wake-outage.ts for the predicate and cron-failure-alert.ts for where it gates the card.
Settings
Both settings live at Settings → Notifications, defined as optional booleans on AppSettings in src/shared/types.ts (search cronFailureAlerts). The UI is in src/renderer/src/features/settings/sections/notifications/NotificationSettings.tsx (search cron-failure-alerts-enabled). Search-index entries: cron-failure-alerts-enabled and cron-failure-alerts-always-fire in settings-search-index.ts.
| Setting | Default | What it does |
|---|---|---|
cronFailureAlertsEnabled |
true |
Master toggle. When off, both the Windows toast AND the inbox card are suppressed. Existing pending rows stay in the inbox until the user acknowledges them; future failures of any job stop alerting. |
cronFailureAlertsAlwaysFire |
false |
Bypass Focus Mode batching for the OS toast. When on, the producer passes bypassFocusMode: true so the toast pierces Focus Mode's count/time threshold and fires immediately. Does NOT bypass silenceUntil, the muted gate, or the master off. |
The "Always fire" nested toggle is gated on the master being true — toggle the master off and the nested row is hidden.
Interaction with self-healing
When self-healing is enabled per-job and a heal lands in pending or spawned state, the failure-alert producer suppresses the alert insert for the same job (Skip condition 4 above). The heal pipeline owns the failure for the duration of its lifetime. The reasoning: a heal card in the inbox already tells the user "this job failed and Omniscio is trying to fix it," and a failure card next to it would be noise telling them the same thing twice.
The opposite direction also holds: a cron.failure_alert card in the inbox does not suppress heal creation. The order in notifyJobResult is insertOrBumpFailureAlert → createHealAttemptIfEligible, so the alert insert sees pre-existing heal attempts (the ones it must defer to) but the subsequent heal-eligibility gate sees no alert at all (alerts are not on the heal pipeline's radar).
If self-healing is off for a job, every failure produces an alert card; the heal hook never fires. If self-healing is on but the heal eligibility gate rejects (e.g. 3-strike cap, system-marked failure, one-off mode), no heal lands, so no heal card — and the alert card fires normally.
The detail pane's inline Turn on auto-fix control is the user-facing bridge from "I see the alert and want to set up auto-recovery" to actually turning the feature on — it flips healingEnabled: true on the job in place (no trip to the editor) and confirms with a success toast. It's only shown when the job currently has healingEnabled === false. Separately, the Fix with AI footer button opens a Start-session dialog to fix the job by hand right now — a normal Claude session pre-filled with the error (it pauses the job while you work). It's the hands-on counterpart to automatic self-heal; the sibling escalation card's own Fix it is what still triggers a one-off manual heal (see cron-self-healing.md § "Manual Fix it").
See cron-self-healing.md for the full heal pipeline flow.
For agents
Files (for agents with repo access)
Producer (main process)
- src/main/services/cron/cron-engine-notify.ts —
notifyJobResult+notifyPermanentFailure: the success/retry/permanent-failure notification arm that bumps the inbox card (insertOrBumpFailureAlert), gates the chime (showSystemNotification), and auto-dismisses on success (autoDismissOpenAlertOnSuccess). The engine keeps thin same-named delegates in cron-engine-service.ts that thread its injectedemitPush; the three alert helpers themselves are free functions in cron-failure-alert.ts. - src/main/services/notification-service.ts —
fireSystemNotification(category: 'cron-failure', opts)(line 619). - src/main/db/queries-cli-pending.ts —
findOpenAlertByTarget,bumpAlertPayload,findIdempotent,insertPending,markApproved,countPendingByKind. - src/main/db/migrations/20260715051230-resolve-orphaned-cron-failure-alerts-on-job-delete.ts — the
AFTER DELETE ON cron_jobstrigger + one-time backfill that close a deleted job's open failure/heal cards (invariant I9). TheCRON_JOB_DELETEhandler in src/main/ipc/cron-job-handlers.ts emitsCLI_PENDING_CHANGEDafterward so the inbox drops the closed card live.
Schema and types
- src/shared/cli-pending-types.ts —
CronFailureAlertPayloadinterface,CLI_PENDING_MAX_OPEN_PER_KIND_CRON_FAILURE = 5, the'cron.failure_alert'actionKind inCliActionKind, and the per-kind status semantics docblock. - src/shared/ipc-schemas.ts —
cronFailureAlertPayloadSchema(line 4788) — Zod payload validated at insert time AND at approve-time revalidation in the dispatcher. - src/shared/ipc-channels/index.ts —
IPC.CRON_FAILURE_ALERT_CHANGED = 'cron:failure-alert-changed'push channel. Payload shape:{ jobId: string, action: 'inserted' | 'updated' | 'auto-dismissed' }. - src/shared/types.ts —
cronFailureAlertsEnabledandcronFailureAlertsAlwaysFireonAppSettings(line 1389), defaults at line 2123.
Dispatcher (approve / reject)
- src/main/services/cli/cli-pending-dispatcher.ts —
'cron.failure_alert'arms in both the payload-revalidation switch (line 333) and the dispatch switch (line 681). The dispatch arm is intentionally a no-op: acknowledge is justmarkApproved, no downstream side effect to invoke. The dispatcher then callsmarkDispatchedso the orphan reconciler doesn't pick it up as crashed-mid-dispatch on next boot. Reject flips the row torejectedwith the supplied reason — the X-key / generic dismiss path (snoozing, from the header clock or H, never rejects; it routes through the universal snooze palette instead, which leaves the rowpendingbut snoozed).
Renderer
- src/renderer/src/features/cli-pending/CronFailureAlertPane.tsx — inline detail pane (
ApprovalPaneShell) with the action buttons; each hotkey is advertised per-button viatitle+aria-keyshortcuts(a hover key-cap), not a static footer hint. The inbox row preview is a generic cli-pending row carrying thefailureAlertCarddiscriminator (see below); the whole row is the click target and routes through the inbox dispatcher (activateUnifiedItem→setActiveApproval({ kind: 'cli-pending', id })). There is no footer Dismiss button — the shared top-right Archive removes the card; Snooze dispatches the universalsnooze-entityevent and the pane auto-closes once the row is snoozed (subscribesuseInboxSnoozeStore+isInboxItemSnoozed). Registers its N/P/O/H keys (plusSwhenever the Auto-fix row is shown) via approval-pane-hotkeys.ts, consulted by the approval intercept in useKeyboardShortcuts.ts. Fix with AI opens the sharedStartSessionDialog(components/inbox/StartSessionDialog.tsx) pre-filled with the job's error (launch sourcecron-failure-alert-start-session); the dialog's optionalonLaunchedfires on a successful launch, and the pane then pauses the job (toggleJob) + resolves the alert (resolveInboxApprovalapprove) + toasts (addToast, grandfathered/baselined). The sibling CronHealEscalationView.tsx escalation view keeps its own footer (Fix it / View details / Dismiss — no Pause, its job is already paused; noS) and still routes Fix it throughuseCronStore.healNow→IPC.CRON_JOB_HEAL_NOW. - src/renderer/src/features/settings/sections/cli-pending-approval/CliPendingApprovalModal.tsx — dispatches the active cli-pending row to
CronFailureAlertPanewhenrow.actionKind === 'cron.failure_alert'; falls through to the generic JSON-payload pane otherwise. - src/renderer/src/stores/cli-pending-approval-items.ts — wires the
failureAlertCarddiscriminator onto eachUnifiedInboxItemderived from acron.failure_alertrow, so the inbox row renderer mounts the failure-alert card preview variant. The discriminator is preview-only — activation always goes through the standard cli-pending approval path.
Feature events (declared in registry)
src/shared/feature-registry/index.ts declares three event ids:
cron_failure_alert_inserted(allowList:jobIdHash,isUpdate,failureCount)cron_failure_alert_acknowledged(allowList:jobIdHash)cron_failure_alert_snoozed(allowList:jobIdHash)
The jobIdHash field is a stable SHA-256 prefix over the job UUID — non-PII, lets analytics correlate inserted → acknowledged → snoozed without persisting raw ids.
Tests
- tests/unit/stores/session-navigation-cron-failure-alert.test.ts and tests/unit/shared/cron-failure-alert-payload.test.ts — the card's navigation and payload shape.
- tests/unit/services/cron-engine-alert-insert.test.ts —
insertOrBumpFailureAlertskip-condition matrix. - tests/unit/services/cron-engine-alert-dismiss.test.ts —
autoDismissOpenAlertOnSuccesspush-on-update gating. - tests/unit/services/cli-pending-dispatcher-cron-failure.test.ts — dispatcher's no-op approve arm + reject-with-reason snooze path.
- tests/unit/db/queries/queries-cli-pending-alert.test.ts —
findOpenAlertByTarget/bumpAlertPayload/findIdempotent/markApprovedSQL behavior. - tests/unit/db/migration-20260715051230-resolve-orphaned-cron-failure-alerts.test.ts — the resolve-on-delete trigger closes the deleted job's cards (and spares other jobs / non-cron kinds / null targets), the one-time backfill, missing-table guards, and idempotency. The
CLI_PENDING_CHANGEDemit is asserted in tests/unit/cron-job-handlers.test.ts. - tests/unit/features/inbox/inbox-cron-failure-visibility.test.tsx — inbox row visibility under various pending-row states.
- tests/unit/features/cli-pending/CronFailureAlertPane.test.tsx — the footer actions (Fix with AI opens the pre-filled
StartSessionDialog; launching it pauses viatoggleJob(id, false)+ resolves (approve) + toasts; cancelling the dialog does neither; Fix disabled whenhealPending; Pause →toggleJob(id, false)+ toast; View job), parse-failure fallback, recent-runs strip, heal-pending banner, row-id swap reset, the Auto-fix toggle (reflectshealingEnabled, flips BOTH ways, honest ON/OFF state), plus: the footer hotkeys advertised on the buttons (no static hint line) with Fix onN, h/n/p/o/s hotkey registration (S registered whenever the Auto-fix row shows — including when healing is already on — and omitted while a fix runs), unmount clears the registry, auto-close-on-snooze, the honest Auto-fix state (no false "handled" check, row hidden while a heal is in flight), the truthful recent-run summary (summarizeRecentFailures), the error empty-state, and the Job field surfacing the full humanized job name (slug → Title Case with the raw slug on hover; a plain name verbatim). - tests/unit/hooks/approval-pane-hotkeys.test.ts — the pane-hotkey registry (set/get/clear, case-insensitive, last-write-wins).
- tests/unit/hooks/useKeyboardShortcuts-approval.test.ts — the approval intercept: Enter/X plus the pane-registered keys fire, H wins over
snoozeSession, guards (input/repeat/modifier), cleared-registry no-op. - tests/unit/stores/cli-pending-store-cron-failure.test.ts —
cliPendingSelectItemsdecoration offailureAlertCard. - tests/unit/stores/session-navigation-cron-failure-alert.test.ts — failure-alert rows route through
setActiveApproval({ kind: 'cli-pending', id })like every other cli-pending row (single activation path shared by click + arrow-key + swipe + archive auto-advance). - tests/unit/shared/cli-pending-types-cron-failure.test.ts — payload type narrowing.
- tests/unit/shared/feature-registry-cron-failure.test.ts — registry shape for the three event ids.
Where to look when it goes wrong
- electron-log. The producer logs three notable lines under
[cron-engine]:Job "<name>" failed permanently: <error>— every final failure logs this once. Missing this means the failure didn't reachnotifyJobResultat all (engine crash, mid-flight tick, etc.).cron.failure_alert per-kind cap (5) reached; skipping insert for job <id>— five+ pending alert rows means the queue is full. Dismiss or snooze older rows.cron.failure_alert daily key already used for job <id> (status=<status>); skipping re-insert— the daily idempotency key is occupied. Status will berejectedfor a previously-snoozed row orapprovedfor a previously-acknowledged one. The card won't re-fire until tomorrow's date.insertOrBumpFailureAlert failed <error>— the producer's outer catch fired. Should be very rare — only DB-corruption / SQLite-died conditions reach this branch. The engine tick continues regardless.
cli_pending_actionstable. FilterWHERE action_kind = 'cron.failure_alert'to see every alert row.target_idis the failing job's id;payload_jsonis theCronFailureAlertPayloadJSON;statustells youpending(visible in inbox),approved(acknowledged or auto-dismissed), orrejected(snoozed).AMC_INSTANCE_ID. If alerts mysteriously aren't firing, check that you're running the default install —npm run devwithout env overrides — and not a sandbox or e2e instance, which suppress the entire pipeline.
Related
The opt-in pipeline that can repair the job — and that suppresses this alert while it is working — is on the cron self-healing page. The inbox queue these cards share is on the CLI pending actions page, the batching gate the toast passes through is on the Focus Mode page, and the global silence and mute rules it respects are on the notifications and silence page.
- cron-self-healing.md — opt-in heal pipeline that can suppress the failure alert when active.
- cli-pending-actions.md — the inbox queue these alerts share with cron / automation / project-delete approvals.
- focus-mode.md — the batching gate the OS toast goes through (the inbox card is unaffected).
- notifications-and-silence.md — global silence, sound, and mute behavior the OS toast respects.
- create-cron-job-with-ai.md — how cron jobs are created in the first place.
Last verified 2026-10-05