---
title: Cron self-healing
---

# Cron self-healing

## What it is

Omniscio's cron scheduler runs scripts and recipes on a fixed schedule. When one of those jobs fails for the last time (after every retry has been exhausted), Omniscio hands the failure to a fresh Claude session that has been pre-loaded with the failing job's command, schedule, exit code, the last 50 lines of stderr and stdout, and a list of regex hits like `ENOENT` or `ECONNREFUSED`. The heal session's job is to figure out what broke, edit the project files to fix it, verify the fix, and resume the schedule. Verification is **side-effect-aware**: the on-demand `/run` is a real production execution (real environment, real side effects — not a dry run), so the heal prompt tells the agent to re-run only jobs that are safe to repeat (idempotent, no external side effects) and to verify a side-effect-bearing job — one that sends email, fires a webhook, calls an API, moves money, or pushes a commit — _without_ re-executing it, so an auto-fired heal can't repeat those irreversible effects (audit finding F001). The cron stays paused the whole time, so a broken job cannot keep firing alongside the agent that's trying to fix it.

The feature is **on by default for every job** (since 2026-08-21 — new jobs default on via `createCronJob`, and every existing job was backfilled on by migration `20260821175646`). A job opts OUT by turning Self-Healing off (`healingEnabled: false`); session-type jobs never heal regardless (Guard 0). There is **one** toggle: Self-Healing on means heal sessions auto-fire on every final failure of this job — no inbox approval, no category gate, no second toggle to set. The reasoning is the user's words from 2026-05-11: "self-healing that requires approval is not self-healing". Heal sessions can edit code in the project but cannot rewrite the cron's own definition or touch other jobs, and after three consecutive failed heals Omniscio stops trying and posts an escalation card so you notice.

> **2026-05-11 redesign.** Earlier versions of this feature shipped a nested **Auto-fire** toggle and an inbox approval card so the user could review each heal before it spawned. The Approve button was broken (the dispatcher arm threw on the legacy payload shape it had stopped emitting) and — independent of the bug — the design itself missed: an approval-gated heal that won't run until a human clicks is not self-healing. The toggle, the pane's Approve/Reject buttons, the bug, and the entire approval-gated code path were removed in one change. The DB column `cron_jobs.healing_auto_fire` is preserved for backward compatibility but is no longer read or written — Guard 1 reads `healing_enabled` only. Pre-redesign `pending` heal rows in `cli_pending_actions` are auto-fired by the `autoFireExistingPendingHeals` startup migration so users don't lose any in-flight work.

## Where to find it

The switch lives on the job itself. In Omniscio's **Cron Jobs** virtual project, open the job you
care about to bring up its editor dialog, and Self-Healing sits below **Require Approval** in that
same form. Everything after that happens without you: the heal session shows up as an ordinary
session in the project the job belongs to, and the job's row in the Cron Jobs view turns paused
while it works.

## How it behaves

### Interaction with cron failure alerts

Cron failure alerts ([cron-failure-alerts.md](cron-failure-alerts.md)) are the other half of the same story: when a cron job fails for the last time, Omniscio always considers BOTH (a) inserting a `cron.failure_alert` row in the inbox so you see "this job broke" and (b) creating a `cron.failure_heal` attempt so an agent can fix it. They are two views of the same failure — the alert tells you, the heal fixes it — and they coordinate so you don't see both for the same incident.

- **An open heal suppresses the alert card.** When a `cron.failure_heal` heal-attempt row exists in `pending` or `spawned` for the same job, the failure-alert producer's skip conditions short-circuit and no alert card is inserted. The reasoning: if Omniscio has already escalated to "an agent is going to fix this", surfacing a separate "this is broken" card is duplicate noise. The OS toast still fires (the toast is the user's "something happened" cue, regardless of who's handling it) but the inbox stays clean. (After the 2026-05-11 redesign, every new heal attempt goes straight from `pending` to `spawned` in the same atomic transaction — there is no human-gated `pending` window — so the practical effect of this rule is "spawned heal suppresses alert".)
- **Once the heal closes, alerts are eligible again.** When the heal ends — fixed, abandoned, or escalated — the suppression lifts. If the job fails again on a future run, the alert card is allowed to insert.
- **Bridge from alert to fixing.** The alert card offers two routes. (a) **Fix with AI** — opens a Start-session dialog to fix the job by hand right now (a normal Claude session pre-filled with the error; it pauses the job while you work), NOT the `healNow` manual-heal path. (b) An inline **Auto-fix** toggle flips `healingEnabled` **in place** with a success toast, so future failures go through the auto heal pipeline (this is the route that actually engages self-healing). (**View job** navigates to the Cron Jobs editor; both routes above stay in the inbox.)

### How to use it

1. **Open the job's editor.** In Omniscio's **Cron Jobs** virtual project, click any job to bring up its editor dialog. Self-healing lives below "Require Approval" in the same form.
2. **Leave Self-Healing on (or flip it off to opt out).** The toggle is labelled **Self-Healing** with the description "Detect failures and auto-launch a heal session to fix them." **On by default** — every new job self-heals and all existing jobs were backfilled on; turn it off for a specific job you don't want auto-fixed. There is no second toggle: Self-Healing on means the heal session auto-fires on every final failure.
3. **Save and let it run.** When the job next fails for good (i.e. after every retry has been exhausted), Omniscio pauses the cron and spawns the heal session into the project the cron belongs to. The cron's row in the Cron Jobs view will switch to paused. There is no approval card and nothing to click — the heal is already running.
4. **Watch for the resume.** The heal session has been told the cron is paused and needs to call a `/toggle` endpoint to resume the schedule once the fix verifies green. When that happens Omniscio marks the heal "fixed" and resets the consecutive-failure counter to zero — back to normal scheduled runs.
5. **Look for the escalation card after three strikes.** If three consecutive heals abandon (the agent gave up, the session was archived without resuming the cron, the app was restarted mid-heal, etc.), the gate stops creating new heal attempts and Omniscio posts a single inbox card — `<job name> needs human attention`, `3 failed self-heals` — so you take over manually. Approving the escalation card is acknowledge-only — it dismisses the card without spawning a heal session, because the cap means the agent has already proven it can't fix this one without you. **The card is not permanent, though: if the job's next scheduled run simply SUCCEEDS on its own, Omniscio treats that green run as proof it recovered — it resets the consecutive-failure counter to zero (re-arming self-healing) and auto-dismisses the escalation card**, exactly as a successful heal would (`clearHealEscalationOnSuccess`; see § "Recovery on a successful run"). So a job whose underlying problem cleared up — a transient outage, or a heal that was only interrupted by an app restart — stops nagging you the moment it runs clean again, instead of leaving behind a stale "needs attention" card while every run goes green.

The cron-job editor dialog UI lives in [src/renderer/src/features/cron/JobEditorDialog.tsx](../../src/renderer/src/features/cron/JobEditorDialog.tsx). After the 2026-05-11 redesign there is a single Self-Healing toggle row — the previously-nested **Auto-fire** toggle row was removed.

## For agents

### How it works

#### The trigger

Every cron run goes through the engine's `notifyJobResult` callback. When a run lands in `failed` after exhausting `retryCount` retries, the engine fires a system notification and a `CRON_RUN_COMPLETED` push event, then calls `createHealAttemptIfEligible(job, freshRun)` as fire-and-forget so the engine tick is never blocked by the heal pipeline. See `notifyPermanentFailure` in [src/main/services/cron/cron-engine-notify.ts](../../src/main/services/cron/cron-engine-notify.ts) (the engine keeps a thin `notifyJobResult` delegate in [cron-engine-service.ts](../../src/main/services/cron/cron-engine-service.ts)).

#### The eligibility gate

Before doing anything, the orchestrator in [src/main/services/cron-heal-orchestrator.ts](../../src/main/services/cron/cron-heal-orchestrator.ts) walks six guards in order — the first failing guard short-circuits and no heal attempt is created:

1. **`job.healingEnabled === false`** — healing defaults ON, so this short-circuits only a job that has been explicitly opted OUT (`healingEnabled: false`).
2. **`run.systemMarkedFailure === true` AND the run carries no failure evidence** — `system_marked_failure` has TWO writers with different meaning. The startup reconciler (`recoverStaleRuns`) flips runs still `running` after a crash — no real outcome, no `failure_evidence_json` — and healing "the app crashed" is meaningless, so those are skipped. But the zombie-run watchdog (`reapOverdueRuns`) also sets the flag when it ends a run that overran its timeout; if that run's executor settled (even late) it carries REAL failure evidence, which is a genuine failure the user expects self-healing to engage. So only the **evidence-less** case is skipped. The discriminator is airtight: `failure_evidence_json` and `status='failed'` are written in one executor transaction, so a still-`running` crash-orphan can never carry evidence. (This is why a job that fails by **overrunning its time limit** now self-heals instead of being silently skipped — the 2026-06-22 postmortems-archive failure that motivated the fix.)
3. **`job.runMode === 'one_off'`** — v1 excludes one-off jobs. A one-off has fired its single intended occurrence, so "fixed" doesn't have a clean meaning.
4. **`job.consecutiveHealFailures >= 3`** — the cap. Three abandoned heals in a row means the agent isn't getting it; we stop looping and insert the escalation inbox card instead. The escalation card is a `cli_pending_actions` row with `actionKind: 'cron.failure_heal'` and payload `{escalation: true, jobId}` — distinct from the legacy `{healAttemptId}` shape that the autoFireExistingPendingHeals migration drains. The dispatcher's `cron.failure_heal` arm branches on `'escalation' in payload`: escalation rows are acknowledge-only (no spawn).
5. **An open heal already exists for this job** — `getOpenHealForJob` returns any heal in `pending` or `spawned` state. One in flight at a time, job-scoped on purpose.
6. **The fleet-wide open AUTO-heal count has hit the cap** (`countOpenAutoHeals() >= MAX_CONCURRENT_AUTO_HEALS`, default 3) — the GLOBAL bound (F119). Guards 4 and 5 are _per job_; a fleet-wide failure (a bad deploy, an expired shared credential, a network outage) fails many DISTINCT jobs at once, and each independently passes its own per-job gate — so without this nothing limits the TOTAL concurrent paid heal spawns. Over the cap the heal is deferred (it re-fires on the job's next failed run once an in-flight heal terminates and frees a slot). It counts only AUTO heals (`is_manual = 0`); the manual "Fix it" path (the escalation card) is user-initiated and exempt. The count is a snapshot, so a brief overshoot under a concurrent burst is accepted (a cost-pacing guard, not a correctness one — same class as NightyTidy2's cap-accuracy race).

Past the gate, evidence is collected (or re-read from the run's `failure_evidence_json` if the engine populated it) and classified into one of `timeout`, `network`, `missing-file`, `permission`, `script-error`, `recipe-step-failed`, or `unknown`. The classifier is a regex panel — see [src/main/services/cron-failure-evidence.ts](../../src/main/services/cron/cron-failure-evidence.ts).

#### Always auto-fire (post 2026-05-11)

There is only one path now. After the eligibility gate accepts a failure the orchestrator inserts a `cron_heal_attempts` row in `pending` and immediately calls `startCronHealSession(healAttemptId)` — no `cli_pending_actions` row, no inbox card, no Approve button. If the spawn itself fails the row is marked `abandoned` with reason `auto-fire-spawn-failed` so the cap counter advances and you don't end up with a phantom `spawned` row. The category classifier still runs (the regex panel in [cron-failure-evidence.ts](../../src/main/services/cron/cron-failure-evidence.ts)) because the escalation card's preview text reads "3 failed self-heals — last category: `<category>`", but it no longer gates anything.

The `cron_jobs.healing_auto_fire` column is still in the schema for backward compatibility but Guard 1 reads `healing_enabled` only. Saving a job from the editor no longer writes `healing_auto_fire`, so the column drifts to whatever value it had at save time. Inert in practice — nothing reads it anymore — but worth knowing if you're inspecting the DB directly.

`startCronHealSession` (in [cron-heal-orchestrator.ts](../../src/main/services/cron/cron-heal-orchestrator.ts)) builds the prompt from job + last run + evidence + last 10 runs of history via [src/main/services/cron-heal-prompt.ts](../../src/main/services/cron/cron-heal-prompt.ts), spawns a session via `createSessionWithPrompt({ injectInAppToken: true })`, then in a single SQLite transaction marks the heal `spawned` and calls `toggleCronJob(job.id, false)` — atomic, so a crash between those writes can't leave the cron firing alongside a live heal session.

#### Manual "Fix it" (user-triggered override)

The **escalation** card offers a **Fix it** button that spawns a heal on demand — a DELIBERATE user override, distinct from the automatic gate above. It routes `useCronStore.healNow(jobId)` → `IPC.CRON_JOB_HEAL_NOW` → `startManualHeal(jobId)` in [cron-heal-orchestrator.ts](../../src/main/services/cron/cron-heal-orchestrator.ts). (The **failure-alert** card's own **Fix with AI** button no longer uses this manual-heal path — it opens a normal Start-session dialog pre-filled with the error and pauses the job; see [cron-failure-alerts.md](cron-failure-alerts.md).)

`startManualHeal` reuses the SAME prompt + spawn as the auto path — both call the shared `buildAndSpawnFixSession(job, run, evidence)` helper, so the two entry points can't drift on prompt shape or spawn args. That shared helper also **resolves the target project's real DB id before spawning**: a project-less job (`job.projectId` null) heals into the Claude workspace, and the fallback maps the `__claude__` folder_path **sentinel** to the Claude project's real row id via `getProjectByFolderPath(CLAUDE_PROJECT_ID)?.id` (mirroring `email-inbound-service` / `workflow-engine` / `stuck-task-helper`). Passing the raw sentinel as the projectId made `createSessionWithPrompt`'s `getProject('__claude__')` return null and throw `Project not found (__claude__)`, silently breaking BOTH the auto heal AND the manual **Fix it** for every project-less job (surfaced to the user as the toast "Could not start the fix session. Please try again."; fixed 2026-07-08). `startManualHeal`'s eligibility differs from the auto gate on purpose:

- **Bypasses auto Guard 1 (`healingEnabled === false`) and Guard 4 (the 3-strike cap).** The user is explicitly asking, so a manual fix runs even when auto-fix is off or has already given up — exactly the state the escalation card is in. (Without this, "Fix it" would silently no-op in the states the user most wants it — the same class as the old "Approve Heal does nothing" bug.)
- **Keeps the sane refusals:** a `session`-type or `one_off` job has no script/recipe to repair (friendly reject); `getOpenHealForJob` still blocks a second fix while one is in flight (auto or manual); and it needs a real failed run (`getLatestFailedRun`) to diagnose.
- **Records an `is_manual` `cron_heal_attempts` row (the fix for the strand-forever bug).** After the spawn, `startManualHeal` (via `recordManualHealAndPause`) inserts a heal_attempt marked `is_manual` and — ONLY if the row was actually written — calls `markHealSpawned` to record `session_id` + pause the cron (a `CRON_JOB_UPDATED` push follows). That row is what makes the manual fix RECOVERABLE: the three un-pause paths (session archive, startup reconciler, hung-reaper) AND the in-app `/run` + `/toggle` auth scope all key on it — before this, a manual fix left the job paused forever and its own fix session got 403 trying to resume. The `is_manual` flag keeps it OUTSIDE the auto-heal escalation accounting: the archive + hung-reaper abandons pass `countsAsFailure:false`, so a manual fix never pushes the job toward the 3-strike cap. On a rare `run_id` collision (the latest failed run already carries a heal_attempt) the insert no-ops and the job is left RUNNING rather than paused-without-a-record — `markHealSpawned`'s cron-pause is unguarded, so pausing a phantom row would re-strand it.
- Returns a discriminated `ManualHealResult` (`{ ok, sessionId }` | `{ ok: false, reason }`) so the `CRON_JOB_HEAL_NOW` handler surfaces a friendly, already-humanized reason to the renderer.

Invariants locked by [tests/unit/cron-heal-spawn.test.ts](../../tests/unit/cron-heal-spawn.test.ts) (`startManualHeal` suite): overrides Guards 1 & 4, refuses session / one-off / no-failed-run / already-running, records an `is_manual` `spawned` row + pauses, spawns-WITHOUT-pausing on a `run_id` collision, and — with the hung-reaper case in [cron-heal-reconciler.test.ts](../../tests/unit/cron-heal-reconciler.test.ts) — never counts a manual heal toward the cap. See [cron-heal-idempotency-contract.md](../../.claude/memory/contracts/cron-heal-idempotency-contract.md) (I5).

#### Startup migration for pre-redesign pending heals

Users running the pre-2026-05-11 build with Auto-fire off had `cron.failure_heal` rows sitting in `cli_pending_actions.status = 'pending'` waiting on an Approve click. The Approve button never worked (the dispatcher arm has not accepted that legacy payload shape since the redesign), and even if it did the new design would still convert them to direct spawns. So `autoFireExistingPendingHeals` runs once at startup ([cron-heal-orchestrator.ts](../../src/main/services/cron/cron-heal-orchestrator.ts)) and:

1. Selects every `cli_pending_actions` row where `action_kind = 'cron.failure_heal' AND status = 'pending'`.
2. Skips any row whose payload is the new escalation shape (`{escalation: true, jobId}` — has no `healAttemptId` to spawn against).
3. For each remaining legacy `{healAttemptId}` row: marks the row `dispatched` (`markApprovedAndDispatched`) BEFORE spawning, then calls `startCronHealSession`. If the spawn throws the row stays `dispatched` (already approved) but the underlying heal_attempt is `markHealAbandoned`'d so the cap counter still advances.
4. After all rows are processed, emits a single coalesced `CLI_PENDING_CHANGED` push so the renderer refreshes the inbox once.

This is one-shot drainage — once the queue is empty for legacy rows, the migration is a no-op on subsequent boots.

#### What the heal session can do

The prompt block tells the heal session it has two HTTP endpoints, scoped to its own cron job, authenticated with the in-app bearer token that was injected into the prompt:

- `POST /cron/jobs/<this job's id>/run` — execute the cron one time on demand and return the exit code + output.
- `POST /cron/jobs/<this job's id>/toggle` with body `{ "isActive": true }` — resume the schedule.

The CLI server in [src/main/services/cli/cli-server-cron-routes.ts](../../src/main/services/cli/cli-server-cron-routes.ts) verifies on every call that the request comes from a session that owns an open `spawned` heal for this job (`classifyCronJobAuthScope`). Any other in-app session calling these endpoints for this job gets a `403 session not authorized for this job`. The user's external CLI control token continues to work normally.

The heal session is a regular Claude session in the cron's project, so it can also `Read` and `Edit` files in the project's working directory — that's how it actually fixes things.

#### What the heal session cannot do

- **Rewrite the cron's definition.** The heal session's bearer token only authorizes `/run` and `/toggle` for this specific job. It cannot edit the cron expression, change the command, or alter any of the cron's other fields. Updating the cron definition still requires you in the Omniscio UI.
- **Touch other jobs.** Other jobs' `/run` and `/toggle` endpoints reject this session's bearer with 403.
- **Loop indefinitely.** After three consecutive abandoned heals (`job.consecutiveHealFailures >= 3`) Guard 4 stops creating new heal attempts and posts the escalation card.

#### Loop prevention — three layers

1. **Job-scoped open-heal guard** (Guard 5, [cron-heal-orchestrator.ts:160](../../src/main/services/cron/cron-heal-orchestrator.ts#L160)) — `getOpenHealForJob` blocks a new heal while any existing heal for the same job is still `pending` or `spawned`. A run-scoped UNIQUE constraint on `cron_heal_attempts(run_id)` is the backstop.
2. **Consecutive-failure cap of 3** (Guard 4, [cron-heal-orchestrator.ts:151-154](../../src/main/services/cron/cron-heal-orchestrator.ts#L151-L154)) — `cron_jobs.consecutive_heal_failures` increments on every failing `markHealAbandoned` and resets to 0 on `markHealFixed` (both wrapped in a single transaction with the heal-attempt status flip — see [queries-cron-heal-attempts.ts:108-147](../../src/main/db/queries-cron-heal-attempts.ts#L108-L147)) — AND, since the recovery fix, on any successful scheduled RUN (`clearHealEscalationOnSuccess`; see § "Recovery on a successful run"). Once it hits 3, Guard 4 short-circuits and writes the escalation inbox card instead of creating a new heal attempt — until the counter is cleared by one of those paths.
3. **Startup reconciler that abandons spawned heals across restarts** ([src/main/services/cron-heal-reconciler.ts:25-42](../../src/main/services/cron/cron-heal-reconciler.ts#L25-L42)) — the in-app HMAC secret rotates on every process boot, so any session token from a prior run can no longer authenticate to `/run` or `/toggle`. On startup Omniscio abandons every `spawned` heal with reason `app restart invalidated session token` and toggles the parent cron back to active so the next scheduled run fires normally. Pending heals older than the inbox display TTL are also abandoned with reason `TTL expired`.

#### Recovery on a successful run

The three layers above STOP a runaway heal loop; this is the complementary UN-stick. A job can hit the 3-strike cap and then recover on its OWN — the underlying problem was transient (a flaky network, an outage that passed), or its heals were only abandoned by app restarts (a lifecycle abandon, which does not itself mean the job is broken). Before this fix, nothing cleared the give-up counter or the escalation card in that case: the counter reset ONLY inside `markHealFixed` (a heal session verifying the fix and calling `/toggle`), so a job that quietly started passing again stayed capped forever, its escalation card falsely reading "keeps failing / paused" while every run went green (the 2026-08-07 stale-card incident).

`clearHealEscalationOnSuccess` ([cron-heal-escalation.ts](../../src/main/services/cron/cron-heal-escalation.ts)), called from `notifyJobResult`'s success branch ([cron-engine-notify.ts](../../src/main/services/cron/cron-engine-notify.ts)), closes the gap: a green run is treated as ground-truth proof the job is healthy, so it resets `consecutive_heal_failures` to 0 (re-arming self-healing) AND dismisses any open `cron.failure_heal` escalation/wedged card for the job (`dismissOpenActionsByTarget`) — in ONE transaction so a crash can't leave the counter cleared with the card stranded. A healthy job (counter already 0) is a no-op; an isolated secondary instance (`AMC_INSTANCE_ID`) skips entirely. It is the exact success-branch mirror of the failure alert's `autoDismissOpenAlertOnSuccess`, and needs no migration — every already-stuck job self-clears on its next green run. See [cron-heal-idempotency-contract.md](../../.claude/memory/contracts/cron-heal-idempotency-contract.md) (I6).

#### Inbox card variants

The inbox dispatches `actionKind === 'cron.failure_heal'` rows to [src/renderer/src/features/cli-pending/CronHealApprovalPane.tsx](../../src/renderer/src/features/cli-pending/CronHealApprovalPane.tsx) (`CliPendingApprovalModal` early-returns this pane on that actionKind). Two variants emit in the post-2026-05-11 design (the legacy `pending` heal-attempt variant only appears for rows that pre-date the redesign and is drained on next boot by the `autoFireExistingPendingHeals` migration):

- **`escalation`** (the only variant created post-redesign) — three consecutive failed heals. Payload is `{escalation: true, jobId}`. Reads "`<job>` needs human attention" + "3 failed self-heals." No spawn happens when you approve — the dispatcher's escalation arm just stamps `dispatched_at` and the row exits the queue. The card's job is to make sure you notice.
- **`pending`** _(legacy, migration-only)_ — pre-redesign rows with payload `{healAttemptId}` waiting on the broken Approve button. The startup migration drains these by spawning the heal session directly. They no longer appear in steady state, but the schema still accepts the shape so any row that survives the migration (e.g. crashed mid-drain) renders correctly until the next boot picks it up.

#### The approval pane (post-redesign)

When the user activates any `cron.failure_heal` row, Dashboard renders [CronHealApprovalPane](../../src/renderer/src/features/cli-pending/CronHealApprovalPane.tsx) inline (via the `ApprovalPaneShell` primitive — same shell every other approval pane uses). After the 2026-05-11 redesign the pane has three view modes that branch on the parsed payload shape:

- **Escalation view** (the common case) — payload `{escalation: true, jobId}`. Title reads **This cron job needs your attention**; the body shows the job summary (name, the **humanized schedule** via `humanizeCronExpression` — e.g. `0 */6 * * *` renders as "Every 6 hours", and raw cron only as a fallback when the expression can't be parsed — plus the consecutive self-heal-failure count) and a plain-English "what you can do" note. Three buttons: **Fix it** (`F`, a manual-override heal via `useCronStore.healNow(jobId)` — see § "Manual Fix it"; confirmed first since it spawns a paid session), **View details** (`O`, `activateVirtualProject(CRON_PROJECT_ID)` + selects the job), and the primary **Dismiss** (`Enter`) which calls `resolveInboxApproval('cli-pending', row.id, { kind: 'approve' })`. There is no **Pause** (the escalation card's job is already paused) and no footer **Snooze** (the header clock snoozes; `H` still works). Dismiss is acknowledge-only — the dispatcher branches on `'escalation' in payload` and stamps `dispatched_at` without spawning.
- **Legacy view** (migration safety net) — payload `{healAttemptId}`. Same five-section detail layout the original approval pane shipped with (job summary, failed run, evidence chips + stderr/stdout, recent history). The Approve button is a no-op label that explains the row is being drained by the startup migration; in steady state these rows never reach the user because the migration runs before the renderer mounts the inbox.
- **Invalid view** — payload doesn't match either union arm. Renders a generic error explaining the row's payload couldn't be parsed and offering Reject as the only action. Should never be reachable because `cronFailureHealPayloadSchema` is enforced at insert time, but defensive against schema drift.

Reject across all three views fires `resolveInboxApproval('cli-pending', row.id, 'reject', reason)` from an inline reason textarea, same convention as every other approval pane. `closeDisabled={isResponding}` matches the AgentDrivenSpawn pane convention so the user can't escape mid-IPC and leave the inbox in a stuck state. Reject reason is read from `reasonRef.current` at click time (not React state) to avoid the stale-closure bug where a rapid-typing user's last keystroke is dropped from the dispatched payload.

### Where to look when it goes wrong

- **electron-log.** Heal pipeline events log under `[CronHeal]` (reconciler, session-lifecycle archive hook) and `[cron-heal]` (orchestrator). Search for either prefix in `main.log` (location: see [logs-and-debugging.md](logs-and-debugging.md)). Notable lines:
  - `[CronHeal] Reconciler abandoned <N> spawned + <M> aged-out pending heal_attempts on startup` — startup reconcile result.
  - `[CronHeal] Session <id> archived without resume — heal <id> abandoned, job <id> restored to active` — heal session was archived before the agent called `/toggle` ([session-lifecycle-service.ts:179-187](../../src/main/services/session/session-lifecycle-service.ts#L179-L187)).
  - `[cron-heal] auto-fire spawn failed for job <id>` — auto-fire path could not start the heal session.
  - `[cron-heal] manual heal spawn failed for job <id> <error>` — the manual **Fix it** path's `buildAndSpawnFixSession` threw; surfaces to the renderer as the toast "Could not start the fix session. Please try again." A historical cause (fixed 2026-07-08) was the `__claude__` folder_path sentinel reaching `createSessionWithPrompt` unresolved → `Project not found (__claude__)`; see § "Manual Fix it" for the projectId resolution.
  - `[cron-heal] manual heal for job <id>: run <id> already has a heal_attempt (run_id collision) — spawning the fix session WITHOUT pausing so the job is never stranded` — the rare re-heal of an already-healed run; the fix session still runs but the cron is left active (never paused without a recoverable row).
  - `[cron-heal] manual heal finalize failed for job <id> <error>` — the manual heal recorded its row but `markHealSpawned` threw; the row is abandoned non-counting so it can't soft-block future heals.
  - `[cron-heal] queue at cap, abandoning heal-attempt for job <id>` — the CLI Pending queue is full, so the approval-gated heal could not be queued.
- **`cron_heal_attempts.abandoned_reason` column.** Every abandoned heal records why. The values currently emitted are:
  - `auto-fire-spawn-failed` — auto-fire path's `startCronHealSession` threw.
  - `app restart invalidated session token` — startup reconciler.
  - `TTL expired` — pending heal aged out beyond the inbox display window.
  - `session ended without resume` — you archived the heal session before the agent called `/toggle` (an AUTO heal counts toward the cap here; a MANUAL "Fix it" does not — `is_manual`).
  - `manual-heal-finalize-failed` — a manual **Fix it** recorded its heal_attempt but the follow-up `markHealSpawned` threw; abandoned non-counting so it can't block future heals.
- **Inbox cards.** An escalation card clears three ways: you dismiss it, a manual **Fix it** succeeds, or the job's next scheduled run succeeds on its own (`clearHealEscalationOnSuccess` resets the counter + auto-dismisses the card — § "Recovery on a successful run"). So a card that PERSISTS means the job is still genuinely failing and needs you — a card that lingers while the job's runs are actually green is a bug (that was the 2026-08-07 stale-card incident, now fixed). Pending cards that disappear without a resume usually mean a startup reconcile or a TTL aging — the `abandoned_reason` column will tell you which.

## Related

The other half of the same story — the toast and inbox card that tell you a job broke, which a
running heal deliberately suppresses — is on the [cron failure alerts](cron-failure-alerts.md)
page. How cron jobs are created and where their editor lives is on the
[create a cron job with AI](create-cron-job-with-ai.md) page, the approval queue these cards share
is on the [CLI pending actions](cli-pending-actions.md) page, and finding the logs is on the
[logs and debugging](logs-and-debugging.md) page.

- [cron-failure-alerts.md](cron-failure-alerts.md) — the toast + inbox-card surface that fires when a cron fails and there is no open heal in flight. Suppressed by an open heal. Its **Fix with AI** button opens a Start-session dialog to fix the job by hand (a normal session pre-filled with the error, not this manual-heal path); its inline **Auto-fix** toggle flips self-healing on/off **in place** for the job.
- [create-cron-job-with-ai.md](create-cron-job-with-ai.md) — how cron jobs are created, where the editor lives, what the cron tick does in normal operation.
- [cli-pending-actions.md](cli-pending-actions.md) — the inbox approval queue heal cards share with cron / automation / project-delete approvals.
- [logs-and-debugging.md](logs-and-debugging.md) — finding `main.log` and the in-app Debug Log Viewer.
