Cron self-healing
How Omniscio reacts when a scheduled job fails for good: it pauses the job and hands the failure to a fresh Claude session pre-loaded with the error and logs, so the agent can fix the cause, verify it, and resume the schedule — plus the three-strike cap that stops it looping.
What it is
Omniscio's cron scheduler runs scripts and recipes on a fixed schedule. When one of those jobs fails for the last time (after every retry has been exhausted), Omniscio hands the failure to a fresh Claude session that has been pre-loaded with the failing job's command, schedule, exit code, the last 50 lines of stderr and stdout, and a list of regex hits like ENOENT or ECONNREFUSED. The heal session's job is to figure out what broke, edit the project files to fix it, verify the fix, and resume the schedule. Verification is side-effect-aware: the on-demand /run is a real production execution (real environment, real side effects — not a dry run), so the heal prompt tells the agent to re-run only jobs that are safe to repeat (idempotent, no external side effects) and to verify a side-effect-bearing job — one that sends email, fires a webhook, calls an API, moves money, or pushes a commit — without re-executing it, so an auto-fired heal can't repeat those irreversible effects (audit finding F001). The cron stays paused the whole time, so a broken job cannot keep firing alongside the agent that's trying to fix it.
The feature is on by default for every job (since 2026-08-21 — new jobs default on via createCronJob, and every existing job was backfilled on by migration 20260821175646). A job opts OUT by turning Self-Healing off (healingEnabled: false); session-type jobs never heal regardless (Guard 0). There is one toggle: Self-Healing on means heal sessions auto-fire on every final failure of this job — no inbox approval, no category gate, no second toggle to set. The reasoning is the user's words from 2026-05-11: "self-healing that requires approval is not self-healing". Heal sessions can edit code in the project but cannot rewrite the cron's own definition or touch other jobs, and after three consecutive failed heals Omniscio stops trying and posts an escalation card so you notice.
2026-05-11 redesign. Earlier versions of this feature shipped a nested Auto-fire toggle and an inbox approval card so the user could review each heal before it spawned. The Approve button was broken (the dispatcher arm threw on the legacy payload shape it had stopped emitting) and — independent of the bug — the design itself missed: an approval-gated heal that won't run until a human clicks is not self-healing. The toggle, the pane's Approve/Reject buttons, the bug, and the entire approval-gated code path were removed in one change. The DB column
cron_jobs.healing_auto_fireis preserved for backward compatibility but is no longer read or written — Guard 1 readshealing_enabledonly. Pre-redesignpendingheal rows incli_pending_actionsare auto-fired by theautoFireExistingPendingHealsstartup migration so users don't lose any in-flight work.
Where to find it
The switch lives on the job itself. In Omniscio's Cron Jobs virtual project, open the job you care about to bring up its editor dialog, and Self-Healing sits below Require Approval in that same form. Everything after that happens without you: the heal session shows up as an ordinary session in the project the job belongs to, and the job's row in the Cron Jobs view turns paused while it works.
How it behaves
Interaction with cron failure alerts
Cron failure alerts (cron-failure-alerts.md) are the other half of the same story: when a cron job fails for the last time, Omniscio always considers BOTH (a) inserting a cron.failure_alert row in the inbox so you see "this job broke" and (b) creating a cron.failure_heal attempt so an agent can fix it. They are two views of the same failure — the alert tells you, the heal fixes it — and they coordinate so you don't see both for the same incident.
- An open heal suppresses the alert card. When a
cron.failure_healheal-attempt row exists inpendingorspawnedfor the same job, the failure-alert producer's skip conditions short-circuit and no alert card is inserted. The reasoning: if Omniscio has already escalated to "an agent is going to fix this", surfacing a separate "this is broken" card is duplicate noise. The OS toast still fires (the toast is the user's "something happened" cue, regardless of who's handling it) but the inbox stays clean. (After the 2026-05-11 redesign, every new heal attempt goes straight frompendingtospawnedin the same atomic transaction — there is no human-gatedpendingwindow — so the practical effect of this rule is "spawned heal suppresses alert".) - Once the heal closes, alerts are eligible again. When the heal ends — fixed, abandoned, or escalated — the suppression lifts. If the job fails again on a future run, the alert card is allowed to insert.
- Bridge from alert to fixing. The alert card offers two routes. (a) Fix with AI — opens a Start-session dialog to fix the job by hand right now (a normal Claude session pre-filled with the error; it pauses the job while you work), NOT the
healNowmanual-heal path. (b) An inline Auto-fix toggle flipshealingEnabledin place with a success toast, so future failures go through the auto heal pipeline (this is the route that actually engages self-healing). (View job navigates to the Cron Jobs editor; both routes above stay in the inbox.)
How to use it
- Open the job's editor. In Omniscio's Cron Jobs virtual project, click any job to bring up its editor dialog. Self-healing lives below "Require Approval" in the same form.
- Leave Self-Healing on (or flip it off to opt out). The toggle is labelled Self-Healing with the description "Detect failures and auto-launch a heal session to fix them." On by default — every new job self-heals and all existing jobs were backfilled on; turn it off for a specific job you don't want auto-fixed. There is no second toggle: Self-Healing on means the heal session auto-fires on every final failure.
- Save and let it run. When the job next fails for good (i.e. after every retry has been exhausted), Omniscio pauses the cron and spawns the heal session into the project the cron belongs to. The cron's row in the Cron Jobs view will switch to paused. There is no approval card and nothing to click — the heal is already running.
- Watch for the resume. The heal session has been told the cron is paused and needs to call a
/toggleendpoint to resume the schedule once the fix verifies green. When that happens Omniscio marks the heal "fixed" and resets the consecutive-failure counter to zero — back to normal scheduled runs. - Look for the escalation card after three strikes. If three consecutive heals abandon (the agent gave up, the session was archived without resuming the cron, the app was restarted mid-heal, etc.), the gate stops creating new heal attempts and Omniscio posts a single inbox card —
<job name> needs human attention,3 failed self-heals— so you take over manually. Approving the escalation card is acknowledge-only — it dismisses the card without spawning a heal session, because the cap means the agent has already proven it can't fix this one without you. The card is not permanent, though: if the job's next scheduled run simply SUCCEEDS on its own, Omniscio treats that green run as proof it recovered — it resets the consecutive-failure counter to zero (re-arming self-healing) and auto-dismisses the escalation card, exactly as a successful heal would (clearHealEscalationOnSuccess; see § "Recovery on a successful run"). So a job whose underlying problem cleared up — a transient outage, or a heal that was only interrupted by an app restart — stops nagging you the moment it runs clean again, instead of leaving behind a stale "needs attention" card while every run goes green.
The cron-job editor dialog UI lives in src/renderer/src/features/cron/JobEditorDialog.tsx. After the 2026-05-11 redesign there is a single Self-Healing toggle row — the previously-nested Auto-fire toggle row was removed.
For agents
How it works
The trigger
Every cron run goes through the engine's notifyJobResult callback. When a run lands in failed after exhausting retryCount retries, the engine fires a system notification and a CRON_RUN_COMPLETED push event, then calls createHealAttemptIfEligible(job, freshRun) as fire-and-forget so the engine tick is never blocked by the heal pipeline. See notifyPermanentFailure in src/main/services/cron/cron-engine-notify.ts (the engine keeps a thin notifyJobResult delegate in cron-engine-service.ts).
The eligibility gate
Before doing anything, the orchestrator in src/main/services/cron-heal-orchestrator.ts walks six guards in order — the first failing guard short-circuits and no heal attempt is created:
job.healingEnabled === false— healing defaults ON, so this short-circuits only a job that has been explicitly opted OUT (healingEnabled: false).run.systemMarkedFailure === trueAND the run carries no failure evidence —system_marked_failurehas TWO writers with different meaning. The startup reconciler (recoverStaleRuns) flips runs stillrunningafter a crash — no real outcome, nofailure_evidence_json— and healing "the app crashed" is meaningless, so those are skipped. But the zombie-run watchdog (reapOverdueRuns) also sets the flag when it ends a run that overran its timeout; if that run's executor settled (even late) it carries REAL failure evidence, which is a genuine failure the user expects self-healing to engage. So only the evidence-less case is skipped. The discriminator is airtight:failure_evidence_jsonandstatus='failed'are written in one executor transaction, so a still-runningcrash-orphan can never carry evidence. (This is why a job that fails by overrunning its time limit now self-heals instead of being silently skipped — the 2026-06-22 postmortems-archive failure that motivated the fix.)job.runMode === 'one_off'— v1 excludes one-off jobs. A one-off has fired its single intended occurrence, so "fixed" doesn't have a clean meaning.job.consecutiveHealFailures >= 3— the cap. Three abandoned heals in a row means the agent isn't getting it; we stop looping and insert the escalation inbox card instead. The escalation card is acli_pending_actionsrow withactionKind: 'cron.failure_heal'and payload{escalation: true, jobId}— distinct from the legacy{healAttemptId}shape that the autoFireExistingPendingHeals migration drains. The dispatcher'scron.failure_healarm branches on'escalation' in payload: escalation rows are acknowledge-only (no spawn).- An open heal already exists for this job —
getOpenHealForJobreturns any heal inpendingorspawnedstate. One in flight at a time, job-scoped on purpose. - The fleet-wide open AUTO-heal count has hit the cap (
countOpenAutoHeals() >= MAX_CONCURRENT_AUTO_HEALS, default 3) — the GLOBAL bound (F119). Guards 4 and 5 are per job; a fleet-wide failure (a bad deploy, an expired shared credential, a network outage) fails many DISTINCT jobs at once, and each independently passes its own per-job gate — so without this nothing limits the TOTAL concurrent paid heal spawns. Over the cap the heal is deferred (it re-fires on the job's next failed run once an in-flight heal terminates and frees a slot). It counts only AUTO heals (is_manual = 0); the manual "Fix it" path (the escalation card) is user-initiated and exempt. The count is a snapshot, so a brief overshoot under a concurrent burst is accepted (a cost-pacing guard, not a correctness one — same class as NightyTidy2's cap-accuracy race).
Past the gate, evidence is collected (or re-read from the run's failure_evidence_json if the engine populated it) and classified into one of timeout, network, missing-file, permission, script-error, recipe-step-failed, or unknown. The classifier is a regex panel — see src/main/services/cron-failure-evidence.ts.
Always auto-fire (post 2026-05-11)
There is only one path now. After the eligibility gate accepts a failure the orchestrator inserts a cron_heal_attempts row in pending and immediately calls startCronHealSession(healAttemptId) — no cli_pending_actions row, no inbox card, no Approve button. If the spawn itself fails the row is marked abandoned with reason auto-fire-spawn-failed so the cap counter advances and you don't end up with a phantom spawned row. The category classifier still runs (the regex panel in cron-failure-evidence.ts) because the escalation card's preview text reads "3 failed self-heals — last category: <category>", but it no longer gates anything.
The cron_jobs.healing_auto_fire column is still in the schema for backward compatibility but Guard 1 reads healing_enabled only. Saving a job from the editor no longer writes healing_auto_fire, so the column drifts to whatever value it had at save time. Inert in practice — nothing reads it anymore — but worth knowing if you're inspecting the DB directly.
startCronHealSession (in cron-heal-orchestrator.ts) builds the prompt from job + last run + evidence + last 10 runs of history via src/main/services/cron-heal-prompt.ts, spawns a session via createSessionWithPrompt({ injectInAppToken: true }), then in a single SQLite transaction marks the heal spawned and calls toggleCronJob(job.id, false) — atomic, so a crash between those writes can't leave the cron firing alongside a live heal session.
Manual "Fix it" (user-triggered override)
The escalation card offers a Fix it button that spawns a heal on demand — a DELIBERATE user override, distinct from the automatic gate above. It routes useCronStore.healNow(jobId) → IPC.CRON_JOB_HEAL_NOW → startManualHeal(jobId) in cron-heal-orchestrator.ts. (The failure-alert card's own Fix with AI button no longer uses this manual-heal path — it opens a normal Start-session dialog pre-filled with the error and pauses the job; see cron-failure-alerts.md.)
startManualHeal reuses the SAME prompt + spawn as the auto path — both call the shared buildAndSpawnFixSession(job, run, evidence) helper, so the two entry points can't drift on prompt shape or spawn args. That shared helper also resolves the target project's real DB id before spawning: a project-less job (job.projectId null) heals into the Claude workspace, and the fallback maps the __claude__ folder_path sentinel to the Claude project's real row id via getProjectByFolderPath(CLAUDE_PROJECT_ID)?.id (mirroring email-inbound-service / workflow-engine / stuck-task-helper). Passing the raw sentinel as the projectId made createSessionWithPrompt's getProject('__claude__') return null and throw Project not found (__claude__), silently breaking BOTH the auto heal AND the manual Fix it for every project-less job (surfaced to the user as the toast "Could not start the fix session. Please try again."; fixed 2026-07-08). startManualHeal's eligibility differs from the auto gate on purpose:
- Bypasses auto Guard 1 (
healingEnabled === false) and Guard 4 (the 3-strike cap). The user is explicitly asking, so a manual fix runs even when auto-fix is off or has already given up — exactly the state the escalation card is in. (Without this, "Fix it" would silently no-op in the states the user most wants it — the same class as the old "Approve Heal does nothing" bug.) - Keeps the sane refusals: a
session-type orone_offjob has no script/recipe to repair (friendly reject);getOpenHealForJobstill blocks a second fix while one is in flight (auto or manual); and it needs a real failed run (getLatestFailedRun) to diagnose. - Records an
is_manualcron_heal_attemptsrow (the fix for the strand-forever bug). After the spawn,startManualHeal(viarecordManualHealAndPause) inserts a heal_attempt markedis_manualand — ONLY if the row was actually written — callsmarkHealSpawnedto recordsession_id+ pause the cron (aCRON_JOB_UPDATEDpush follows). That row is what makes the manual fix RECOVERABLE: the three un-pause paths (session archive, startup reconciler, hung-reaper) AND the in-app/run+/toggleauth scope all key on it — before this, a manual fix left the job paused forever and its own fix session got 403 trying to resume. Theis_manualflag keeps it OUTSIDE the auto-heal escalation accounting: the archive + hung-reaper abandons passcountsAsFailure:false, so a manual fix never pushes the job toward the 3-strike cap. On a rarerun_idcollision (the latest failed run already carries a heal_attempt) the insert no-ops and the job is left RUNNING rather than paused-without-a-record —markHealSpawned's cron-pause is unguarded, so pausing a phantom row would re-strand it. - Returns a discriminated
ManualHealResult({ ok, sessionId }|{ ok: false, reason }) so theCRON_JOB_HEAL_NOWhandler surfaces a friendly, already-humanized reason to the renderer.
Invariants locked by tests/unit/cron-heal-spawn.test.ts (startManualHeal suite): overrides Guards 1 & 4, refuses session / one-off / no-failed-run / already-running, records an is_manual spawned row + pauses, spawns-WITHOUT-pausing on a run_id collision, and — with the hung-reaper case in cron-heal-reconciler.test.ts — never counts a manual heal toward the cap. See cron-heal-idempotency-contract.md (I5).
Startup migration for pre-redesign pending heals
Users running the pre-2026-05-11 build with Auto-fire off had cron.failure_heal rows sitting in cli_pending_actions.status = 'pending' waiting on an Approve click. The Approve button never worked (the dispatcher arm has not accepted that legacy payload shape since the redesign), and even if it did the new design would still convert them to direct spawns. So autoFireExistingPendingHeals runs once at startup (cron-heal-orchestrator.ts) and:
- Selects every
cli_pending_actionsrow whereaction_kind = 'cron.failure_heal' AND status = 'pending'. - Skips any row whose payload is the new escalation shape (
{escalation: true, jobId}— has nohealAttemptIdto spawn against). - For each remaining legacy
{healAttemptId}row: marks the rowdispatched(markApprovedAndDispatched) BEFORE spawning, then callsstartCronHealSession. If the spawn throws the row staysdispatched(already approved) but the underlying heal_attempt ismarkHealAbandoned'd so the cap counter still advances. - After all rows are processed, emits a single coalesced
CLI_PENDING_CHANGEDpush so the renderer refreshes the inbox once.
This is one-shot drainage — once the queue is empty for legacy rows, the migration is a no-op on subsequent boots.
What the heal session can do
The prompt block tells the heal session it has two HTTP endpoints, scoped to its own cron job, authenticated with the in-app bearer token that was injected into the prompt:
POST /cron/jobs/<this job's id>/run— execute the cron one time on demand and return the exit code + output.POST /cron/jobs/<this job's id>/togglewith body{ "isActive": true }— resume the schedule.
The CLI server in src/main/services/cli/cli-server-cron-routes.ts verifies on every call that the request comes from a session that owns an open spawned heal for this job (classifyCronJobAuthScope). Any other in-app session calling these endpoints for this job gets a 403 session not authorized for this job. The user's external CLI control token continues to work normally.
The heal session is a regular Claude session in the cron's project, so it can also Read and Edit files in the project's working directory — that's how it actually fixes things.
What the heal session cannot do
- Rewrite the cron's definition. The heal session's bearer token only authorizes
/runand/togglefor this specific job. It cannot edit the cron expression, change the command, or alter any of the cron's other fields. Updating the cron definition still requires you in the Omniscio UI. - Touch other jobs. Other jobs'
/runand/toggleendpoints reject this session's bearer with 403. - Loop indefinitely. After three consecutive abandoned heals (
job.consecutiveHealFailures >= 3) Guard 4 stops creating new heal attempts and posts the escalation card.
Loop prevention — three layers
- Job-scoped open-heal guard (Guard 5, cron-heal-orchestrator.ts:160) —
getOpenHealForJobblocks a new heal while any existing heal for the same job is stillpendingorspawned. A run-scoped UNIQUE constraint oncron_heal_attempts(run_id)is the backstop. - Consecutive-failure cap of 3 (Guard 4, cron-heal-orchestrator.ts:151-154) —
cron_jobs.consecutive_heal_failuresincrements on every failingmarkHealAbandonedand resets to 0 onmarkHealFixed(both wrapped in a single transaction with the heal-attempt status flip — see queries-cron-heal-attempts.ts:108-147) — AND, since the recovery fix, on any successful scheduled RUN (clearHealEscalationOnSuccess; see § "Recovery on a successful run"). Once it hits 3, Guard 4 short-circuits and writes the escalation inbox card instead of creating a new heal attempt — until the counter is cleared by one of those paths. - Startup reconciler that abandons spawned heals across restarts (src/main/services/cron-heal-reconciler.ts:25-42) — the in-app HMAC secret rotates on every process boot, so any session token from a prior run can no longer authenticate to
/runor/toggle. On startup Omniscio abandons everyspawnedheal with reasonapp restart invalidated session tokenand toggles the parent cron back to active so the next scheduled run fires normally. Pending heals older than the inbox display TTL are also abandoned with reasonTTL expired.
Recovery on a successful run
The three layers above STOP a runaway heal loop; this is the complementary UN-stick. A job can hit the 3-strike cap and then recover on its OWN — the underlying problem was transient (a flaky network, an outage that passed), or its heals were only abandoned by app restarts (a lifecycle abandon, which does not itself mean the job is broken). Before this fix, nothing cleared the give-up counter or the escalation card in that case: the counter reset ONLY inside markHealFixed (a heal session verifying the fix and calling /toggle), so a job that quietly started passing again stayed capped forever, its escalation card falsely reading "keeps failing / paused" while every run went green (the 2026-08-07 stale-card incident).
clearHealEscalationOnSuccess (cron-heal-escalation.ts), called from notifyJobResult's success branch (cron-engine-notify.ts), closes the gap: a green run is treated as ground-truth proof the job is healthy, so it resets consecutive_heal_failures to 0 (re-arming self-healing) AND dismisses any open cron.failure_heal escalation/wedged card for the job (dismissOpenActionsByTarget) — in ONE transaction so a crash can't leave the counter cleared with the card stranded. A healthy job (counter already 0) is a no-op; an isolated secondary instance (AMC_INSTANCE_ID) skips entirely. It is the exact success-branch mirror of the failure alert's autoDismissOpenAlertOnSuccess, and needs no migration — every already-stuck job self-clears on its next green run. See cron-heal-idempotency-contract.md (I6).
Inbox card variants
The inbox dispatches actionKind === 'cron.failure_heal' rows to src/renderer/src/features/cli-pending/CronHealApprovalPane.tsx (CliPendingApprovalModal early-returns this pane on that actionKind). Two variants emit in the post-2026-05-11 design (the legacy pending heal-attempt variant only appears for rows that pre-date the redesign and is drained on next boot by the autoFireExistingPendingHeals migration):
escalation(the only variant created post-redesign) — three consecutive failed heals. Payload is{escalation: true, jobId}. Reads "<job>needs human attention" + "3 failed self-heals." No spawn happens when you approve — the dispatcher's escalation arm just stampsdispatched_atand the row exits the queue. The card's job is to make sure you notice.pending(legacy, migration-only) — pre-redesign rows with payload{healAttemptId}waiting on the broken Approve button. The startup migration drains these by spawning the heal session directly. They no longer appear in steady state, but the schema still accepts the shape so any row that survives the migration (e.g. crashed mid-drain) renders correctly until the next boot picks it up.
The approval pane (post-redesign)
When the user activates any cron.failure_heal row, Dashboard renders CronHealApprovalPane inline (via the ApprovalPaneShell primitive — same shell every other approval pane uses). After the 2026-05-11 redesign the pane has three view modes that branch on the parsed payload shape:
- Escalation view (the common case) — payload
{escalation: true, jobId}. Title reads This cron job needs your attention; the body shows the job summary (name, the humanized schedule viahumanizeCronExpression— e.g.0 */6 * * *renders as "Every 6 hours", and raw cron only as a fallback when the expression can't be parsed — plus the consecutive self-heal-failure count) and a plain-English "what you can do" note. Three buttons: Fix it (F, a manual-override heal viauseCronStore.healNow(jobId)— see § "Manual Fix it"; confirmed first since it spawns a paid session), View details (O,activateVirtualProject(CRON_PROJECT_ID)+ selects the job), and the primary Dismiss (Enter) which callsresolveInboxApproval('cli-pending', row.id, { kind: 'approve' }). There is no Pause (the escalation card's job is already paused) and no footer Snooze (the header clock snoozes;Hstill works). Dismiss is acknowledge-only — the dispatcher branches on'escalation' in payloadand stampsdispatched_atwithout spawning. - Legacy view (migration safety net) — payload
{healAttemptId}. Same five-section detail layout the original approval pane shipped with (job summary, failed run, evidence chips + stderr/stdout, recent history). The Approve button is a no-op label that explains the row is being drained by the startup migration; in steady state these rows never reach the user because the migration runs before the renderer mounts the inbox. - Invalid view — payload doesn't match either union arm. Renders a generic error explaining the row's payload couldn't be parsed and offering Reject as the only action. Should never be reachable because
cronFailureHealPayloadSchemais enforced at insert time, but defensive against schema drift.
Reject across all three views fires resolveInboxApproval('cli-pending', row.id, 'reject', reason) from an inline reason textarea, same convention as every other approval pane. closeDisabled={isResponding} matches the AgentDrivenSpawn pane convention so the user can't escape mid-IPC and leave the inbox in a stuck state. Reject reason is read from reasonRef.current at click time (not React state) to avoid the stale-closure bug where a rapid-typing user's last keystroke is dropped from the dispatched payload.
Where to look when it goes wrong
- electron-log. Heal pipeline events log under
[CronHeal](reconciler, session-lifecycle archive hook) and[cron-heal](orchestrator). Search for either prefix inmain.log(location: see logs-and-debugging.md). Notable lines:[CronHeal] Reconciler abandoned <N> spawned + <M> aged-out pending heal_attempts on startup— startup reconcile result.[CronHeal] Session <id> archived without resume — heal <id> abandoned, job <id> restored to active— heal session was archived before the agent called/toggle(session-lifecycle-service.ts:179-187).[cron-heal] auto-fire spawn failed for job <id>— auto-fire path could not start the heal session.[cron-heal] manual heal spawn failed for job <id> <error>— the manual Fix it path'sbuildAndSpawnFixSessionthrew; surfaces to the renderer as the toast "Could not start the fix session. Please try again." A historical cause (fixed 2026-07-08) was the__claude__folder_path sentinel reachingcreateSessionWithPromptunresolved →Project not found (__claude__); see § "Manual Fix it" for the projectId resolution.[cron-heal] manual heal for job <id>: run <id> already has a heal_attempt (run_id collision) — spawning the fix session WITHOUT pausing so the job is never stranded— the rare re-heal of an already-healed run; the fix session still runs but the cron is left active (never paused without a recoverable row).[cron-heal] manual heal finalize failed for job <id> <error>— the manual heal recorded its row butmarkHealSpawnedthrew; the row is abandoned non-counting so it can't soft-block future heals.[cron-heal] queue at cap, abandoning heal-attempt for job <id>— the CLI Pending queue is full, so the approval-gated heal could not be queued.
cron_heal_attempts.abandoned_reasoncolumn. Every abandoned heal records why. The values currently emitted are:auto-fire-spawn-failed— auto-fire path'sstartCronHealSessionthrew.app restart invalidated session token— startup reconciler.TTL expired— pending heal aged out beyond the inbox display window.session ended without resume— you archived the heal session before the agent called/toggle(an AUTO heal counts toward the cap here; a MANUAL "Fix it" does not —is_manual).manual-heal-finalize-failed— a manual Fix it recorded its heal_attempt but the follow-upmarkHealSpawnedthrew; abandoned non-counting so it can't block future heals.
- Inbox cards. An escalation card clears three ways: you dismiss it, a manual Fix it succeeds, or the job's next scheduled run succeeds on its own (
clearHealEscalationOnSuccessresets the counter + auto-dismisses the card — § "Recovery on a successful run"). So a card that PERSISTS means the job is still genuinely failing and needs you — a card that lingers while the job's runs are actually green is a bug (that was the 2026-08-07 stale-card incident, now fixed). Pending cards that disappear without a resume usually mean a startup reconcile or a TTL aging — theabandoned_reasoncolumn will tell you which.
Related
The other half of the same story — the toast and inbox card that tell you a job broke, which a running heal deliberately suppresses — is on the cron failure alerts page. How cron jobs are created and where their editor lives is on the create a cron job with AI page, the approval queue these cards share is on the CLI pending actions page, and finding the logs is on the logs and debugging page.
- cron-failure-alerts.md — the toast + inbox-card surface that fires when a cron fails and there is no open heal in flight. Suppressed by an open heal. Its Fix with AI button opens a Start-session dialog to fix the job by hand (a normal session pre-filled with the error, not this manual-heal path); its inline Auto-fix toggle flips self-healing on/off in place for the job.
- create-cron-job-with-ai.md — how cron jobs are created, where the editor lives, what the cron tick does in normal operation.
- cli-pending-actions.md — the inbox approval queue heal cards share with cron / automation / project-delete approvals.
- logs-and-debugging.md — finding
main.logand the in-app Debug Log Viewer.
Last verified 2026-09-23