---
title: Vendor server fault recovery
---

# Vendor server fault recovery

## What this is for

Every AI provider has outages. When one happens, the same session can behave two completely
different ways depending on something you never chose: whether it had already produced some output
before the outage hit.

A session caught **mid-work** rode the outage out quietly — Omniscio retried on a growing backoff,
the session came back on its own, and nothing was asked of you. A session caught on its **very first
turn** was parked permanently and told to open the model picker and choose a different model. That
advice was wrong: the provider's servers were down for a few minutes, and no model change would have
helped. On 2026-09-29 a DeepSeek outage stranded five brand-new sessions that way.

Both cases now take the same road.

## What changed

A provider reports *why* it refused a request, and Omniscio already kept all of it — the HTTP status,
the text, and the error name the provider itself sent (`server_error`, `overloaded_error`,
`service_unavailable`, `api_error`, and a few more). The **name of a fault on the provider's own
side** is now read as a temporary failure, wherever it arrived: in the reply text, or as the error
name on its own with no text at all.

That single reading drives both decisions — whether to retry and what to tell you — so they can never
disagree.

## What you see

- **A temporary provider fault:** the session says the provider returned a temporary error and that
  Omniscio keeps retrying in the background, and adds that no action is needed. It then retries on a
  growing backoff, and if the outage outlasts those retries it keeps retrying in the background on a
  much longer, restart-safe schedule. Nothing is asked of you, and it is not a red card you have to
  clear.
- **A real, permanent refusal** — a model the provider has withdrawn, a model configuration it
  rejects, an account out of credit, a bad or revoked key, or a conversation that no longer fits:
  unchanged. The session stops, names the real cause, and tells you the one thing that fixes it
  (choose another model, add credit, sign in again).
- **Nothing reported at all:** unchanged. If the provider produced no output and said nothing about
  why, the session still stops and points at the model — the honest reading of a hard, silent
  rejection is still the model, and it is deliberately not treated as retryable, because a genuinely
  dead model id would then retry for ever.

## Why the temporary case is worth knowing about

A provider blip is the one failure where doing nothing is correct. If you see a session resting with
"it'll keep retrying automatically", you can leave it alone — it is not waiting on you, and switching
its model would not have helped.

## Where this lives

The reading is one predicate, `isTransientVendorFault`, in
[placeholder-vendor-evidence.ts](../../src/main/process/placeholder-vendor-evidence.ts). It is read by
the placeholder filter that arms the retry and by the sink that writes the ending; the rules it
serves are in provider-spawn-model-availability-contract and transient-spawn-retry-contract.
