---
title: Session handoff (carry a long chat into a fresh session)
---

# Session Handoff

## What it is

**Session handoff** turns one long, cluttered session into a **fresh session that
is already caught up** — same project, same folder, same branch, with a link back
to the one it came from. One click on **"Hand off to a fresh Omniscio session…"**
in a session's **⋯** menu does the whole thing.

Here is the problem it solves. A session you have been working in for hours is
carrying everything that happened in it: the three approaches you tried and threw
away, the files you opened that turned out to be irrelevant, the error messages
you already fixed, and the reasoning you built on an assumption that changed an
hour later. Omniscio's AI reads **all of it, every turn**, and it cannot tell
which parts you have mentally discarded. To it, a dead end you abandoned at 2pm
looks exactly as live as the thing you are actually doing now.

A handoff is an **edit**. It writes a summary that keeps the conclusions and drops
the search, then starts a new session holding that summary. The old session is not
deleted, archived, stopped, or changed in any way — it stays right where it is in
your sidebar, so if the handoff turns out to have missed something you can go
straight back and read it.

**Nothing happens automatically, and nothing is ever spent without a click.**
Omniscio can show a short note in the chat when a session gets long, suggesting a
handoff — that proactive note is **off by default**; you can turn it on in
**Settings → Sessions** (see below). The note is only a
suggestion: showing it costs nothing at all — no summary is written, no file is
saved, no new session is started. Only the button on the note, and then the **Hand
off** button in the box that opens, spends anything. The **⋯ menu** handoff is
always available regardless of this setting.

The same button also appears at the **one place a session used to be a dead end**.
When a session's context fills to the point that Omniscio's automatic recovery
tries a few times and still cannot get an answer out, it gives up rather than burn
money in a loop — and it used to leave a plain sentence there with no way forward.
Now that give-up note carries the same **Hand off to a new chat →** button too, so
a wedged session is never a dead end: one click summarizes it and continues in a
fresh one. Showing that button still costs nothing; only clicking it spends.

### Handoff vs. compaction — they are not the same thing

Claude's own answer to a full session is **compaction**: it writes a summary of
the conversation so far and swaps the old turns out, keeping one session alive.
That works, but it carries **everything** forward in condensed form — the dead
ends come along with the useful parts.

|              | Compaction (automatic)                           | Handoff (you choose it)                   |
| ------------ | ------------------------------------------------ | ----------------------------------------- |
| When         | Automatically, when the session is nearly full   | Only when you click                       |
| Result       | The **same** session, with its history condensed | A **new** session, linked to the old one  |
| Dead ends    | Carried forward, condensed                       | **Deliberately listed as "do not retry"** |
| The old chat | Replaced in place                                | Left completely untouched                 |
| Cost         | Included in the session                          | Free on a Claude plan, otherwise one small paid summary call |

Handing off does not replace compaction. A session you never hand off still
compacts exactly as it does today. See [Compaction Summary](compaction-summary.md).

## Where to find it

### How to use it

1. **Open a session's ⋯ menu** (top right of the session) and choose **"Hand off
   to a fresh Omniscio session…"**. It sits right next to the **Context NN%** row,
   because the number in that row is what makes the action worth taking.

   You will occasionally see two other versions of that item instead:
   - **"Already handed off: open the new Omniscio session"** — this session has
     been handed off once already. Omniscio only does it once per session, so the
     item takes you to the session it created instead of making a second one.
   - **Greyed out** — either the session has no conversation yet (there is nothing
     to summarize), or it is already at the end of a long chain of handoffs. Hover
     it and the tooltip says which.

2. **Read the box that opens.** It tells you, before anything is spent:
   - **How much context the session is using**, e.g. "Using 220k tokens of context
     (22% of a 1M window)". If the session has not run since you last started
     Omniscio, this says **unknown** — that is normal, and the handoff still works.
   - **What Omniscio will do**, as a numbered list.
   - **Both places it writes**, spelled out in full.
   - **What the summary call costs**, also printed on the button itself. If you
     are signed in with a Claude subscription, Omniscio runs the summary on that
     plan, the same way it runs your sessions, and the box says so instead of
     naming a price. Signed in with an API key instead, it estimates the charge
     in dollars.

3. **Optionally tell it what to emphasize.** There is a free-text box: _"Anything
   Omniscio should emphasize?"_ Whatever you type is folded into the summary
   request — for example _"keep the decision about the database, drop the CSS
   detour"_. Leave it blank and the standard summary is written.

4. **Optionally tick "Also append a short entry to this project's Auto Context."**
   Off by default. Ticked, Omniscio adds a ten-ish-line dated entry to a file
   every future session in this project reads. Useful for a running project
   history; it does cost a little context in every session from then on, which is
   why it is off unless you ask. Whichever way you leave it is remembered for the
   next handoff **in this project** — but only if you actually go through with the
   handoff. Ticking it and then cancelling changes nothing.

5. **Click Hand off.** That is the only point where anything is spent. Omniscio
   writes the summary, saves it, starts the new session, and takes you to it. The
   new session opens with a pinned **"Spawned by …"** note at the top pointing
   back at the parent, and the old session gets a **"Handed off to → …"** note
   pointing forward.

6. **If the box says the word-for-word tail did not fit**, read that line. It
   means the new session has the summary but not the last stretch of the
   conversation verbatim, because there was no room for both. If a recent detail
   matters, paste it in yourself once the new session opens. Omniscio always tells
   you when this happens — it never quietly drops it.

### An agent-requested handoff never takes over your screen

A handoff can also be **requested by an agent** rather than clicked by you. By
default it just happens — a session that fills up hands itself off and keeps
working, with no approval stop in between, because a handoff only continues work
you already set in motion. If you would rather review each one first, turn on
**Session handoff** under Settings → CLI Control → approval toggles
(`requireApprovalForCliSessionHandoff`); with that on, agent-requested handoffs
queue in your inbox as approvals again.

Either way, an agent-started handoff behaves differently from your own click in
exactly one way: **it never moves your view.**

- **You clicked Hand off yourself** → Omniscio takes you straight to the new
  session, because you asked for it and that is where you want to be.
- **An agent started the handoff** (instantly, or after your approval when the
  toggle is on) → the new session appears quietly in the sidebar and your screen
  stays exactly where it was.

This matters most when you **approve a batch at once**. Approving fifty queued
handoffs used to drag you into fifty new sessions one after another; now the whole
batch lands silently. You still get the confirmation for the batch, each request
visibly flips to approved, every new session shows up in the sidebar, and each
original session carries its **"Handed off to → …"** note pointing at its
successor — so nothing is hidden, it just does not interrupt you.

The same rule covers every session an approval can start, not only handoffs — a
re-run from an edited message, or a "let's get you unstuck" coaching session. If
one of them ever genuinely needs your attention, Omniscio shows a toast you can
click, never a jump you did not ask for.

### Turning the suggestion note on or off, or moving when it appears

The proactive note is **off by default** — turn it on (or tune when it appears) in
either of two places, same two settings:

- **Settings → Sessions → Advanced**, next to the 75% / 90% context warnings — which are
  now **off by default too**, so a fresh install gets no context nag of either kind.
- **Right on a note itself** (once you have switched it on) — the **⋯** button at the
  end of the note line (or right-click the note) opens a small popover that explains
  what the note is and carries both settings.

The settings are: whether to suggest at all, and **how the size is chosen** — by
default **Automatic**, which scales the trigger to the model (see below). Turn
Automatic off to pin your own exact token number instead. Changes affect **future**
notes only, so the one already on screen stays put. Pinning a very high custom number
is a way to **mute** the suggestion without switching it off — and the ⋯ menu item is
always there regardless of the note.

## How it behaves

### How the size is chosen (it scales with the model)

By default the trigger is **Automatic**: Omniscio sets it to **75% of the model's own
context window**, as a real token count. So a session on a 1M-window model (Opus 4.8,
Sonnet 5, Fable 5) is nudged at **750,000 tokens**, a 2M model at 1.5M, and an older
200k model at **150,000** — the very number the feature used to use for everyone. You
can turn Automatic off and pin an exact number if you prefer.

**Why it scales, and why it is still a real token count (not a "% full" gauge).** The
thing that degrades an AI's answers is how **long** its input is, not how full its
window is — but _how long is too long_ is larger for a larger, more capable model. A
single frozen 150,000 was the right number back when every model held 200,000 tokens
(0.75 × 200k), and it silently became far too eager when windows grew to 1,000,000: a
"start fresh?" nudge at 150k on a 1M model fires at **15% full**, which is absurdly
early. Scaling to 75% of each model's window fixes that while keeping the trigger an
absolute, auditable token count — a percentage is the right gauge for _"will this
compact soon?"_ but the wrong thing to *show* here, so you always see the real tokens.

**Per-model tuning where we have data.** A model whose answer quality is known to
degrade earlier gets a hand-tuned override instead of a flat 75%. **DeepSeek** is the
first: measured retrieval already slips well before its 1M window is full, so its
suggestion fires at **400,000** (its own quality ceiling) rather than 750k. Every other
model uses the 75% rule until there is a reason not to.

**What the number counts.** It is _context in use_, not the length of your conversation.
It includes the system prompt, the tool and skill definitions, and any project documents
folded into the session, alongside the actual turns. That is the correct quantity, because
the model reads all of it — and it means a heavier setup reaches the threshold sooner,
which is true and worth seeing rather than hiding. The wording is always "using … of
context", never "this conversation is …", for exactly that reason.

### The three-dead-ends rule of thumb

A size threshold is only half the signal. The other half you have to judge
yourself:

> **Once a session has produced three genuinely rejected approaches to the same
> problem, consolidating usually beats continuing.**

By then the context is holding three plausible-looking wrong paths that the model
cannot tell apart from the live one, and the cost of re-deriving from a clean
summary has dropped below the cost of reasoning around the wreckage. Three is also
Omniscio's existing number for "this isn't working" in two unrelated places, both
of which arrived at 3 independently.

**This is deliberately not automated, and it is not going to be.** A "dead end" is
a judgment about _intent_ — telling _abandoned because it was wrong_ apart from
_parked to come back to_ apart from _worked but looked messy_ means reading what
you meant. Every cheap stand-in available (turn count, turns that ended in an
error, stall-recovery firings, tool-call churn) misfires on ordinary hard work,
and a false alarm here would nag you toward **paid** new sessions you did not
need. A confident-looking detector that is wrong a third of the time is worse than
no detector, so the number lives here and in the note's tooltip as something you
apply, not something the app computes.

### Limitations worth knowing

- **The note needs a session that has run since Omniscio last started.** Context
  size is only ever held in memory — it is never saved to disk — so a session that
  has not streamed anything since the app launched has **no** token count for
  Omniscio to compare against the threshold. Practical consequence: **a session
  that was already past 150k when you restarted Omniscio will not get a
  suggestion.** The box shows "unknown" for its size, and the ⋯ menu item still
  works normally — this costs you a nudge, not the feature.
- **Ended and paused sessions** have no live number for the same reason. You can
  still hand them off; the summary is read from the database, not from a running
  process.
- **One handoff per session.** Deliberate. If you want to split work three ways,
  hand off, then hand off again from the successor.
- **Same project only.** The successor lands in the same project. Moving work
  between projects is a separate thing.
- **Handing off to a non-Claude engine — the sizing is settled, the rest is not.**
  The summary is database-based, so it works everywhere. The word-for-word tail is
  measured against the **successor's** window, including when you pick a different
  provider without naming a model — a case that used to measure against the window of
  the session you are *leaving*, and so could hand a smaller-window vendor far more
  text than it can hold. A provider Omniscio has no default model for now falls back
  to a small, safe budget rather than to your global model. What remains unverified is
  the rest of the non-Claude path end to end, not the sizing.

### When a handoff summary fails — what each message means

The handoff box names the actual reason, in your language, and says what to do. When the
paid fallback also runs and also fails, the message still names the subscription run's reason,
because that is the one you can act on.

| The message starts with | What happened | What to do |
| --- | --- | --- |
| "A handoff summary needs a working Claude sign-in or an Anthropic API key." | No Claude subscription login is available, and no API key is set up, so there was nothing to run the summary on. | Sign in or add a key in Settings → Accounts, then hand off again. |
| "Claude rejected your sign-in…" | Claude refused the login the summary ran on — for example, it was revoked or expired. | Sign in again in Settings → Accounts, then retry. |
| "Your Claude plan's usage limit is reached…" | The subscription hit its usage cap. | Retry after the limit resets. |
| "This conversation is too long for Claude to summarize in one pass…" | Claude refused the request as too large. | Start a fresh session by hand and paste in what matters; retrying the same session will not help. |
| "Claude is overloaded or unreachable right now…" | Claude's service was busy or could not be reached. | Retry in a few minutes. |
| "Writing the handoff summary took too long…" | The summary ran past Omniscio's time limit and was stopped. | Retry, or hand off a shorter conversation. |
| "Couldn't write the handoff summary." | Something else went wrong. | Retry once; if it keeps happening, report it. Omniscio already sends this one to its crash reports. |

The "too long", "took too long" and "something else" cases are reported to Omniscio's crash
reporting automatically, each under its own name. A sign-in, usage-limit or busy-service failure
is logged but not reported as a crash, because it is not a fault in the app.

## For agents

### How it works

**The suggestion note.** At the end of each turn, Omniscio checks the session's
context size against the threshold and — if the session has crossed it — writes
one system note into the chat with `metadata.kind = 'handoff-notice'`. The
decision goes through a pure helper, `decideHandoffNotice()` in
[ndjson-context-decisions.ts](/src/main/process/ndjson-context-decisions.ts), which
applies four suppression rules on top of the raw threshold, mirroring the existing
75% / 90% warnings:

- **Once per fill-up.** Tracked per session in memory, cleared only by a real
  compaction. An app restart, a respawn, or a rate-limit account switch does not
  re-nag you.
- **Never while the agent is deliberately parked waiting** — a waiting agent is not
  burning context, so the note would be a false alarm.
- **Never on a session that has already been handed off.**
- **Never on a session at the end of a long handoff chain.**

The note itself renders as its own row in the chat (never folded into a
neighbouring message) so the **Hand off** button and the ⋯ popover have something
to attach to.

**The give-up dead-end button.** When the automatic context-stall recovery has
tried its maximum number of times and still cannot get an answer, the give-up
step routes through the **same** emitter as the proactive note, so the dead-end
message becomes a real `handoff-notice` row with a working **Hand off to a new
chat →** button instead of an inert sentence. Its once-per-fill-up tracking is kept
**separate** from the proactive note's on purpose: a session only gets stuck this
way by filling its context, so it has nearly always seen the earlier note already, and
sharing the tracking would hide the button in exactly the situation it exists for.
So the dead-end row always appears, even if you saw the earlier note — while a
repeat give-up before the next compaction still will not stack a second button on
top. It does share the same visibility gate, so a user for whom the feature is
switched off never gets a paid button written into their transcript — they get a
plain explanatory note instead. Like everything else here, showing it spends
nothing; only the click does.

**What the handoff does, in order.** All of the checks come first, so a handoff
that cannot possibly work never spends anything:

1. Writes the summary. This is the **one AI call** — a side call that reads the
   conversation from Omniscio's database. It prefers your Claude subscription (a
   one-off headless run on the same login your sessions use, so it costs nothing
   beyond the plan) and falls back to a metered API-key call only when an Anthropic
   API key is set up and the subscription run is unavailable or fails. Without a key
   the paid call is never tried — it could only fail, and would hide the real reason
   the subscription run gave. It is **not** a message sent
   into the session, so the old session's context is not touched at all. The
   summary is asked for in **seven sections**: what we set out to do · decisions
   made and why · what is done and verified · what is in flight · open questions ·
   files and branches touched · and **what NOT to redo**.
2. Saves the **full** summary as a document in Omniscio's own data folder — on
   purpose **outside** your project. Anything inside a project's Auto Context
   folder is read by every future session in that project, and the full document
   is meant for exactly one successor.
3. If you ticked the box, appends the short ten-line digest to the project's
   `HANDOFFS.md`. That file is **capped at 20 entries** — the 21st drops the
   oldest. Without the cap it would grow forever and become a permanent tax on
   every session in the project.
4. Records the link from the old session to the new one **before** starting the new
   session. That ordering is what makes a double-click, or a crash halfway
   through, unable to produce two paid sessions.
5. Starts the successor in the same project, folder and branch. It uses the **same
   model by default**, but you can **pick a different provider or model** in the box
   before handing off — the reason to do so is usually that the model you are on is
   the thing that broke (it is stuck, or its login has failed), so continuing on the
   same one would just wedge again. The model is checked against the provider you are
   moving **to**, before anything is spent: a model that provider cannot run is refused
   up front with no charge, instead of failing after the summary was paid for (fixed
   2026-09-26 — a DeepInfra → DeepSeek pick used to fail that way). It carries the
   summary and — when there is room — the last stretch of the old conversation **word
   for word**.
6. Pins a "Handed off to → …" note on the old session.

**The seventh section is the one that matters most.** "What NOT to redo" is the
field a hand-written summary reliably forgets, and it is the feature's one
plain-eye quality test: **if the new session re-tries something the old one
already ruled out, the summary failed.** Nothing else about summary quality is
reliably judgeable by looking at it.

**The word-for-word tail.** A summary is lossy by construction, and the last few
turns are the ones most likely to matter and least likely to be captured well by
one. So the new session gets the summary **plus** a verbatim tail, sized against
the new session's own window after the summary and the standard project context are
accounted for. When there is no room left, the tail is dropped **and Omniscio says
so on screen** — it never trades detail for space quietly.

**The complete transcript is always preserved.** Separately from the bounded tail,
every handoff also writes the **entire** parent conversation, word for word, to a
file in the project's own `.claude/handoffs/` folder, and tells the successor where
it lives. So even when the tail does not fit, nothing is actually lost — the new
session can open the full transcript and read any earlier stretch it needs, rather
than that detail being gone for good. The file is historical reference the successor
reads on demand, **not** standing project context, so it is never folded into every
future session the way an Auto Context entry would be.

**Length is not the limit.** The summary is one pass, and the model that writes it is
pinned to a large-context one rather than inherited from whatever your sessions happen to
be set to — so the longest sessions, which are exactly the ones worth handing off, are
summarized in a single call instead of being the ones that fail. The transcript is still
trimmed to fit that model's window (oldest middle first, keeping the start and the end),
so a very long chat is condensed rather than refused for its size — and if Claude refuses
it anyway, the handoff says so in plain words (see the failure messages below).

**If that one call fails,** Omniscio says so plainly and stops there. Nothing is written,
no successor is started, and the old session is left exactly as it was — including the
"already handed off" marker, which is only set *after* the summary succeeds. So a failed
handoff costs you the attempt and nothing else, and its message names the cause — see **When
a handoff summary fails** above.

## Related

- [Context usage warnings (75% and 90%)](context-warnings.md) — the older
  percentage-based warnings. They are **also off by default now**, so no context nag of
  either kind reaches a fresh install; the handoff note is still a separate row with its
  own separate setting, sized per model in absolute tokens for the reason in **How the
  size is chosen** above.
- [Compaction Summary](compaction-summary.md) — what Claude does automatically when
  a session fills up, and how to read the summary it writes. Handoff is the
  deliberate alternative to letting that happen.
- [Session provenance](session-provenance.md) — the "Spawned by …" note pinned at
  the top of the successor is the same mechanism that traces any spawned session
  back to its origin.
- [Keep my computer smooth under load](smooth-load.md) — governs how often
  Omniscio takes an authoritative context reading; the handoff note forces a fresh
  reading once a session is near the threshold, so it never fires off a stale
  number.
