Session handoff (carry a long chat into a fresh session)
Session handoff turns one long, cluttered session into a fresh one that is already caught up: same project, folder and branch, carrying a summary that keeps the conclusions and lists the dead ends as things not to retry. One click in a session's overflow menu does it, and the old session is left completely untouched.
What it is
Session handoff turns one long, cluttered session into a fresh session that is already caught up — same project, same folder, same branch, with a link back to the one it came from. One click on "Hand off to a fresh Omniscio session…" in a session's ⋯ menu does the whole thing.
Here is the problem it solves. A session you have been working in for hours is carrying everything that happened in it: the three approaches you tried and threw away, the files you opened that turned out to be irrelevant, the error messages you already fixed, and the reasoning you built on an assumption that changed an hour later. Omniscio's AI reads all of it, every turn, and it cannot tell which parts you have mentally discarded. To it, a dead end you abandoned at 2pm looks exactly as live as the thing you are actually doing now.
A handoff is an edit. It writes a summary that keeps the conclusions and drops the search, then starts a new session holding that summary. The old session is not deleted, archived, stopped, or changed in any way — it stays right where it is in your sidebar, so if the handoff turns out to have missed something you can go straight back and read it.
Nothing happens automatically, and nothing is ever spent without a click. Omniscio can show a short note in the chat when a session gets long, suggesting a handoff — that proactive note is off by default; you can turn it on in Settings → Sessions (see below). The note is only a suggestion: showing it costs nothing at all — no summary is written, no file is saved, no new session is started. Only the button on the note, and then the Hand off button in the box that opens, spends anything. The ⋯ menu handoff is always available regardless of this setting.
The same button also appears at the one place a session used to be a dead end. When a session's context fills to the point that Omniscio's automatic recovery tries a few times and still cannot get an answer out, it gives up rather than burn money in a loop — and it used to leave a plain sentence there with no way forward. Now that give-up note carries the same Hand off to a new chat → button too, so a wedged session is never a dead end: one click summarizes it and continues in a fresh one. Showing that button still costs nothing; only clicking it spends.
Handoff vs. compaction — they are not the same thing
Claude's own answer to a full session is compaction: it writes a summary of the conversation so far and swaps the old turns out, keeping one session alive. That works, but it carries everything forward in condensed form — the dead ends come along with the useful parts.
| Compaction (automatic) | Handoff (you choose it) | |
|---|---|---|
| When | Automatically, when the session is nearly full | Only when you click |
| Result | The same session, with its history condensed | A new session, linked to the old one |
| Dead ends | Carried forward, condensed | Deliberately listed as "do not retry" |
| The old chat | Replaced in place | Left completely untouched |
| Cost | Included in the session | Free on a Claude plan, otherwise one small paid summary call |
Handing off does not replace compaction. A session you never hand off still compacts exactly as it does today. See Compaction Summary.
Where to find it
How to use it
Open a session's ⋯ menu (top right of the session) and choose "Hand off to a fresh Omniscio session…". It sits right next to the Context NN% row, because the number in that row is what makes the action worth taking.
You will occasionally see two other versions of that item instead:
- "Already handed off: open the new Omniscio session" — this session has been handed off once already. Omniscio only does it once per session, so the item takes you to the session it created instead of making a second one.
- Greyed out — either the session has no conversation yet (there is nothing to summarize), or it is already at the end of a long chain of handoffs. Hover it and the tooltip says which.
Read the box that opens. It tells you, before anything is spent:
- How much context the session is using, e.g. "Using 220k tokens of context (22% of a 1M window)". If the session has not run since you last started Omniscio, this says unknown — that is normal, and the handoff still works.
- What Omniscio will do, as a numbered list.
- Both places it writes, spelled out in full.
- What the summary call costs, also printed on the button itself. If you are signed in with a Claude subscription, Omniscio runs the summary on that plan, the same way it runs your sessions, and the box says so instead of naming a price. Signed in with an API key instead, it estimates the charge in dollars.
Optionally tell it what to emphasize. There is a free-text box: "Anything Omniscio should emphasize?" Whatever you type is folded into the summary request — for example "keep the decision about the database, drop the CSS detour". Leave it blank and the standard summary is written.
Optionally tick "Also append a short entry to this project's Auto Context." Off by default. Ticked, Omniscio adds a ten-ish-line dated entry to a file every future session in this project reads. Useful for a running project history; it does cost a little context in every session from then on, which is why it is off unless you ask. Whichever way you leave it is remembered for the next handoff in this project — but only if you actually go through with the handoff. Ticking it and then cancelling changes nothing.
Click Hand off. That is the only point where anything is spent. Omniscio writes the summary, saves it, starts the new session, and takes you to it. The new session opens with a pinned "Spawned by …" note at the top pointing back at the parent, and the old session gets a "Handed off to → …" note pointing forward.
If the box says the word-for-word tail did not fit, read that line. It means the new session has the summary but not the last stretch of the conversation verbatim, because there was no room for both. If a recent detail matters, paste it in yourself once the new session opens. Omniscio always tells you when this happens — it never quietly drops it.
An agent-requested handoff never takes over your screen
A handoff can also be requested by an agent rather than clicked by you. By
default it just happens — a session that fills up hands itself off and keeps
working, with no approval stop in between, because a handoff only continues work
you already set in motion. If you would rather review each one first, turn on
Session handoff under Settings → CLI Control → approval toggles
(requireApprovalForCliSessionHandoff); with that on, agent-requested handoffs
queue in your inbox as approvals again.
Either way, an agent-started handoff behaves differently from your own click in exactly one way: it never moves your view.
- You clicked Hand off yourself → Omniscio takes you straight to the new session, because you asked for it and that is where you want to be.
- An agent started the handoff (instantly, or after your approval when the toggle is on) → the new session appears quietly in the sidebar and your screen stays exactly where it was.
This matters most when you approve a batch at once. Approving fifty queued handoffs used to drag you into fifty new sessions one after another; now the whole batch lands silently. You still get the confirmation for the batch, each request visibly flips to approved, every new session shows up in the sidebar, and each original session carries its "Handed off to → …" note pointing at its successor — so nothing is hidden, it just does not interrupt you.
The same rule covers every session an approval can start, not only handoffs — a re-run from an edited message, or a "let's get you unstuck" coaching session. If one of them ever genuinely needs your attention, Omniscio shows a toast you can click, never a jump you did not ask for.
Turning the suggestion note on or off, or moving when it appears
The proactive note is off by default — turn it on (or tune when it appears) in either of two places, same two settings:
- Settings → Sessions → Advanced, next to the 75% / 90% context warnings — which are now off by default too, so a fresh install gets no context nag of either kind.
- Right on a note itself (once you have switched it on) — the ⋯ button at the end of the note line (or right-click the note) opens a small popover that explains what the note is and carries both settings.
The settings are: whether to suggest at all, and how the size is chosen — by default Automatic, which scales the trigger to the model (see below). Turn Automatic off to pin your own exact token number instead. Changes affect future notes only, so the one already on screen stays put. Pinning a very high custom number is a way to mute the suggestion without switching it off — and the ⋯ menu item is always there regardless of the note.
How it behaves
How the size is chosen (it scales with the model)
By default the trigger is Automatic: Omniscio sets it to 75% of the model's own context window, as a real token count. So a session on a 1M-window model (Opus 4.8, Sonnet 5, Fable 5) is nudged at 750,000 tokens, a 2M model at 1.5M, and an older 200k model at 150,000 — the very number the feature used to use for everyone. You can turn Automatic off and pin an exact number if you prefer.
Why it scales, and why it is still a real token count (not a "% full" gauge). The thing that degrades an AI's answers is how long its input is, not how full its window is — but how long is too long is larger for a larger, more capable model. A single frozen 150,000 was the right number back when every model held 200,000 tokens (0.75 × 200k), and it silently became far too eager when windows grew to 1,000,000: a "start fresh?" nudge at 150k on a 1M model fires at 15% full, which is absurdly early. Scaling to 75% of each model's window fixes that while keeping the trigger an absolute, auditable token count — a percentage is the right gauge for "will this compact soon?" but the wrong thing to show here, so you always see the real tokens.
Per-model tuning where we have data. A model whose answer quality is known to degrade earlier gets a hand-tuned override instead of a flat 75%. DeepSeek is the first: measured retrieval already slips well before its 1M window is full, so its suggestion is pinned to 750,000 — kept equal to its compaction point — rather than the flat 75% default, which would resolve to 786,432 and fire after compaction had already happened. Every other model uses the 75% rule until there is a reason not to.
What the number counts. It is context in use, not the length of your conversation. It includes the system prompt, the tool and skill definitions, and any project documents folded into the session, alongside the actual turns. That is the correct quantity, because the model reads all of it — and it means a heavier setup reaches the threshold sooner, which is true and worth seeing rather than hiding. The wording is always "using … of context", never "this conversation is …", for exactly that reason.
The three-dead-ends rule of thumb
A size threshold is only half the signal. The other half you have to judge yourself:
Once a session has produced three genuinely rejected approaches to the same problem, consolidating usually beats continuing.
By then the context is holding three plausible-looking wrong paths that the model cannot tell apart from the live one, and the cost of re-deriving from a clean summary has dropped below the cost of reasoning around the wreckage. Three is also Omniscio's existing number for "this isn't working" in two unrelated places, both of which arrived at 3 independently.
This is deliberately not automated, and it is not going to be. A "dead end" is a judgment about intent — telling abandoned because it was wrong apart from parked to come back to apart from worked but looked messy means reading what you meant. Every cheap stand-in available (turn count, turns that ended in an error, stall-recovery firings, tool-call churn) misfires on ordinary hard work, and a false alarm here would nag you toward paid new sessions you did not need. A confident-looking detector that is wrong a third of the time is worse than no detector, so the number lives here and in the note's tooltip as something you apply, not something the app computes.
Limitations worth knowing
- The note needs a session that has run since Omniscio last started. Context size is only ever held in memory — it is never saved to disk — so a session that has not streamed anything since the app launched has no token count for Omniscio to compare against the threshold. Practical consequence: a session that was already past 150k when you restarted Omniscio will not get a suggestion. The box shows "unknown" for its size, and the ⋯ menu item still works normally — this costs you a nudge, not the feature.
- Ended and paused sessions have no live number for the same reason. You can still hand them off; the summary is read from the database, not from a running process.
- One handoff per session. Deliberate. If you want to split work three ways, hand off, then hand off again from the successor.
- Same project only. The successor lands in the same project. Moving work between projects is a separate thing.
- Handing off to a non-Claude engine — the sizing is settled, the rest is not. The summary is database-based, so it works everywhere. The word-for-word tail is measured against the successor's window, including when you pick a different provider without naming a model — a case that used to measure against the window of the session you are leaving, and so could hand a smaller-window vendor far more text than it can hold. A provider Omniscio has no default model for now falls back to a small, safe budget rather than to your global model. What remains unverified is the rest of the non-Claude path end to end, not the sizing.
When a handoff summary fails — what each message means
The handoff box names the actual reason, in your language, and says what to do. When the paid fallback also runs and also fails, the message still names the subscription run's reason, because that is the one you can act on.
| The message starts with | What happened | What to do |
|---|---|---|
| "A handoff summary needs a working Claude sign-in or an Anthropic API key." | No Claude subscription login is available, and no API key is set up, so there was nothing to run the summary on. | Sign in or add a key in Settings → Accounts, then hand off again. |
| "Claude rejected your sign-in…" | Claude refused the login the summary ran on — for example, it was revoked or expired. | Sign in again in Settings → Accounts, then retry. |
| "Your Claude plan's usage limit is reached…" | The subscription hit its usage cap. | Retry after the limit resets. |
| "This conversation is too long for Claude to summarize in one pass…" | Claude refused the request as too large. | Start a fresh session by hand and paste in what matters; retrying the same session will not help. |
| "Claude is overloaded or unreachable right now…" | Claude's service was busy or could not be reached. | Retry in a few minutes. |
| "Writing the handoff summary took too long…" | The summary ran past Omniscio's time limit and was stopped. | Retry, or hand off a shorter conversation. |
| "Couldn't write the handoff summary." | Something else went wrong. | Retry once; if it keeps happening, report it. Omniscio already sends this one to its crash reports. |
The "too long", "took too long" and "something else" cases are reported to Omniscio's crash reporting automatically, each under its own name. A sign-in, usage-limit or busy-service failure is logged but not reported as a crash, because it is not a fault in the app.
For agents
How it works
The suggestion note. At the end of each turn, Omniscio checks the session's
context size against the threshold and — if the session has crossed it — writes
one system note into the chat with metadata.kind = 'handoff-notice'. The
decision goes through a pure helper, decideHandoffNotice() in
ndjson-context-decisions.ts, which
applies four suppression rules on top of the raw threshold, mirroring the existing
75% / 90% warnings:
- Once per fill-up. Tracked per session in memory, cleared only by a real compaction. An app restart, a respawn, or a rate-limit account switch does not re-nag you.
- Never while the agent is deliberately parked waiting — a waiting agent is not burning context, so the note would be a false alarm.
- Never on a session that has already been handed off.
- Never on a session at the end of a long handoff chain.
The note itself renders as its own row in the chat (never folded into a neighbouring message) so the Hand off button and the ⋯ popover have something to attach to.
The give-up dead-end button. When the automatic context-stall recovery has
tried its maximum number of times and still cannot get an answer, the give-up
step routes through the same emitter as the proactive note, so the dead-end
message becomes a real handoff-notice row with a working Hand off to a new
chat → button instead of an inert sentence. Its once-per-fill-up tracking is kept
separate from the proactive note's on purpose: a session only gets stuck this
way by filling its context, so it has nearly always seen the earlier note already, and
sharing the tracking would hide the button in exactly the situation it exists for.
So the dead-end row always appears, even if you saw the earlier note — while a
repeat give-up before the next compaction still will not stack a second button on
top. It does share the same visibility gate, so a user for whom the feature is
switched off never gets a paid button written into their transcript — they get a
plain explanatory note instead. Like everything else here, showing it spends
nothing; only the click does.
What the handoff does, in order. All of the checks come first, so a handoff that cannot possibly work never spends anything:
- Writes the summary. This is the one AI call — a side call that reads the conversation from Omniscio's database. It prefers your Claude subscription (a one-off headless run on the same login your sessions use, so it costs nothing beyond the plan) and falls back to a metered API-key call only when an Anthropic API key is set up and the subscription run is unavailable or fails. Without a key the paid call is never tried — it could only fail, and would hide the real reason the subscription run gave. It is not a message sent into the session, so the old session's context is not touched at all. The summary is asked for in seven sections: what we set out to do · decisions made and why · what is done and verified · what is in flight · open questions · files and branches touched · and what NOT to redo.
- Saves the full summary as a document in Omniscio's own data folder — on purpose outside your project. Anything inside a project's Auto Context folder is read by every future session in that project, and the full document is meant for exactly one successor.
- If you ticked the box, appends the short ten-line digest to the project's
HANDOFFS.md. That file is capped at 20 entries — the 21st drops the oldest. Without the cap it would grow forever and become a permanent tax on every session in the project. - Records the link from the old session to the new one before starting the new session. That ordering is what makes a double-click, or a crash halfway through, unable to produce two paid sessions.
- Starts the successor in the same project, folder and branch. It uses the same model by default, but you can pick a different provider or model in the box before handing off — the reason to do so is usually that the model you are on is the thing that broke (it is stuck, or its login has failed), so continuing on the same one would just wedge again. The model is checked against the provider you are moving to, before anything is spent: a model that provider cannot run is refused up front with no charge, instead of failing after the summary was paid for (fixed 2026-09-26 — a DeepInfra → DeepSeek pick used to fail that way). It carries the summary and — when there is room — the last stretch of the old conversation word for word.
- Pins a "Handed off to → …" note on the old session.
The seventh section is the one that matters most. "What NOT to redo" is the field a hand-written summary reliably forgets, and it is the feature's one plain-eye quality test: if the new session re-tries something the old one already ruled out, the summary failed. Nothing else about summary quality is reliably judgeable by looking at it.
The word-for-word tail. A summary is lossy by construction, and the last few turns are the ones most likely to matter and least likely to be captured well by one. So the new session gets the summary plus a verbatim tail, sized against the new session's own window after the summary and the standard project context are accounted for. When there is no room left, the tail is dropped and Omniscio says so on screen — it never trades detail for space quietly.
The complete transcript is always preserved. Separately from the bounded tail,
every handoff also writes the entire parent conversation, word for word, to a
file in the project's own .claude/handoffs/ folder, and tells the successor where
it lives. So even when the tail does not fit, nothing is actually lost — the new
session can open the full transcript and read any earlier stretch it needs, rather
than that detail being gone for good. The file is historical reference the successor
reads on demand, not standing project context, so it is never folded into every
future session the way an Auto Context entry would be.
For a built-in project such as the default Claude project, "the project's own
folder" is the real folder that project works in (for Claude, ~/Claude), the same
folder the new session starts in. A session whose project has no folder on this
computer at all (an OpenClaw session, which runs on a remote gateway) can't be handed
off; Omniscio says so before anything is spent.
Length is not the limit. The summary is one pass, and the model that writes it is pinned to a large-context one rather than inherited from whatever your sessions happen to be set to — so the longest sessions, which are exactly the ones worth handing off, are summarized in a single call instead of being the ones that fail. The transcript is still trimmed to fit that model's window (oldest middle first, keeping the start and the end), so a very long chat is condensed rather than refused for its size — and if Claude refuses it anyway, the handoff says so in plain words (see the failure messages below).
If that one call fails, Omniscio says so plainly and stops there. Nothing is written, no successor is started, and the old session is left exactly as it was — including the "already handed off" marker, which is only set after the summary succeeds. So a failed handoff costs you the attempt and nothing else, and its message names the cause — see When a handoff summary fails above.
Related
- Context usage warnings (75% and 90%) — the older percentage-based warnings. They are also off by default now, so no context nag of either kind reaches a fresh install; the handoff note is still a separate row with its own separate setting, sized per model in absolute tokens for the reason in How the size is chosen above.
- Compaction Summary — what Claude does automatically when a session fills up, and how to read the summary it writes. Handoff is the deliberate alternative to letting that happen.
- Session provenance — the "Spawned by …" note pinned at the top of the successor is the same mechanism that traces any spawned session back to its origin.
- Keep my computer smooth under load — governs how often Omniscio takes an authoritative context reading; the handoff note forces a fresh reading once a session is near the threshold, so it never fires off a stale number.
Last verified 2026-10-05