---
title: Real Conversation Layout (hide the tool noise, keep your first message pinned)
---

# Real-conversation layout

## What it is

### What it does

When you open a Claude Code session in Omniscio, the chat panel — by default — shows just **the real conversation**: the messages you actually typed, the agent's actual final replies, and inline "Conversation compacted" landmarks where they happened. Every agent turn's intermediate work — tool calls, file reads, "let me check X" connective prose, the whole stream of internal steps — folds into a collapsed **activity toggle in the agent bubble's header** (next to the bot icon and timestamp). Click anywhere on that header to expand the per-turn tool detail into a panel that drops INSIDE the same bubble, between the header and the final answer. Leave it closed if you only care about the dialogue. Once a turn finishes, the final answer below it is always visible — you never have to expand anything to see the agent's reply. (While a turn is still _streaming_ the header is already expanded and the body shows the live work — see "While a turn is in flight" below.)

Auto-continues — the "Please continue" prompts Omniscio sends on your behalf when an answer is interrupted, when a rate-limit retry resumes, or when a scheduled send fires — are not part of the conversation you had, so they are hidden from the default view too.

The result: a 4,000-message tool-heavy thread reads as the 50-or-so messages of the conversation that actually happened, top to bottom. Your first message is pinned at the very top with **nothing above it** — no "Show more messages" pill, no spacer, no gap. The first user message _is_ the top of the thread.

## Where to find it

### The settings toggle and the kill switch

- **Settings → Performance → "Real-conversation layout"** — user-facing toggle. Default `true`. Off reverts to the prior render. Stored as `AppSettings.realConversationLayoutEnabled` and respected on next cold mount.
- **`VITE_DISABLE_MESSAGE_BODY_CULL=1`** — build-time kill switch for the off-screen-cull layer only (see Performance above). Set it at build time to drop `content-visibility: auto` from the message bodies while keeping the layout itself on — useful for A/B-ing the cull's smoothness or ruling it out in a regression hunt. It does NOT disable the layout; only the **Settings → Performance → "Real-conversation layout"** toggle does that.
- **`AMC_DISABLE_INCREMENTAL_TURNS=1`** — environment-variable kill switch for the incremental turn-rebuild fast path only (see Performance above). Set it to force every render to rebuild the whole turn list from scratch (the pre-2026-05-29 behavior) while keeping the layout itself on. Because the incremental path is byte-for-byte identical to the full rebuild, this switch can only change _speed_, never what renders — it exists to rule the optimization out in a regression hunt. Default ON; any value other than `"1"` keeps it on.

## How it behaves

### What the user sees

From top to bottom in a session that has the layout on (the default):

- **Your first message**, anchored at the top. Nothing renders above it.
- Then, oldest to newest, **every real exchange**: a message you typed, then for each agent turn **one merged bubble** whose header is the **"▸ 🤖 activity label · took Xs · 10:27 PM"** toggle and whose body is the **agent's real reply** — one bordered card per turn, header on top, answer beneath. Clicking the header reveals the per-turn tool detail in a panel between header and body. A turn whose agent never ran a tool renders a plain bubble (no merged header — just the standard `Bot · Agent · timestamp` header and the answer). A turn whose agent's final act was a tool call shows the merged bubble with the tool detail in the expansion panel and an empty body.
- **`> Compaction N` expanders** appear inline where each compaction happened — they are landmarks, not truncation points. Visually each one is a **full-width centered section divider**: thin horizontal line on the left, the `▸ Compaction N` toggle in the middle, thin horizontal line on the right — so the boundary reads as a clean break across the chat, not as a left-aligned button. Real messages render both above and below them. Each expander is closed by default; click it once to open it AND fetch the model-generated summary. The opened body is exactly that summary, rendered as markdown, centered under the divider — nothing else. The first click lazy-fires a `COMPACTION_SUMMARY_READ` IPC to read the saved summary file; subsequent collapse + re-open cycles reuse the cached result without re-fetching. (Rev 9, 2026-05-22 — supersedes the rev-3 two-level shape that also wrapped pre-compact tool work + a nested `📜 Compaction summary` sub-chevron inside the expander. Rev-9 divider styling, 2026-05-23 — supersedes the original left-aligned button shape.)
- While a compaction is **in progress** (after the agent has decided to compact but before the divider has landed), a pulsing amber **"Compacting conversation…"** status row renders below the most recent agent answer so you can see it's happening. It is replaced inline by the settled `> Compaction N` expander once the divider arrives.
- This continues all the way down to **your most recent real message** and the **agent's latest reply**.

What you do NOT see by default:

- **Auto-continues.** "Please continue", rate-limit retries, scheduled-send prompts. Real DB rows (they're recorded as `is_auto_response = 1` operator messages), just hidden from the default render.
- **Intermediate tool calls.** Every file read, command run, "let me check X" preface — folded into that turn's single activity summary. The header's activity label shows the count even while collapsed (`5 comments · 5 tool calls`), so you can gauge the turn's weight at a glance; the full step-by-step is one click away when you expand the turn.
- **Bootstrap noise.** Plain unkinded `Session ready` system rows are dropped at source. Kinded system markers — rate-limit notices, snooze markers, seed-context, sub-agent results — stay visible as inline rows between turns. If a kinded marker falls _inside_ a compaction window, it lives inside that window's expander and only appears when you open the expander.

  **What still surfaces (notable-system carve-out).** Diagnostic system rows that announce a state change you need to see — placeholder-stuck pauses, CLI process loss, `Session stopped by you`, CLI errors, `Plan auto-approved`, the expired-token "please re-authenticate" prompt — render as plain bubbles. The producer tags them with `metadata.kind = 'notable-system'` when it writes the row, so they survive the unkinded drop the same way rate-limit and snooze markers do. This is what tells you _why_ a session in your inbox is paused: without these rows you'd see your last message, an empty agent bubble, and no context for the pause. (Hot-fixed 2026-05-23 after the prior default dropped these rows along with the bootstrap noise — see the [postmortem](../../.claude/memory/postmortems/notable-system-messages-dropped-postmortem.md). A small follow-up sweep remains: ~25 other diagnostic emissions in the same area aren't kinded yet and will still paint empty bubbles if they're the latest row — they'll be converted as that code is touched.)

- **Recovery & lifecycle events fold into the agent's prose, not their own bubble.** When a session quietly recovers or flips state — a rate limit clears and the session auto-resumes, the app restarts after a crash or a normal quit and reconnects, a stalled turn reconnects and re-sends your last message, you pause or resume a session, the app is closed mid-work and the session is saved to resume on next launch, or an expired login token is refreshed mid-turn and your last turn is silently re-run, or a brief API error triggers a silent automatic retry — Omniscio folds a one-line italic note into that turn's collapsed activity instead of painting a standalone centered row. You see it only when you expand the turn: e.g. _— Session paused, will resume on restart —_, _— Rate limit cleared, session auto-resumed —_, _— App restarted, session reconnected —_, _— Session paused —_ / _— Session resumed —_, _— Auth token refreshed (attempt 1/4), last turn re-run —_, _— API error (502), retrying automatically… —_. The event still happened and is still recorded in the database; it just rides inside the agent's thinking area so the dialogue stays clean. This is the deliberate opposite of the notable-system bubbles above: those announce a state you must act on (so they stay visible even collapsed), while these quietly annotate something that already self-healed (so they hide until you look). The "paused — will resume on restart" note is the most common one — it's what you'll see folded into the last turn of any session that was mid-work when you last quit Omniscio.

Want to see the hidden plumbing for one session — auto-continues, recovery prompts, away-mode auto-replies, and the unkinded system events — without leaving the layout? Flip the per-session **Show system messages** toggle (⋯ menu → More → Filter Messages). It reveals exactly these rows for that one session and nothing else. See [show-system-messages-toggle.md](show-system-messages-toggle.md).

Compaction landmarks are never hidden by those status folds. If a turn contains one or more `> Compaction` dividers, whole-turn auto-wait / background-check-in / woken-no-op folding is vetoed so the dividers stay visible in the timeline. This prevents a stored sequence like compactions 1–9 from rendering only the ninth divider because the earlier dividers lived inside older auto-wait-collapsed turns.

### Expanding the activity panel

Clicking the merged bubble's **header** (the row with the chevron + bot icon + descriptive activity label + "took Xs" + timestamp) expands the per-turn tool detail into a panel between header and body. The descriptive activity label is **always visible — collapsed settled headers included** (R67, 2026-06-21): a glance at `🤖 5 comments · 5 tool calls` tells you whether a turn did enough work to be worth opening, without expanding it first. Only the **"took Xs" duration** stays reserved for the expanded view (or while streaming) — a collapsed settled header shows the count but no duration. The label is built from cheap counts recorded once when the row was finalized — no content parsing — so it survives compaction unchanged.

**The activity label adapts to your screen width.** On a desktop-width window it shows both halves of the turn's activity joined by a dot — `N comments · N tool calls` — and drops whichever half is zero (a tool-only turn reads just `5 tool calls`; a reply-only turn reads just `2 comments`). On a narrow / mobile window there is only room for one metric, so the header **prefers comments**: it reads `N comments` whenever the turn produced any agent prose, and falls back to `N tool calls` only when the turn had no comments at all. Turns that included extended thinking prefix `Thinking · ` to whichever metric applies (or read a bare `Thinking` when there is no count to qualify it). Older turns recorded before Omniscio began persisting the per-row comment / tool counts fall back to the previous `N tool calls` / `Thinking` wording.

**The same mobile-prefers-comments rule governs the between-compaction summary blocks.** Each `> Compaction N` divider can carry a sibling activity block summarizing the agent work in that inter-compaction window (see "How the chat panel decides what to render"). On desktop that block reads `C comments, A actions in Z` (e.g. `7 comments, 19 actions in 11m`); on a narrow / mobile window it drops the action count and reads `C comments in Z` — unless the window was pure tool work with no comments (`C` is 0), in which case it falls through to `A actions in Z` so the block is never blank or reads "0 comments". Older inter-compaction blocks recorded before Omniscio persisted the comment / tool counts keep the legacy `N actions took Z` wording on both widths. (This is why an old session reopened today still shows `N actions took Z` for its compaction blocks — the comment count is derived once at finalize and never back-filled onto rows that predate it.)

The full interleaved body — the chevrons + preface prose, one block per hidden row — loads on demand the first time you expand a settled (older) turn: Omniscio fires one per-row content fetch over IPC, joins the responses, and renders them into the expansion panel. The fetch is per-row and per-turn, so opening one turn's panel does not load any other turn's bodies. Closing the panel releases the cache; re-opening re-fetches. If you never click, you never pay for it.

Some turns are **pure tool work** with no final text reply — the agent's last act was a tool call, not a message to you. For those turns the body is empty and the expansion is the _only_ content, so Omniscio opens it automatically and never lets it render blank. While the fetch is in flight you briefly see **"Loading actions…"**; if the content genuinely can't be retrieved (the underlying rows were pruned, or the session was deleted) you get an honest **"Activity details failed to load."** line with a **Retry** link, not an invisible empty bubble. This floor is why a tool-only reply can never "disappear" into a header-only shell.

For the **most recent** in-flight turn the header is auto-expanded as it streams, and the body renders the live work directly from what's already in memory — no fetch needed (see "While a turn is in flight" below). The instant the turn settles, the header auto-collapses to its compact form (chevron + activity label + timestamp — the count stays, only the live body and "took Xs" drop away) and the body switches from live-work to the agent's final answer. The next click on the header brings the tool detail back, lazy-fetched this time; the activity label was already on the collapsed header.

Each row inside the expanded view is the natural shape of one Anthropic assistant message: preface prose (italic, visible) plus a nested `> N tools used` sub-chevron that batches that row's tool invocations behind one more click. A row with zero preface shows only the sub-chevron; a row with zero tools shows only the preface. One row may stream text and tool calls together, and the sub-chevron sits at the row boundary.

Once the activity panel is open, you can collapse it in two ways: click the header chevron again, **or** click the thin clickable vertical rail that runs down the left edge of the expanded panel. The rail gives you a second collapse target so you don't have to scroll back up to the header to close a long expanded panel. It's only present on settled turns — while a turn is still streaming the body IS the live work, and the rail is hidden so a stray click can't hide what you're watching.

### While a turn is in flight (streaming)

While a turn is streaming, Omniscio shows **one** agent bubble whose header is **auto-expanded** and reads `streaming… <elapsed>`. The word _streaming…_ is painted in the green agent-accent color with a left-to-right shimmer that sweeps across the letters on a 2-second loop (a horizontally-translating gradient highlight, not a pulsing dot — `prefers-reduced-motion: reduce` flattens it to solid green text). The bubble's body shows the live work: the tool calls and the connective "let me check X" prose as the agent writes them. The expansion panel between header and body stays empty during streaming — the body already shows the live work, so a separate panel would duplicate the same prose. **You can collapse a streaming bubble, too:** click the header chevron while the agent is still working and the live body folds away to a header-only row — the `streaming…` shimmer and the elapsed timer keep running in the header so you can still see the turn is alive, while the in-progress work tucks out of the way. Click the chevron again to bring the live body back; the collapse holds as new content keeps streaming in behind the collapsed header. The instant the turn settles, the header auto-collapses to its minimal form (unless you collapsed it yourself, which makes the collapsed state sticky for the rest of that turn), the body switches from live-work to the agent's final answer, and the per-turn tool detail moves to the expansion panel where it lives for any future click.

**If you turn on the "Show only the final message" marker** (the opt-in final-message-marker feature), there is one refinement to the above (2026-08-10): the instant the agent emits its terminal `[[OMNISCIO_FINAL]]` marker with a real answer after it, the *still-streaming* body switches to showing **only** that final message — the pre-marker narration folds into a reachable collapsed block right away, instead of waiting for the turn to fully settle. This is what makes the marker feel like it "works" while you watch: under heavy multi-session load a finished-but-not-yet-settled turn would otherwise keep showing its working narration for a while. It applies on both desktop and your phone, and the folded work is always one click away (nothing is hidden unreachably). Without the marker enabled (the default) the streaming body shows all live work until settle, exactly as described above.

**A live elapsed-time timer sits next to "streaming…"** so you can see how long the agent has been working on the current turn. It starts at `0s` the moment the turn begins and ticks once a second: under a minute it counts seconds (`5s`, `42s`); from one to ten minutes it shows minutes and seconds (`1m 45s`); past ten minutes it drops the seconds (`12m`); past an hour it reads `1h 2m`. The timer stays visible even if you expand the bubble to watch the work — expanding does not stop the clock. The moment the turn settles, the live timer is replaced by a frozen **`took <elapsed>`** badge on the now-collapsed pill (same compact format, e.g. `took 1m 30s`), so a glance at any past turn tells you how long that answer took to produce. The badge is computed from the turn's recorded start and end timestamps, so it is stable across reloads and compaction. This timer is part of the real-conversation layout only — the legacy stream (toggle off) shows no timer.

The user's words: "I still want to see the agent's displayed responses as it's doing its thinking work … that's how it's been for months." and "for the thinking action I want it to show a timer so the user can see how long."

### When the layout applies — and when it doesn't

The layout is the **default cold-mount render** on desktop and mobile when `realConversationLayoutEnabled` is on (which it is, out of the box).

It does **not** affect:

- **Search.** Ctrl+K still indexes the full content of every row, auto-continues and tool output included. If a search hit lands on a tool-output line inside a not-yet-expanded turn, Omniscio auto-expands that turn's pill before scrolling.
- **Export to markdown, share-publish, audit-export.** Those keep calling the full-content path and see every row.
- **The legacy stream.** Turn the toggle off (Settings → Sessions → "Real-conversation layout") and the chat panel reverts to the prior render — every row, every auto-continue, the per-row activity overview. (The unified "Load older messages" pill that legacy layout used to render at the top is disabled by default for every user as of 2026-05-21 via a code-level flag in `SessionPanel.tsx` — see [load-older-pill-disabled postmortem](../../.claude/memory/postmortems/load-older-pill-disabled-postmortem.md).) Useful for debugging or if you really want the firehose for a specific session.

### Why this exists

Two goals, both required:

1. **Readability.** A long agent session, rendered as every row, is a wall of noise — file reads, status pings, retried operator turns, partial thinking, tool output blocks. The dialogue you actually had with the agent is buried inside it. The new layout strips the chat panel down to the conversation and folds everything else behind one collapsed control per turn.
2. **Performance.** A 4,000-message tool-heavy thread must open instantly on both desktop and mobile. The real conversation inside such a thread is typically 50–150 messages, so rendering it whole is cheap. The cost was always the tool chatter, and tool chatter is now deferred per-turn behind a click. One catch: the layout renders those 50–150 bubbles _unwindowed_ (no virtualizer), and while that mounts cheaply, the browser re-painted every off-screen bubble on each scroll frame, which made scrolling a long finished session laggy. The fix (2026-05-28) is browser-native off-screen culling — `content-visibility: auto` on each message body, which skips layout + paint of off-screen bubbles without unmounting their DOM. See the [scroll contract § Off-screen render culling](../../.claude/memory/contracts/frontend-scroll-contract.md) for the full mechanism and invariants.

   A second cost lived in _streaming_. The visible panel rebuilds its turn list whenever new content arrives, so a session that was actively streaming rebuilt **all** of its turns on every token-batch — and with 10–20+ sessions open at once, that rebuild was the main reason switching between sessions and scrolling felt sluggish on desktop (mobile only ever has one such panel mounted, so it never paid the cost). The fix (2026-05-29) reuses the unchanged earlier turns and rebuilds only the one turn currently being written, producing a byte-for-byte identical result far faster (measured ~33× faster per streaming update in the isolated harness). It is internal-only — nothing about what you see changes — and can be turned off with `AMC_DISABLE_INCREMENTAL_TURNS=1` if it ever misbehaves.

The layout _is_ the presentation half of the lazy-content-load architecture ([lazy-content-load.md](lazy-content-load.md)). The lazy plumbing — `summary_prose`, `has_tool_activity`, the function-pair `getSessionHistoryLite()`, the per-turn `SESSION_GET_TURN_CONTENT` IPC — was the data layer; this is the layer that consumes it visibly. Without the layout, the lazy plumbing was only half the feature.

### What didn't change

- **Search index.** `conversation_messages_fts` is unaffected by this layout change — it indexes per the [search rule](../../src/main/db/fts-index-text.ts) (operator messages full content, agent messages their final trailing prose only; mid-turn narration, tool calls/results, streaming partials and system rows are never indexed), and Ctrl+K still finds everything that rule indexes.
- **Export, share-publish, audit-export, FTS.** Still call `getSessionHistory()` directly and see every row.
- **The legacy render path.** Turning the setting off restores the prior chat panel exactly — every row, every auto-continue, the per-row activity overview, the top-of-thread "Load older messages" pill.

## Related

- [lazy-content-load.md](lazy-content-load.md) — the data-layer foundation this layout consumes.
- [show-system-messages-toggle.md](show-system-messages-toggle.md) — per-session toggle that reveals the plumbing rows this layout folds away.
- [scroll-position-memory.md](scroll-position-memory.md) — the other layer that affects how a session paints on return.
- [agent-message-display.md](agent-message-display.md) — the per-row structured-prose render; still used inside each real-reply bubble.
- Build contract in the repo: `.claude/memory/project_real_conversation_layout.md` — the R-numbered build rules + AC acceptance criteria + verbatim approving-session user quotes. Rev-7 additions (R14 two-bubble shape, R15 compaction-survival contract, R16 in-progress compaction indicator), Rev-9 additions (R19 thinking timer), Rev-11 additions (R23 inter-compaction activity bubble, R23.a `preserveProseOnExpand`, R23 in-flight fetch gate), Rev-13 additions (R24 merged-bubble-per-turn, R25 header-is-the-toggle, R26 expansion-inside-the-bubble, R27 streaming preserves R8, R28 rev-11 inter-compaction untouched — supersedes R14), and Rev-32 additions (R56 viewport-responsive comment/action label on BOTH the per-turn merged bubble and the inter-compaction activity bubble — mobile prefers the comment count, desktop keeps both) are appended at the bottom of that file.
