Omniscio documentation
Browse all documentation
  1. Getting Started13
  2. Sessions & Agents115
  3. Inbox & Notifications59
  4. Projects & Tasks95
  5. Automation & Scheduling75
  6. Knowledge & Memory26
  7. AI Features60
  8. Integrations100
  9. Plugins & Marketplace33
  10. Cloud & Teams56
  11. Settings & Customization58
  12. Account & Billing28
  13. Troubleshooting84
  14. CLI & API Reference22
  15. Legal & Policies4
  16. Uncategorised22

Doc Token Alerts (heads-up when agent-instructions grow large)

The amber inbox card Omniscio raises when a project's always-loaded agent documents grow past a token threshold — what it counts, the three ways to act on it, and how to change or silence the threshold.

What it is

The files an AI coding agent (Claude Code, Cursor, Copilot, etc.) reads on every turn — the project's root CLAUDE.md and AGENTS.md, plus the .claude/memory/MEMORY.md index — bloat silently over a project's life. Each adds context-window tokens to every single message the agent sends. Past about 40,000 tokens the always-loaded bundle starts crowding out the user's actual question, slowing responses and degrading reasoning quality. There is no in-editor signal when you cross that line — the agent just gets quietly worse.

Omniscio watches those three files per project and, when their combined estimated token count crosses a threshold, drops an amber card into the inbox, titled "Agent docs over threshold — <project name>", so the user knows it's time to prune. Opening the card offers three ways to act — Archive (acknowledge at the current size), Start session (let an agent propose a slim-down), or Edit (open the biggest of the three files yourself). The alert is opt-out per project (a per-project rule can disable it) and the re-alert is tiered: acknowledging a 50k-token alert silences the card until the docs grow past a configurable growth % (default 50 %, so ~75k) or a configurable cadence of days elapses (default 30) — whichever comes first. A big jump surfaces immediately; a slow creep still gets a periodic nudge.

Three concepts shape the feature:

  1. The always-loaded bundle is the alert basis. Omniscio measures three files per project — the root CLAUDE.md, the root AGENTS.md, and .claude/memory/MEMORY.md. These are the files an agent reads on every turn. The token count for the alert is the sum of the counted files (see concept 4 — a synced mirror is excluded so the total isn't double-counted).
  2. The .claude/skills/**and.claude/memory/** totals are computed for context too, but they are NOT the alert basis in v1 — and they are no longer surfaced in the UI. The scan still records them (the project_doc_token_scans row carries skills_tokens / memory_tokens), but neither the inbox row nor the opened card displays them: against this alert's purpose — the always-loaded bundle — they were noise. Skill / memory files are read on-demand, not every turn, so they don't crowd the live context the same way.
  3. A per-doc breakdown lives in the card (not the inbox row, which is a bare title). It lists each always-loaded file and its token estimate, so you can see which file is heavy — the CLAUDE.md, the AGENTS.md, or the MEMORY.md. The headline total turns red when you're at or over the threshold.
  4. Sync-aware counting. If you use Omniscio's agent-instructions-sync feature, your AGENTS.md is kept as a byte-for-byte mirror of your CLAUDE.md (or vice versa), and your Claude Code agent auto-loads only one of the pair per turn. Counting both would double-count a whole file. So when sync is ON, the synced mirror is excluded from the always-loaded total (and skipped by the Edit button). The breakdown still lists the mirror — dimmed, labeled "synced mirror, not counted" — so the total is honest. When sync is OFF, all three files count. MEMORY.md is never a sync mirror, so it always counts.
  5. The threshold, the re-alert knobs, and the dismiss-ack are per-project, layered on a global default. A row called 'global' lives in doc_token_alert_rules and seeds the default threshold (40,000 tokens), growth % (50), and cadence (30 days). Adding a row keyed to a project id overrides the defaults for that project only. The 'global' row can be edited but never deleted — the queries module throws if you try.

The token estimate is byte-cheap by design: Math.ceil(file.size / 4) from a single lstat per file. Omniscio never reads the contents of these files to count tokens. Walking a project takes O(files visited) syscalls, no I/O against the file bodies, and the walk is asynchronous — it yields to the event loop between directories, so a scan never holds the main thread. That combination is what lets the scan run live-ish on every session turn.

Where to find it

How to use it

  1. Wait for the card. When a project's always-loaded bundle crosses its threshold (default 40,000 tokens), an amber card appears in your inbox alongside your other alerts, titled Agent docs over threshold — <project name>.

  2. Open the card. It shows the headline always-loaded tokens total (rendered red when you're at/over threshold), the threshold (default 40,000 or your per-project override), how far over you are, and a per-doc breakdown — each existing always-loaded file (root CLAUDE.md, root AGENTS.md, .claude/memory/MEMORY.md) with its own token estimate, so you can see which file to prune first. A file excluded as a synced mirror (see concept 4) is shown dimmed and labeled "synced mirror, not counted".

  3. Archive — the card's Archive button (top of the card; also the keyboard E shortcut, middle-click, bulk archive, or an inbox rule) acknowledges the docs at their current size. It writes the current always-loaded total into acknowledged_tokens and the instant into acknowledged_at on the per-project scan row. The alert then re-fires on the tiered rule: when the bundle grows at least the growth % beyond what you acknowledged (default 50 %), or when the cadence in days elapses (default 30) while still over threshold — whichever trips first. It's still sticky-to-size (you can't "shut it up forever"), but it also won't let a slowly-creeping bloat hide indefinitely under the growth gate. Setting the cadence to 0 disables the periodic nudge, leaving growth-only.

  4. Start session — the card's universal Start-session button, pre-filled with a stock prompt that names the three always-loaded files, reports the current tokens / threshold / overage, and asks the agent to assess the docs, propose a slim-down plan, wait for your approval, then apply it. Use this when you'd rather have the agent do the pruning. The card stays in the inbox until a later scan confirms the docs dropped below threshold.

  5. Edit — jump straight into the largest counted always-loaded file in the built-in markdown editor (a synced mirror is skipped — editing it would just be overwritten by sync, so the button opens the canonical instead). Saving the trimmed file automatically re-runs the scan, so the card self-clears once you've dropped below the threshold (no manual re-scan needed). Use this for a quick hand-prune. If none of the always-loaded files exist (or the only one is an excluded mirror), Edit is unavailable and the card says so — Start a session to create them, or Archive.

  6. Snooze instead, if you want it back tomorrow. Snooze on the card (right-click, or the card's own Snooze button) puts it away for the chosen duration without touching acknowledged_tokens, so when the snooze expires you'll see the card again at whatever the docs are then. You can also mute Context Size alerts entirely — from the card, or from Settings → Notifications → Alert types.

  7. Tune the per-project ceiling and re-alert behavior. Open Settings → Notifications → Doc Alerts. The top cards hold the global defaults — threshold (default 40,000), re-alert growth % (default 50), and re-alert cadence (days) (default 30). Below that you can add a per-project override — pick a project and set its own threshold, growth %, and cadence (or disable the alert for that project entirely). Editing a global default applies to every project that doesn't have an override; editing an override applies to just that one. Changing a threshold also clears that project's dismiss anchor so a new ceiling doesn't get silently suppressed by an old ack — but tuning only the growth % or cadence preserves the existing dismiss, so adjusting a knob never itself re-nags you.

  8. Disable the alert globally. Toggle the master switch in Settings → Notifications → Doc Alerts off. This is the enabled flag on the 'global' rule. With it off, no alerts fire for any project (unless a per-project override has enabled: true).

How it behaves

How it works

Storage (base tables migration v247 + tiered-realert ledger migration)

Two SQLite tables back the feature. The base tables landed in migration v247; the tiered re-alert knobs + dismiss timestamp were added additively later by the ledger migration 20260530204917-doc-token-tiered-realert.ts:

  • doc_token_alert_rules holds the threshold + enabled flag + the two re-alert knobs (realert_growth_pct, realert_cadence_days), keyed by project id. The reserved id 'global' is seeded with threshold_tokens = 40000 and enabled = 1; the ledger migration added the knobs defaulting to 50 / 30 (including for the existing 'global' row). Per-project rows override the global; lookups via resolveDocTokenAlertRule(projectId) in /src/main/db/queries-doc-token-alert.ts fall back to the global row when the project has no override OR when the project is soft-deleted (projects.is_deleted = 1). Deleting the global row throws "cannot delete global doc-token-alert rule" — there's no way to break the system into a fall-back-less state.
  • project_doc_token_scans caches the most recent scan per project — always_loaded_tokens, skills_tokens, memory_tokens, full_tree_tokens, file_count, acknowledged_tokens + acknowledged_at (the dismiss anchor — size and instant), and scanned_at. One row per project, upserted on each scan. There is no per-file breakdown column — the card's per-doc breakdown is computed fresh on mount (see "The card body and its three actions" below), not persisted. The ledger migration backfilled acknowledged_at for already-dismissed scans so their cadence clock starts at upgrade time (nobody is nagged the instant they update).

upsertDocTokenAlertRule runs in a transaction and clears the dismiss anchor (acknowledged_tokens + acknowledged_at) only when the effective threshold changes — a lowered or raised threshold is a real event the user should see re-surface, but tuning only the growth % or cadence preserves the existing dismiss (editing a knob must never itself re-nag). deleteDocTokenAlertRule always clears the anchor, since dropping an override flips the threshold back to the global value.

Scan walker

/src/main/services/doc-scan/doc-token-scan.ts walks a project's working directory and returns the three-bucket scan result. Constraints:

  • MAX_DEPTH = 8 — recursion stops eight directory levels deep so a misconfigured project with a deep node_modules symlink can't hang the scan.
  • MAX_FILES = 5000 — the walker bails after 5000 files visited. Both caps are intentionally pessimistic; a typical Omniscio project visits well under 100 files for the three buckets.
  • Skip rules — dotfiles (except .claude/), symlinks (avoids node_modules-style infinite recursion), and the OS-junk set (HIDDEN_FILES = new Set(['.ds_store', 'thumbs.db', 'desktop.ini'])).
  • Token estimate — Math.ceil(file.size / 4), off the lstat the walk already performed. The walker never reads file contents. This is the same bytes/4 heuristic the Claude Files sidebar uses, which keeps the two surfaces consistent.
  • Bucket attribution — uses POSIX forward-slash relpaths so the rule works identically on Windows and macOS. The three "always-loaded" files are CLAUDE.md and AGENTS.md at the project root, plus .claude/memory/MEMORY.md. Anything else under .claude/skills/ lands in the skills bucket; anything else under .claude/memory/ lands in the memory bucket. Files outside those prefixes don't count toward any bucket.
  • Virtual-project sentinels — the walker returns all-zeros for sentinel project paths (__inbox_pilot__, __tags__, etc.) since those projects have no real filesystem. The one carve-out is CLAUDE_PROJECT_ID (__claude__), which resolves to ~/Claude via the shared resolveProjectWorkDir() and gets scanned normally.

Scan triggers

/src/main/services/doc-scan/doc-token-scan-service.ts exposes two entry points:

  • requestDocTokenScan(projectId) — debounced. Stores a timeout in a per-projectId Map; the actual scan + persist runs after DEBOUNCE_MS = 60_000 of quiet. Calls during the quiet window reset the timer. This is what drives live-ish scanning without re-scanning on every single turn of an active session.
  • scanDocTokensNow(projectId) — async, awaited by the DOC_TOKEN_SCAN_REQUEST IPC handler, and always a fresh walk (it never takes the skip-if-unchanged shortcut). The one renderer caller is the built-in markdown editor: when the user saves one of the three always-loaded docs from the card's Edit action, /src/renderer/src/stores/file-peek-store.ts fires this IPC fire-and-forget so a slimmed-down file clears the inbox row immediately, without waiting for the next session-turn debounce.

Both paths persist the result via upsertDocTokenScan (which auto-clears acknowledged_tokens if the new always-loaded total drops below the threshold — a clean slate when the user actually prunes) and then emit IPC.DOC_TOKEN_ALERT_UPDATED so the renderer can refresh.

Scans are kicked off from two places in /src/main/services/doc-scan/doc-token-scan-triggers.ts:

  1. Startup fan-out — at app boot, listProjects() filtered through isSpawnableProjectPath() (real filesystem paths plus the CLAUDE_PROJECT_ID carve-out) gets one requestDocTokenScan each, so every project warms its scan cache in the background.
  2. Session-turn ended — a sessionEvents.on('turnEnded', ...) listener resolves the session's project id and calls requestDocTokenScan(projectId). This is the live trigger: every time a Claude turn ends, Omniscio schedules a debounced re-scan of that project's docs.

Inbox row + decision

The pure decision lives in /src/shared/alert-features/doc-token-alert.ts:

export const ALERT_BUCKET: DocTokenBucket = 'always-loaded'
export const DEFAULT_REALERT_GROWTH_PCT = 50
export const DEFAULT_REALERT_CADENCE_DAYS = 30
export const DOC_TOKEN_ALERT_DEFAULT_THRESHOLD = 40_000
const MS_PER_DAY = 86_400_000

export function shouldFireDocTokenAlert(scan, rule, now: string): boolean {
  if (!rule.enabled) return false
  const value = pickAlertValue(scan)
  if (value < rule.thresholdTokens) return false
  if (scan.acknowledgedTokens == null) return true

  // Big-jump trigger (integer cross-multiplication — no float drift).
  if (value * 100 >= scan.acknowledgedTokens * (100 + rule.realertGrowthPct)) return true

  // Cadence trigger (0 disables; missing acknowledgedAt fails quiet).
  if (rule.realertCadenceDays > 0 && scan.acknowledgedAt != null) {
    if (Date.parse(now) - Date.parse(scan.acknowledgedAt) >= rule.realertCadenceDays * MS_PER_DAY) {
      return true
    }
  }
  return false
}

The gates run in order: the rule must be enabled, the value must be over threshold, and — once an ack is on file — re-alert is tiered so EITHER trigger fires it: the value grew at least realertGrowthPct past the dismissed size, OR realertCadenceDays have elapsed since the dismiss. The clock is injected (now) rather than read inside the function, so the cadence branch is testable without faking time; the one production caller (handleDocTokenAlertList) computes new Date().toISOString() once and passes the same snapshot to every row. The big-jump integer cross-multiplication (value * 100 >= ack * (100 + pct)) sidesteps floating-point drift — pinned in /tests/unit/doc-token-alert-decision.test.ts so a "let's just use * 1.5" refactor will fail CI.

The decision walk moved to the main process: /src/main/services/doc-scan/doc-token-alert-items.ts computes which projects are firing right now (shouldFireDocTokenAlert), and /src/main/services/doc-scan/doc-token-alert-card.ts turns that into a central inbox alert card — one per project, dedup key doc-token-alert:<projectId> — via the shared reconcileAlertFamily/registerAlertArchiveListener building blocks every alert family uses. Snoozing the card works exactly like any other inbox item; there is no per-feature snooze branch.

The card body and its three actions

The row itself carries no inline buttons — it is a plain inbox row. Clicking it opens the central alert card in AlertInboxViewer, which mounts /src/renderer/src/features/doc-token-alert/DocTokenAlertCardBody.tsx for the doc-token-alert:<projectId> dedup key. The card body reads its live item straight from useDocTokenAlertStore keyed on projectId, and falls back to the alert's own stored text (no numbers, no action) once that project's row vanishes from the store — push-resolved by another client, dropped below threshold by a fresh scan, or dismissed on mobile. The shared frame supplies two of the three actions universally; the card body itself owns only the third:

  • Dismiss is the shared frame's Archive button — AlertInboxViewer renders the common InboxArchiveButton for every central alert card, so doc-token-alert no longer registers its own dispatcher handler. Archiving a still-firing card runs /src/main/services/doc-scan/doc-token-alert-card.ts's archive listener, which stamps the same sticky-to-size acknowledged_tokens acknowledge server-side, then auto-advances to the next attention row. doc-token-alert is still a reject-only integration — there is no approve handler.

  • Start session is the frame's universal Start-session button. The stock prompt is no longer built at click-time in the renderer — /src/main/services/doc-scan/doc-token-alert-card.ts builds it once, when the card is created, via buildDocTokenStockPrompt from /src/shared/doc-token-stock-prompt.ts, and stores it on the alert row for the frame to spawn from via IPC.QUICK_LAUNCH_SUBMIT. The prompt still names the three always-loaded files, reports current tokens / threshold / overage, and asks the agent to propose a slim-down plan and wait for approval.

  • Edit opens the largest counted always-loaded file in the FilePeekOverlay markdown editor (startInEdit: true). The card body resolves that target on mount by calling IPC.DOC_TOKEN_ALERT_RESOLVE_EDIT_TARGET, whose handler runs resolveDocTokenEditTarget(folderPath, getDocTokenSyncState()) in /src/main/services/doc-scan/doc-token-scan.ts — an lstat-only, largest-wins pick (alphabetical tie-break) over ALWAYS_LOADED_DOC_PATHS that skips a synced mirror (so you edit the canonical, not a copy sync would overwrite). The target is resolved at mount-time and never persisted (no migration, no stored path that could go stale). If none of the files exist yet — or the only one is an excluded mirror — the button is disabled and the card body says so. On save, /src/renderer/src/stores/file-peek-store.ts re-fires DOC_TOKEN_SCAN_REQUEST so a slimmed file clears the row automatically.

  • The per-doc breakdown is fetched on mount via the shared useAlertBreakdown hook (channel IPC.DOC_TOKEN_ALERT_RESOLVE_BREAKDOWN, extracted so every card family shares one fetch-and-retry pattern), whose handler still runs resolveDocTokenBreakdown(folderPath, getDocTokenSyncState()) — the same one lstat pass as the Edit-target resolver (shared listAlwaysLoadedFiles helper), returning each existing always-loaded file as { relPath, tokens, counted }. counted: false marks a synced mirror; the card body lists it dimmed and labeled "synced mirror, not counted". The counted-sum equals the headline always_loaded_tokens (the fire basis). Computed fresh every mount, never persisted.

The store stays a pure CRUD producer: it never imports session-store, inbox-actions, or the navigation modules (locked by /tests/unit/lint/approval-store-producer-consumer.test.ts). That is why the card body — not the store — owns the edit-target resolve; Dismiss and Start session now live in the shared AlertInboxViewer frame instead.

Push channel + app-level sync

IPC.DOC_TOKEN_ALERT_UPDATED is listed in INBOX_ALLOWED_CHANNELS in /src/main/services/web/web-access-push-filter.ts so mobile gets the push (the per-channel allow-list is the contract for what reaches mobile vs desktop-only). The renderer-side listener is registered in the app-level useCliPushSync hook (/src/renderer/src/hooks/useCliPushSync.ts) so it runs view-independent — the user gets the row no matter which view they're on when the scan completes.

Settings UI

/src/renderer/src/features/settings/sections/notifications/DocTokenAlertSettings.tsx renders the global defaults plus per-project overrides:

  • Global defaults — a master enable toggle, then (when enabled) three number inputs: threshold (default 40,000), re-alert growth % (default 50, range 10–500), and re-alert cadence (days) (default 30, range 0–365; 0 = off). Edits write to the 'global' row via DOC_TOKEN_RULE_UPSERT.
  • Per-project overrides — a list of project rows, each with its own enable toggle + a "Reset to global" button and a three-column grid of threshold / growth % / cadence inputs. Adding a new override picks any project that doesn't already have one and writes a new row (starting at the defaults). Removing one calls DOC_TOKEN_RULE_DELETE for the project id (the 'global' row is unreachable from this UI — the queries module throws if anyone tries).

The settings screen has no manual "Re-scan" control — scans are driven entirely by the startup fan-out, the session-turn debounce, and the Edit-then-save re-fire described above. There is no button to force a scan from here.

Related

  • claude-files-sidebar.md — sister feature using the same bytes/4 token-estimate heuristic; collapsible sidebar section gives one-click peek-overlay access to the same three always-loaded files plus their globals
  • inbox-overview.md — how unified-inbox rows surface across all integrations; the doc-token-alert row lives in this list
  • project-docs-auto-injection.md — separate token budget for <project>/.claude/docs/ files that auto-inject into every session's first message (a 125k cap there, not the always-loaded bundle this page covers)

Last verified 2026-09-28