Omniscio documentation
Browse all documentation
  1. Getting Started13
  2. Sessions & Agents115
  3. Inbox & Notifications59
  4. Projects & Tasks95
  5. Automation & Scheduling75
  6. Knowledge & Memory26
  7. AI Features60
  8. Integrations100
  9. Plugins & Marketplace33
  10. Cloud & Teams56
  11. Settings & Customization58
  12. Account & Billing28
  13. Troubleshooting84
  14. CLI & API Reference22
  15. Legal & Policies4
  16. Uncategorised22

Drive Omniscio Cockpit

A developer and QA tool for watching an AI operate Omniscio's own interface — clicking, navigating and exercising features toward a goal you type, with its reasoning narrated beside the app and the ability to take the wheel.

What it is

A developer/QA tool that lets you watch an AI operate Omniscio's interface — clicking buttons, opening Settings, navigating sessions, exercising a feature like a person would — toward a goal you type, with its reasoning narrated right beside the app, and the ability to grab the wheel and redirect it anytime. It is the Playwright-MCP "watch it drive a browser" experience, but the app being driven is Omniscio itself.

This is a dev tool, not a shipped end-user feature. It lives in the test tooling (the drive-amc skill + the sandbox scripts). It is invoked explicitly by a developer in a Claude session, never auto-triggered, and never available to end users.

What you see — the "cockpit" is two windows, both already built

There is no new cockpit UI. The cockpit is two existing Omniscio windows side by side:

┌─ YOUR REAL Omniscio ──────────────────┐     ┌─ SANDBOX Omniscio (headed) ───────────┐
│  Driving session (popped out)    │     │  The app being operated:         │
│  • its CHAT = the live action    │ ──▶ │  real clicks land, views change  │
│    log (what it's doing + why)   │drive│  (its OWN database — NOT your     │
│  • its COMPOSER = your steering  │     │   real sessions / data)          │
│    box (type a redirect anytime) │     │                                  │
└──────────────────────────────────┘     └──────────────────────────────────┘
  • The brain is a normal Claude session in your real Omniscio. Its chat stream is the live action log; its composer is the steering wheel.
  • The stage is a headed sandbox — a second, fully isolated Omniscio (its own database, app name omniscio-claude-sandbox) launched visible instead of hidden. It docks to one half of your screen so the driving session's window sits beside it.
  • The driving session operates the sandbox over Chrome DevTools Protocol — the same remote-control channel the E2E harness uses — so clicks land on real elements, not guessed pixels.

Where to find it

How to use it

  1. In a Claude session inside the Omniscio project, invoke the skill with a goal: /drive-amc test that snoozing an inbox item works.
  2. The skill boots the headed sandbox (npm run sandbox:launch:headed) if it isn't already up. The window appears, docked to the right half of your screen.
  3. Pop this driving session out beside it — right-click the project in the sidebar → Open in new window, snap it to the left half. Omniscio remembers that position per project, so it's a one-time setup.
  4. Watch: the agent narrates each step in its chat ("Opening Settings → Appearance → flipping the theme…") and you see it happen in the sandbox window.
  5. Steer anytime by typing in the composer — "now try the other toggle", "go check the inbox instead". The agent reads it as its next turn and adapts. Esc interrupts a run mid-loop.

How it behaves

Safety — sandbox-only, and it costs real tokens

  • It can only ever drive the sandbox. Every action goes through scripts/sandbox-cdp-driver.cjs, which attaches only to the sandbox's CDP endpoint (127.0.0.1:19222) and refuses to run unless the sandbox CLI server confirms instanceId=claude-sandbox. Your real install exposes no CDP endpoint at all, so the agent cannot touch your real app or data. This is the bright line from Omniscio's data-safety rules, made structural.
  • The driving session is a normal paid Claude session — that's expected. Free actions: navigating, opening settings, toggling, screenshotting. Paid action: clicking "New session" / submitting a prompt inside the sandbox spawns a real Claude session and bills real tokens (the sandbox uses a real account). Default mode never does this; spawning is an explicit, per-run, cost-gated opt-in (the agent must state which session, what it's verifying, and the expected cost first — the amc-sandbox rule).
  • No surprise outbound. The sandbox's email/SMS/Drive integrations stay unarmed unless you deliberately enable one, so "test the email feature" can't fire a real email by accident.

While you're away

  • It keeps drawing when its window is covered or your screen is off, so an agent can drive it — and send you screenshots — with nobody at the desk. A minimized sandbox window still stops drawing.
  • Only the sandbox is started this way; your own Omniscio keeps saving power when it is covered.

For agents

How the headed launch + dock work

The sandbox normally runs hidden (AMC_SANDBOX_HEADLESS=1 → BrowserWindow({ show: false }), driven by CDP only). npm run sandbox:launch:headed sets AMC_SANDBOX_HEADED=1, which the launcher turns into a visible launch and a longer readiness wait (a cold headed boot under load can take >45s, so the old fixed 45s poll bailed before writing the PID file). On show, the main window — instead of maximizing — docks to one half of the primary display's work area (computeHeadedSandboxBounds), default the right half; set AMC_SANDBOX_HEADED_SIDE=left to flip. The driving session's popped-out window takes the other half; its position persists per project, so after the first setup the two land side by side on their own.

Every launch also passes Chromium's --disable-backgrounding-occluded-windows (sandboxElectronArgs), the switch Playwright and Puppeteer use: Chromium stops drawing a page it judges occluded, and the driver's clicks wait on animation frames, so a covered or screen-off sandbox would otherwise stall them. It was added after a drive stalled with the user away on 2026-09-26; that stall could not be recreated, so its cause is unconfirmed — if a click still hangs at "waiting for element to be visible, enabled and stable", capture document.visibilityState and the window state before retrying.

The driver — the agent's hands

npm run sandbox:cdp -- <op> [args] performs ONE UI operation per call so the agent can act and narrate step by step:

  • snapshot — print the on-screen accessibility tree + the data-ui-anchor ids currently visible.
  • screenshot [file] — capture a PNG of the window.
  • role <role> "<name>" / anchor <id> / text "<text>" / click "<css>" — click an element (prefer role/anchor — Omniscio ships data-ui-anchor ids precisely so automation is reliable).
  • type "<text>" / press <Key> / wait <ms> — type, press a key, or let an animation settle.

Every action leaves a fresh screenshot at a known path so the agent can "look" at the result before its next step. It never closes the window (it only disconnects the CDP client).

What's reused vs new

Reused (~90%): the sandbox's isolation + real-credential sourcing (the amc-sandbox skill), a normal Omniscio session's chat / composer / pop-out window, and the CDP remote-control the E2E harness already uses. New (small): the first-class headed launch + half-screen dock, the reusable CDP driver, and the drive-amc skill that runs the snapshot→decide→act→narrate loop and takes your steering.

Known follow-up: the dock is one-time-manual (pop out + snap once; position then persists). A CLI route into the pop-out feature now ships — POST /window/open opens (or focuses) any project or integration in its own window by id (see the omniscio-control windows surface / cli-control.md). What's still manual for THIS dev tool is narrower: that route pops out a whole project/integration window, not a single detached session, and there is no CLI positioning route — so half-screen placement stays a one-time manual snap. Auto-popping the driving-session window specifically + positioning it remains the open follow-up; the generic project pop-out itself is no longer out of scope.

Where it lives in the code

  • Skill (the brain's instructions + loop): .claude/skills/drive-amc/SKILL.md.
  • Headed launch + readiness widening: scripts/sandbox-launch.cjs (resolveLaunchMode) + scripts/sandbox-shared.cjs. npm scripts: sandbox:launch:headed, sandbox:up:headed, sandbox:cdp.
  • Half-screen dock: src/main/app/window.ts (the ready-to-show handler) + src/main/window-bounds.ts (computeHeadedSandboxBounds).
  • The driver (hands): scripts/sandbox-cdp-driver.cjs.
  • Tests: tests/unit/scripts/sandbox-launch-mode.test.ts, tests/unit/scripts/sandbox-cdp-driver.test.ts, tests/unit/window-bounds.test.ts, tests/unit/sandbox-headless-gate.test.ts.

Related

  • The built-in terminal — the other way an agent runs commands on this computer.
  • Running apps — the dev-server preview panel the launch path shares its routes with.
  • The amc-sandbox skill — the Claude-driven sandbox this cockpit is a fourth, narrow, watch-and-steer mode layered on (its cost guardrail applies here too).

Last verified 2026-10-06