---
title: AI-Drivable Browser (RETIRED — see My Real Chrome)
---

# AI-Drivable Browser (RETIRED — see My Real Chrome)

## What it is

> **RETIRED (owner decision 2026-08-02).** Omniscio consolidated its in-app browsers down to
> **[My Real Chrome](real-chrome-bridge.md)** (drive your own signed-in Chrome via a companion
> extension) + [Browser Logins](browser-logins.md). The AI-drivable browser is disabled + hidden;
> its code is kept for reference only. This page describes the historical feature.

> **This ships OFF.** The AI-drivable browser is gated behind the `ai-browser` Lab flag (Settings → Lab, off by default). While it's off, the Browser pane is exactly the plain [Embedded Browser](embedded-browser.md) — one page, no tabs, no agent access. Nothing on this page happens until you turn the flag on.

The AI-drivable browser turns Omniscio's embedded Browser pane into a real, multi-tab browser that an Omniscio **agent can drive for you** — open a page, read it, click, type, screenshot — all inside the pane, while you watch. It's Omniscio's answer to the "agent that can use a browser" idea, built so the agent works in its **own isolated tab** and can never touch your logged-in cookies.

## Where to find it

### Turning it on

1. **Settings → Lab → enable `ai-browser`.** (Or launch with `AMC_SHOW_AI_BROWSER=1` to reveal it.)
2. Make sure the **Browser** panel is available: Settings → Features → **Enable Browser**.
3. Open **Browser** in the Omniscio sidebar. You'll now see a tab strip instead of a single page.

## How it behaves

### Autofill vault (saved logins for your own tabs)

The **Passwords** tab in **History** is a small password manager for the browser pane's own tabs — not the agent's. Log into a site in one of your own tabs and Omniscio offers to save it: you type the username/password into a save dialog yourself (it doesn't capture what you typed on the page). Come back to that exact site later and a toast offers to fill it in for you. The list shows the site, username, and last-used date; remove a saved login any time. The password itself is encrypted through your OS's own credential store (Keychain / Credential Manager / libsecret) — the management panel only ever shows metadata, never the password.

This is a different feature from [Browser Logins](browser-logins.md), and the two don't share data:

|              | Autofill vault                                  | Browser Logins                             |
| ------------ | ----------------------------------------------- | ------------------------------------------ |
| Stores       | An encrypted password, in your OS keyring       | A cookie/session — never a password        |
| Works in     | Your own tabs, in the AI-drivable Browser panel | Any spawned agent (via the CLI) — and now the in-app agent tab / your own in-app tab (cookies-only, best-effort) |
| Reached from | Browser → History → Passwords                   | Settings → Browser Logins                  |
| Agent access | None — no tool, no CLI route reaches it         | The whole point — `amc-browser use <name>`, or seed the in-app agent tab |

The agent's own isolated tab never sees your saved passwords, by design — there's no tool and no CLI route that reaches this vault, the same isolation that keeps your cookies out of the agent's tab keeps your passwords out of it too. A saved credential only fills on the exact site (scheme + host + port) it was saved on — it never carries over to a different subdomain.

### Memories (ask a session about pages you've already seen)

Omniscio quietly indexes the pages you visit in your **own** tabs — never the agent's — chunking and embedding the text with the same local embedding model Omniscio uses elsewhere (nothing leaves your machine; a duplicate visit to the same content isn't re-indexed). A session can then search across everything you've browsed instead of you re-finding and re-opening it, with `browser_search_memory` — give it a query, get back the best-matching page chunks with URL, title, matched text, and when you visited, ranked by relevance. "What was that pricing page I looked at last week" becomes one tool call instead of digging through history.

Manage it from **History → Memories**: browse what's indexed, delete individual entries, clear everything, or turn indexing off entirely (it's on by default).

### Journeys (your browsing, auto-grouped by topic)

**History → Journeys** groups your recent browsing into topic clusters automatically — nothing to save by hand. It's computed from the same page embeddings Memories uses, bucketing pages that are topically close together and labeling each group after its most-recently-visited page. It's a lightweight, on-the-fly grouping (computed when you open the tab, not a background job or a precise research-session tracker) — good for "what was I looking at when I was researching X."

### Cursor chat & Compose (rewrite or draft text in any field)

The **Write in field** item in the **✨ AI tools** menu captures whatever text field is focused in your own tab. Select text first and it opens in **rewrite** mode (tell it how to change the selection); focus an empty or unselected field and it opens in **compose** mode (tell it what to draft). Type your instruction in the prompt bar and submit — like Design mode, this needs a session bound in the side chat.

The bound session gets the field's context plus your instruction, then writes the result back into the page via `browser_write_inline_text`. The write-back is authorized by a single-use token, scoped to that exact tab and expiring in about two minutes, so nothing else can reach into your page and write to it. On plain `<input>`/`<textarea>` fields, the write-back is inserted the same way real typing is, so **Ctrl+Z undoes it** like anything you typed by hand. Rich text editors (Gmail, Notion, and similar contenteditable surfaces) are best-effort — the exact selection isn't always restored, so treat undo there as a bonus, not a guarantee. Password fields are never captured.

### Workflow recording (turn what just happened into a skill)

Two ways to record a workflow into a skill — same result, different actor:

- **You record.** The **Record workflow** item in the **✨ AI tools** menu captures your clicks, typed values, form submissions, and page navigations in your own tab while it's on. While recording, a **Stop** button surfaces in the toolbar itself (never buried in a menu) so you can always end it. Hit Stop and a review panel lists every typed value it saw; **you** decide which ones become `{{variables}}` in the resulting skill — nothing is tagged automatically. A value that looks like a password shows up redacted and prompts you to tag it if you want it replaced with a variable rather than baked in literally.
- **The agent records.** A session driving its own tab can call `browser_start_recording` / `browser_stop_recording` to capture its own navigate/click/type/submit sequence the same way, ending with the same review-and-tag step before saving.

Either path saves a draft **skill** — the same shape the Skills gallery (below) already understands, so it shows up there ready to run or edit. The honest ceiling: it's a literal, numbered replay of what happened, generalized only by the variables tagged at save time — it doesn't infer a broader workflow from a single recording.

### Side chat pane

A collapsible panel on the right side of the Browser view holds a live chat with whichever session you pick — the same chat surface as the main Dashboard, just docked next to the browser so you're not switching views to talk to the agent that's driving (or co-piloting) the tab. This is the session Visual Edit, Design mode, and Cursor chat all hand off to. Toggle it from the command palette or its own collapse control.

While an agent drives the browser, a small **activity ticker** under the pane header narrates each step in plain English ("Opened example.com", "Clicked e12", "Couldn't click — element not found") — the last five actions, failures tinted red. The text an agent types or fills is **never shown** in these lines (it may be your passwords or personal data); only the target field and the outcome appear. Below the chat, quiet **suggestion chips** appear when the composer is empty — rule-based shortcuts for the current page ("Get the transcript" on a YouTube video, "Review this PR" on GitHub, "Summarize this page" anywhere). Clicking a chip only pre-fills the composer; **nothing is ever sent until you press send**, so a stray click never costs a paid call.

### Skills (save a workflow as a one-click command)

A **skill** is a browser workflow you save in plain English and re-run in one click — "Summarize this page", "Find the cheapest flight and add it to my cart", "Pull the action items off this doc". It's the Omniscio take on Dia's Skills / Comet's Shortcuts.

Open it from the **History toolbar button → Skills tab** in the browser. There you can create a skill (a name, the plain-English steps, and an optional emoji for its card), edit or delete existing ones, and run any of them. Typing one out isn't the only way in — see [Workflow recording](#workflow-recording-turn-what-just-happened-into-a-skill) above for recording one instead.

Under the hood a skill **is a recipe** — there's no separate "skill engine". Saving a skill writes a browser-scoped recipe (a normal Omniscio recipe flagged as a browser skill), so it shows up in the Skills gallery. Running one goes through the **exact same gated run path** as any recipe: it opens the standard Run dialog (pick a project, fill any inputs, confirm) and starts a recipe run — it never spawns anything directly. That means every safety rule above still applies while a skill runs: an acting skill routes its plan through `request_plan_approval` (the inbox approval described in "Plan preview / approval"), and each action it takes still obeys attended-only writes, category watch-mode, per-site tiers, and the Pause switch. A skill is a shortcut to a workflow, not a way around the guardrails.

v1 is your own skills only — there's no community gallery or marketplace yet.

### Video pages (YouTube Q&A)

A "what does this video say" / "summarize this YouTube video" request from a session driving the agent tab: `browser_navigate` to the video URL, then `browser_get_video_transcript` — one call, no clicking through the page. It reads YouTube's own caption-track list off the page and fetches the chosen track's caption data directly, so it works even when the "Show transcript" panel is never opened. Only works on a YouTube **watch** page (`youtube.com/watch?v=…`) with at least one caption track, and only while `ai-browser` is enabled (Settings → Lab) — anything else (a non-YouTube site, a channel/Shorts page, a video with no captions) is a friendly error, not a best-effort scrape.

Two other paths do **not** cover this, for context:

1. **The `summarize` skill does not cover this.** For a URL it does a plain `WebFetch` of the page HTML looking for "the main article text" — on a YouTube watch page that's just the title/description/JS shell; the transcript loads client-side and is never in the initial HTML `WebFetch` sees.
2. **`POST /download` (yt-dlp) does not fetch captions either.** Its `DownloadMode` is `video | audio | original` only (`src/main/services/conversion/download-types.ts`) — it saves a media _file_, never subtitle/transcript text, and nothing pipes a downloaded video through the app's whisper transcriber (that pipeline only transcribes local screen recordings — a separate feature).

If `browser_get_video_transcript` ever fails on a page that visibly has a transcript (a YouTube shape change), the old manual workaround still applies as a fallback: `browser_snapshot` (or `browser_get_page_content`) to locate the **"Show transcript"** control under the description (may need expanding "...more" first), `browser_click` its ref, then `browser_get_dom` with a CSS selector on the transcript panel (e.g. `ytd-transcript-segment-list-renderer`) to read the rendered timestamp+text segments — or `browser_network` to spot the transcript/timedtext request the click just fired, then fetch that URL directly.

### PDF / docs pages

A "what does this PDF say" / "ask about this PDF" request from a session driving the agent tab: `browser_get_page_content` (and `browser_read_tabs`) detect a PDF — by response content-type or a `.pdf` URL — and refuse with a `PdfPageError` instead of returning garbage, because a PDF has no DOM for Readability to read. The error names the fix: `browser_save_pdf` saves the active tab into Downloads as a real `.pdf` file (it calls CDP's `Page.printToPDF` directly — when the tab is already showing a PDF via Chromium's built-in viewer, that streams back the _original_ PDF bytes, not a re-rendered copy, so selectable text survives), then the session runs the **pdf-extract** skill's `Read` tool on that saved path. Two calls, no new tool needed — `save_pdf` already existed for the Band-1 export pack; PDF pages are just its second use.

Two other paths do **not** cover this, for context:

1. **`POST /download` (yt-dlp) does not fetch PDFs.** Same as the video-transcript case above: its `DownloadMode` is `video | audio | original` only, aimed at yt-dlp's media-site extractors — pointed at a plain PDF URL it has no matching extractor and fails. It's not a generic "download this URL to disk" endpoint.
2. **Google Docs / Office docs opened as a live web page are not handled by `get_page_content` either** — `docs.google.com/document/…` is a JS app, not a PDF, so it won't hit `PdfPageError`, but its editor view often has little or no extractable DOM text (canvas-rendered), so Readability's plain-text fallback can come back thin or empty. Omniscio has a dedicated path for this instead: `POST /gdoc/import` (`src/main/services/cli/cli-server-gdoc-routes.ts`) pulls a Google Doc's real content as Markdown via the Drive API — use that for a Google Doc, not the browser tools.

### Automations & co-work

Three common jobs are shipped as **recipe templates** in the Pattern Library (Recipes → New from pattern), not as new machinery. Each is a normal recipe built over the browser tools above plus the session's own tools; there is no new scheduler, no new engine, and no new file route. Pick one, fill in the blanks, and it becomes a runnable recipe.

#### Recurring browser jobs

"Every morning, open these dashboards, pull the numbers, and write me a summary." That is the **Recurring browser report** pattern (`recurring-browser-report`): a one-step recipe whose prompt tells the agent to `browser_navigate` + `browser_get_page_content` over a list of URLs, extract what you asked for, and save a dated markdown summary via its own `Write` tool. It is read-only by design (no clicking or submitting), which is what makes it safe to run unattended.

To make it _recurring_, wire the instantiated recipe to a cron job on the **existing** cron system — nothing new is built for scheduling:

1. Instantiate the pattern into a recipe (it lands in your recipe list with an id).
2. Create a recurring cron job of `type: "recipe"` pointing at it. Over the CLI control server that is a `POST /cron/jobs` with `runMode: "recurring"`, a `cronExpression` (e.g. `0 8 * * *` for 8am daily), and `jobConfig: { "recipeConfigId": "<the recipe id>" }`. The cron **recipe** executor (`cron-recipe-executor.ts`) already dispatches it on schedule.
3. A cron job created via the API always registers as **pending approval** — you approve it once in the inbox before it starts firing.

Because a scheduled run is unattended, the Band-2 safety layer applies with no extra work: the moment the agent tries to _act_ (not just read) on a **finance / payments / social** site while nobody is watching, the permission gate denies it with `watch-mode-required` (`ai-browser-permission-evaluate.ts`). So a recurring job is fine for reading dashboards, prices, or status pages, and cannot quietly buy or post anything. For a multi-step unattended run, have the recipe call `browser_request_plan_approval` up front so its plan is queued for you to approve.

#### Research to report (co-work)

The **Browser research to report** pattern (`browser-research-to-report`) is the co-work shape: one session researches a topic with `browser_navigate` + `browser_get_page_content` + `browser_read_tabs`, cross-checks across sources, then writes a cited markdown report with its **own** `Write` tool to its working directory. No file route or share pipeline is involved — the report is just a file the session wrote where you can open it. Run it as a normal session (you watch and steer) or on a schedule like above.

#### Guided shopping / booking (confirm-gated)

The **Guided shopping** pattern (`guided-shopping`) is a three-step, confirm-gated recipe for "compare these, then buy the best one":

1. **Compare** — the agent reads listings/prices across the sites you named and produces a shortlist plus a proposed purchase plan. It changes nothing, and it calls `browser_request_plan_approval` so the plan is queued for you.
2. **Approval gate** — the recipe pauses at a human approval step. Nothing proceeds until you approve.
3. **Checkout** — only after approval, the agent adds the item to the cart and drives to the checkout page, then **stops at the payment step**. It never enters card/address details and never places the order; you review and pay by hand.

This is the load-bearing invariant, enforced at two layers: the recipe's own approval gate sits before the checkout step, and underneath it the Band-2 permission gate **never** lets any agent complete a payment unattended (`no unattended payments, ever`). Booking flows follow the same shape — compare, approve, drive up to the confirm/pay step, hand off.

### Privacy + safety notes

- The agent tab is **isolated** — it uses its own profile and never reads your `persist:browser` cookies.
- **History and bookmarks are yours** — the agent's navigation is never recorded to your history, and the agent can't clear it.
- **Memories only indexes your own tabs** — the agent's tab is never included in what gets indexed or what `search_memory` can find.
- **The autofill vault is invisible to the agent** — no tool or CLI route reaches it; it's exclusively for your own tabs.
- **The agent can only navigate to `http:`/`https:` pages** — `navigate` and `tab.new` reject `file:`, `data:`, `about:`, and similar schemes, so the agent can't be tricked into reading a local file or acting from a null-origin page.
- Everything is **local** — no browsing data leaves your machine, and the `/ai-browser/tool` route is bound to `127.0.0.1` behind the CLI bearer token.
- The embedded page is treated as **untrusted** and runs sandboxed, exactly like the plain Embedded Browser — the AI-drivable mode doesn't weaken that.

### What's still coming

Design mode's element reordering is button-stepped (move up/down among siblings), not drag-and-drop. Cursor chat's rewrite/compose is best-effort inside rich text editors (Gmail, Notion, and similar contenteditable surfaces) — plain `<input>`/`<textarea>` fields are the fully-supported case. Workflow recording generalizes only the values explicitly tagged as variables at save time; it doesn't infer a broader workflow from one recording. Everything else on this page — multi-tab browsing, agent driving in an isolated tab, the full permission/safety layer, Design mode, Cursor chat, and workflow recording — is complete behind the flag.

## For agents

### Where it lives in the code

- **UI:** `src/renderer/src/features/browser/BrowserView.tsx` (tab strip, toolbar, bars) + `TabWebview` (per-tab web view), with `BookmarksManagerModal.tsx`, `HistoryModal.tsx` (the six-tab History/Memories/Passwords/Journeys/Permissions/Skills modal), `DownloadShelf.tsx`, `visual-edit.ts` (the injected page inspector), and `DesignModeStylePanel.tsx` (the style panel).
- **Tool registry:** `src/main/services/ai-browser/tools/index.ts` — the single source of truth all 40 tools, the MCP server, and the `/ai-browser/tool` HTTP route derive from.
- **Agent driver:** `src/main/services/ai-browser/ai-browser-driver.ts` (the tool vocabulary over CDP), gated by `ai-browser-policy.ts` (the auto/allowlist/off + pause gate) and `ai-browser-permission-evaluate.ts` (per-site tiers + category watch-mode + attended check).
- **Isolated partition:** `AI_BROWSER_AGENT_PARTITION` in `src/shared/ai-browser.ts`.
- **Autofill vault:** `src/main/services/ai-browser/autofill-vault-service.ts` (human tabs only — no tool or CLI route reaches it).
- **Memories / Journeys:** `src/main/services/ai-browser/browser-index-service.ts` (indexing + `search_memory`) + `journeys-service.ts` (topic clustering).
- **Spaces:** `src/main/services/ai-browser/ai-browser-spaces-service.ts`.
- **Cursor chat / inline write-back:** `src/main/services/ai-browser/tools/write-inline-text.ts` + `inline-write-token.ts`.
- **Workflow recording:** `src/main/services/ai-browser/ai-browser-recording.ts` + `tools/start-recording.ts` / `stop-recording.ts` (agent-tab path) and `src/renderer/src/features/browser/use-browser-recording.ts` (human-tab path).
- **CLI routes:** `src/main/services/cli/cli-server-ai-browser-routes.ts` (`/ai-browser/tool`) + `cli-server-ai-browser-bookmarks-routes.ts`.
- **Flag:** the `ai-browser` entry in `src/shared/unreleased-features.ts` (`status: 'in-development'`).
- **Invariants + safe-change checklist:** `.claude/memory/contracts/ai-browser-contract.md`.

## Related

The surfaces that replaced this retired feature are [My Real Chrome](real-chrome-bridge.md), which drives your own signed-in Chrome through a companion extension, and [Browser Logins](browser-logins.md) for reusing a captured session. The plain single-page pane it was built on is [Embedded Browser](embedded-browser.md), the approval queue its plan preview posts through is [CLI Pending Actions](cli-pending-actions.md), and the local HTTP surface its raw tool route lives on is [cli-control.md](cli-control.md).

Material that outgrew this page — the two point-at-the-page tools (Design mode and Visual Edit), spaces, tabs and quick navigation, and the agent's own isolated tab with its MCP tool vocabulary and permission tiers — continues in [AI Browser (part 2)](ai-browser-part-2.md).
