---
title: AI-Drivable Browser (RETIRED — see My Real Chrome) (part 2)
---

# AI-Drivable Browser (RETIRED — see My Real Chrome) (part 2)

## What it is

This is part 2 of the [AI-Drivable Browser (RETIRED — see My Real Chrome)](ai-browser.md) page. It carries what the feature added over the plain embedded browser, the two point-at-the-page tools (Design mode and Visual Edit), spaces and quick navigation, and the agent's own isolated tab — its MCP tool vocabulary, the plan approval queue, and the permission tiers that decide what it may do.

## Where to find it

The feature lived in the Browser pane, opened from the Omniscio sidebar once the `ai-browser` Lab flag was on; [AI-Drivable Browser (RETIRED — see My Real Chrome)](ai-browser.md) covers the rest of it, including how to turn it on, the autofill vault, Memories, Journeys, cursor chat, side chat, skills, video pages, PDF pages and co-work. The agent-facing controls named here sit in the Browser panel itself, under History → Permissions.

## How it behaves

### What you get (over the plain Embedded Browser)

- **Real tabs.** A tab strip across the top — open, switch, and close as many pages as you like, each its own web view. `window.open()` / `target="_blank"` links open as a new in-app tab instead of being blocked.
- **A bookmarks bar + manager.** Click the **star** in the address bar to bookmark the current page; bookmarks show as a bar under the toolbar. The **Manage** button opens a modal to rename, reorder, and remove them.
- **Browsing history.** Pages you visit in **your own** tabs are recorded; the **History** button opens a modal with six tabs — **History** (click to reopen, "Clear browsing history" with confirmation), **Memories**, **Passwords**, **Journeys**, **Permissions**, and **Skills** — the last five are covered in their own sections below. The address bar suggests from your history + bookmarks as you type.
- **A download shelf.** Downloads appear in a strip at the bottom with progress, "show in folder", and discard. A file type that can run code prompts you to approve before it saves.
- **Find in page** (in the **⋯** overflow menu or Ctrl+F), **zoom** (Ctrl +/−/0), and the usual **keyboard shortcuts** (Ctrl+T new tab, Ctrl+W close, Ctrl+L address bar, Ctrl+R reload).
- **A collapsed, progressive-disclosure toolbar.** At rest the address bar's right side shows just four controls — **✨ AI tools** (the four page-action modes below), **side chat**, **Library** (the History/Memories/… hub), and a **⋯** overflow (Find, command palette, open-in-new-window, tab-strip orientation) — plus the bookmark star inside the address bar. Pause appears only when an agent is in play; a Stop control surfaces while you're recording.
- **Point-and-prompt Visual Edit.** The **Visual edit** item in the **✨ AI tools** menu lets you click any element on the page and describe a change ("make this button blue"); if you have a session bound in the side chat, Omniscio sends the request straight to it — no copy-paste needed. With no session bound, it falls back to copying the same prompt to your clipboard. A separate **Design mode** tool goes further — see [Design mode & Visual Edit](#design-mode--visual-edit) below.
- **Spaces, a command palette, and quick tab switching.** Tab groups, Ctrl+K search, and Ctrl+Tab most-recently-used cycling — see [Spaces, tabs, and quick navigation](#spaces-tabs-and-quick-navigation) below.

### The agent's own tab (the important part)

When the feature is on, the browser runs **two kinds of tabs**:

- **Your tabs** — normal browsing, using the same `persist:browser` cookies as the plain Embedded Browser (so your logins carry over).
- **The agent's tab** — a single tab on a **separate, isolated profile** (`persist:ai-browser-agent`). It has its **own** cookies and never sees yours. It's marked in the tab strip and you can't close it (the agent's controller lives on it).

This is deliberately **safer than a shared-session agent browser**: the agent can log into sites in its own profile, but it can never act as _you_ on a site you're signed into. If the agent needs a login, that's a per-site thing in its own profile.

**The agent's tab stays alive in the background (persistent agent browser).** Once you've opened the Browser panel, it stays loaded (hidden) when you switch to another project or session, instead of being thrown away — so the agent's tab keeps working and can be driven even while you're looking at something else. Without this, switching away used to destroy the agent's tab, so an agent trying to drive the browser would find "no tab" until you came back. You can turn it off with **"Keep the agent's browser running in the background"** in **History → Permissions** (on by default); off means the tab is dropped when you leave the panel, the old behavior. One caveat: an agent `screenshot` needs the panel actually visible (a hidden tab can't be painted) — the agent's normal perception (the accessibility snapshot) works fine while hidden, so navigating and clicking are unaffected.

##### Handing the agent tab a saved Browser Login

If you use [Browser Logins](browser-logins.md) (the separate, also-in-development feature where you
log into a site once and every agent can reuse it), you can point that saved login at the **in-app**
agent tab too — so the in-app agent acts as that login. It's **one saved login at a time**: choosing
another re-seeds the same agent tab, replacing the first. The isolation still holds — the agent tab
is only ever seeded from a login **you** captured, never from your own `persist:browser` browsing, so
handing it a login still can't let the agent act as *you*. You can also open a saved login in one of
**your own** tabs; that opens on its own per-login space, kept apart from your normal browsing and
never reachable by the agent.

**Caveat (best-effort, cookies-only):** the in-app browser is a different engine than the real-Chrome
clone external CLI agents use, so a saved login is applied here by copying its **cookies** (the full
session, http-only cookies included) — not `localStorage`/`IndexedDB`, and it can't refresh a rotating
token. Ordinary cookie-session sites work; some SPAs and **certain Google surfaces** may still show
logged-out in-app and need a manual re-login there. External CLI agents are unaffected (they keep the
full-profile clone). Full detail + the exact partitions: [Browser Logins](browser-logins.md#using-a-saved-login-in-the-in-app-browser-agent-tab--your-own-tabs).

#### How an agent drives it — MCP tools (the normal way) or the raw HTTP route

Any Omniscio session spawned while the flag is on automatically gets a bundled **`ai-browser` MCP server** — no `.mcp.json` edit needed. Its tools appear to the session as native tools named `browser_<name>` (dots become underscores, e.g. the tab-list tool is `browser_tab_list`):

| MCP tool                        | What it does                                                                                                                                                                                                                                                                                                                                         |
| ------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `browser_navigate`              | Navigate the active agent browser tab to a URL.                                                                                                                                                                                                                                                                                                      |
| `browser_snapshot`              | Get an accessibility-tree snapshot of the active agent tab, with ref handles (e.g. "@e1") for interactable elements.                                                                                                                                                                                                                                 |
| `browser_click`                 | Click an element in the active agent tab by its snapshot ref (e.g. "@e3").                                                                                                                                                                                                                                                                           |
| `browser_type`                  | Type text into an element in the active agent tab by its snapshot ref.                                                                                                                                                                                                                                                                               |
| `browser_screenshot`            | Capture a screenshot of the active agent tab as a PNG data URL.                                                                                                                                                                                                                                                                                      |
| `browser_console`               | Get the last 50 console log entries from the active agent tab.                                                                                                                                                                                                                                                                                       |
| `browser_network`               | Get the last 50 network requests from the active agent tab.                                                                                                                                                                                                                                                                                          |
| `browser_get_page_content`      | Get the active agent tab's page content as clean Markdown via Readability, with a plain-text fallback.                                                                                                                                                                                                                                               |
| `browser_get_video_transcript`  | Get a YouTube video's transcript as timestamped `[m:ss] text` lines, from the caption track matching `lang` (or the video's default track). YouTube watch pages only.                                                                                                                                                                                |
| `browser_read_tabs`             | Read content from several agent tabs in one call — same extraction as `get_page_content`, run per tab (default: all open agent tabs, capped at 8; pass `tabIds` for specific ones). Each tab is checked and extracted independently — a tab that's no-access or fails to extract comes back as `{ tabId, error }` instead of failing the whole call. |
| `browser_tab_list`              | List the agent's open browser tabs.                                                                                                                                                                                                                                                                                                                  |
| `browser_tab_new`               | Open a new agent browser tab, optionally navigating to a URL (defaults to Google).                                                                                                                                                                                                                                                                   |
| `browser_tab_select`            | Select which agent tab subsequent tool calls address.                                                                                                                                                                                                                                                                                                |
| `browser_tab_close`             | Close an agent browser tab by tabId.                                                                                                                                                                                                                                                                                                                 |
| `browser_hover`                 | Hover the mouse over an element in the active agent tab by its snapshot ref, without clicking.                                                                                                                                                                                                                                                       |
| `browser_focus`                 | Move keyboard focus to an element in the active agent tab by its snapshot ref, without clicking or typing.                                                                                                                                                                                                                                           |
| `browser_clear`                 | Clear all text from an input or textarea in the active agent tab by its snapshot ref.                                                                                                                                                                                                                                                                |
| `browser_check`                 | Check a checkbox or radio button in the active agent tab by its snapshot ref (no-op if already checked).                                                                                                                                                                                                                                             |
| `browser_uncheck`               | Uncheck a checkbox in the active agent tab by its snapshot ref (no-op if already unchecked).                                                                                                                                                                                                                                                         |
| `browser_select_option`         | Select one or more options in a select element in the active agent tab by its snapshot ref.                                                                                                                                                                                                                                                          |
| `browser_press_key`             | Press a keyboard key in the active agent tab, optionally with modifiers (Ctrl, Shift, Alt, Meta).                                                                                                                                                                                                                                                    |
| `browser_scroll`                | Scroll the active agent tab by direction or bring an element into view by its snapshot ref.                                                                                                                                                                                                                                                          |
| `browser_fill_form`             | Fill several fields in the active agent tab in one call via ref/value pairs.                                                                                                                                                                                                                                                                         |
| `browser_upload_file`           | Set the selected files for a file-input element by snapshot ref.                                                                                                                                                                                                                                                                                     |
| `browser_handle_dialog`         | Arm a one-shot response for the next JS dialog (alert/confirm/prompt) in the active agent tab.                                                                                                                                                                                                                                                       |
| `browser_drag`                  | Drag from one element to another in the active agent tab by snapshot ref.                                                                                                                                                                                                                                                                            |
| `browser_wait_for`              | Wait in the active agent tab until text appears, a CSS selector matches, or the URL contains a substring.                                                                                                                                                                                                                                            |
| `browser_save_pdf`              | Save the active agent tab as a PDF into the Downloads folder.                                                                                                                                                                                                                                                                                        |
| `browser_save_screenshot`       | Save a screenshot of the active agent tab as a PNG into the Downloads folder.                                                                                                                                                                                                                                                                        |
| `browser_get_page_links`        | List every link (href and visible text) on the active agent tab, deduped by URL.                                                                                                                                                                                                                                                                     |
| `browser_get_dom`               | Get the outerHTML of the first element matching a CSS selector on the active agent tab.                                                                                                                                                                                                                                                              |
| `browser_search_dom`            | Find elements in the active agent tab matching text, a CSS selector, or an XPath expression.                                                                                                                                                                                                                                                         |
| `browser_start_dev_session`     | Start streaming the active agent tab's console logs and network request summaries (method/url/status/type/duration, never bodies) into two NDJSON files under Downloads/ai-browser-dev, so they can be grepped or tailed instead of dumped into context.                                                                                             |
| `browser_stop_dev_session`      | Stop a capture started by `start_dev_session` and return the final file paths and line counts (and whether either file hit its ~5MB cap).                                                                                                                                                                                                            |
| `browser_list_dev_servers`      | List Omniscio's running dev-server previews (Running Apps): name/url/port for each, read-only. Pair with `navigate` to open one.                                                                                                                                                                                                                     |
| `browser_search_memory`         | Semantically search your indexed browsing history (your own tabs only, never the agent's) for the best-matching page chunks — returns URL, title, matched text, and visit time. See "Memories" below.                                                                                                                                                |
| `browser_request_plan_approval` | Post a proposed multi-step plan (a title + a numbered checklist of 2-20 steps) for human approval before running it autonomously. Returns `{ status, approvalId }` — see "Plan preview / approval" below.                                                                                                                                            |
| `browser_start_recording`       | Begin recording the agent's own navigate/click/type/submit actions on its own tab, for turning into a draft skill. See "Workflow recording" below.                                                                                                                                                                                                   |
| `browser_stop_recording`        | Stop a capture started by `start_recording` and return the recorded sequence as a draft browser skill.                                                                                                                                                                                                                                               |
| `browser_write_inline_text`     | The return half of Cursor chat / Compose — write session-produced text back into the human field it came from, authorized by the single-use token the prompt bar minted. Not something you'd call standalone; the session calls it as part of that flow. See "Cursor chat & Compose" below.                                                          |

The agent **perceives** with `snapshot` (or `browser_get_page_content` to read the text) and **acts** by `ref` — the same perceive-then-act model the Playwright-MCP tools use. This tool list is generated automatically from Omniscio's internal tool registry, so it's always current — every tool above ultimately calls the exact same code (and obeys the exact same gates below) whether it's reached over MCP or over the raw HTTP route.

An outside tool without an MCP client (or a script you're writing by hand) can call the same tools directly against the CLI control server instead:

```
POST http://127.0.0.1:19519/ai-browser/tool
{ "tool": "navigate", "args": { "url": "https://example.com" } }
```

The `tool` field takes the same names the MCP tools are derived from (`navigate`, `snapshot`, `click`, `type`, `screenshot`, `console`, `network`, `get_page_content`, `start_dev_session`, `stop_dev_session`, `list_dev_servers`, `search_memory`, `request_plan_approval`, `start_recording`, `stop_recording`, `write_inline_text`, or the tab ops `tab.list` / `tab.new` / `tab.select` / `tab.close`). The route returns `404` until the flag is on. (Agent-facing usage notes also live in the omniscio-control skill's `ai-browser` surface.)

#### Plan preview / approval

Before running a multi-step task unattended, a session driving the agent tab **should** call `browser_request_plan_approval` with a short title and a numbered list of the steps it's about to take (2-20 steps — a text checklist, not a graph editor). It posts through the same [CLI Pending Actions](cli-pending-actions.md) approval queue every other AI-triggered change uses — the plan lands as a "browser plan…" row in your inbox — rather than a separate approval mechanism. The call itself never touches the browser and works regardless of whether the Browser panel is open or focused (it isn't subject to the attended/control/pause gates below — it gates _itself_). It waits up to 20 seconds for a quick yes/no, then returns `{ status: 'pending', approvalId }` if nobody has responded yet; the row stays in your inbox until you act on it (there's no auto-expiry), and the session (or any CLI caller with the token) can check the outcome later with `GET /cli-pending/:approvalId`. A plan is never auto-approved — no response means `'pending'`, forever, until you approve or reject it. This is the seam a future recurring/unattended job feature will require before it's allowed to run.

#### Staying in control

You decide how much the agent may do. The user-facing controls live in the Browser panel itself — the Pause button (which appears in the toolbar only when an agent is in play) and the per-site tiers in the **History → Permissions** tab — not in Settings:

- **Agent control** — an underlying `aiBrowserAgentControl` setting exists (`auto` = the agent may navigate/click/type freely, the default; `allowlist` = writes only on origins in `aiBrowserOriginAllowlist`; `off` = read-only), enforced main-side in `ai-browser-policy.ts`. **There is currently no Settings UI for this dial** — it stays at `auto` unless changed programmatically; the everyday controls are the per-site tiers, watch-mode, and Pause below.
- **Pause** — a Pause/Resume button that appears in the toolbar whenever an agent is meaningfully in play (a session bound in the side chat, the agent's own tab active, or recent agent activity) freezes all agent writes instantly so you can take over. It stays visible the whole time it's paused, so a paused agent is always resumable.
- **Per-site permission tiers.** From **History → Permissions**, set an exact permission for any site: `control` (the default — agent can read and act, subject to everything else on this list), `read-only` (agent can read the page but any write is denied), `no-access` (agent can't even read it), or `blocked` (same effect as no-access, meant for sites you never want the agent touching, like your bank). A site's explicit tier always wins over the category watch-mode default below it — set a payment site to `read-only` and it stays read-only even while you're watching. Adding a bare domain (no subdomain) covers all its subdomains too. (Don't confuse this with the tab right-click menu's Pin/Favorite/Today — that's tab housekeeping, not permissions; see [Spaces, tabs, and quick navigation](#spaces-tabs-and-quick-navigation).)
- **Attended-only writes.** Underneath the control setting, Omniscio also checks whether you're actually watching: "attended" means the Browser panel is open, **actually visible** (a kept-warm hidden copy doesn't count — see "The agent's own tab" above), AND its window is focused, confirmed by a live ~5-second heartbeat — not just any Omniscio window being open, and never a signal the agent itself can report. It also requires the session driving the tab to be an interactive one — a scheduled or automated (agent-driven) session is never treated as attended, even with the panel open and focused. A write attempted while unattended is refused regardless of the control setting, as a conservative default.
- **Supervised agent writes (opt-in, default OFF).** Because an *agent-driven* session (one an AMC agent is running, not one you're typing in) is never attended by default, it can read the browser but not act — the common wall when you want the AI to actually do something (e.g. like songs on Spotify). **History → Permissions** has a switch, **"Let the agent act while I'm watching,"** that relaxes exactly this: with it on, your real presence at the *visible* panel + a focused window vouches for an agent session too, so it may click and type on normal sites while you watch. It's a deliberate security relaxation — off by default, and it still requires genuine presence (a hidden/kept-warm panel, an unfocused window, or a call with no session identity is still refused). A second switch, **"Include banks, payments & social sites"** (also default OFF), is required before supervised writes reach the finance/payments/social category below — otherwise those stay locked to genuine interactive supervision.
- **Category watch-mode (finance / payments / social).** A curated, built-in list of well-known bank, brokerage, payment-processor, and social-network domains always requires you watching — even after a future setting relaxes the general attended-only rule for other sites. A write on one of these while unattended is refused with a distinct reason, and Omniscio drops a dismissible "The agent needs you watching to continue" card in your inbox (one card per site, not one per attempt) so you know it's waiting rather than silently failing. This is on top of, not instead of, the "no unattended payments, ever" invariant.

Reads (`snapshot`, `screenshot`, `console`, `network`, `get_page_content`, `tab.list`, `stop_dev_session`, `list_dev_servers`, `search_memory`) are normally allowed regardless of the control setting — with one exception: a site you've set to `no-access` or `blocked` denies reads too, not just writes. Actions that change a page — including `start_dev_session`, since it attaches to the tab's live CDP session — obey, in order, the per-site tier, the control setting, the attended check, the category watch-mode, and the pause switch. `request_plan_approval` is a separate case: it never touches a page, so none of the above applies to it — it always runs, and its own safety mechanism is the approval queue described in "Plan preview / approval" above (it gates itself; a plan is never auto-approved).

### Design mode & Visual Edit

Two point-at-the-page tools, both for **your own tabs** (not the agent's), both items in the **✨ AI tools** menu next to the address bar:

- **Visual edit** (in the ✨ AI tools menu) — click an element, type what you want changed in plain English ("make this button blue"), submit. With a session bound in the side chat, Omniscio sends it the element's selector, its rendered HTML/styles, and a cropped screenshot, with an instruction to make the change in the source code. With no session bound, it falls back to copying that same prompt to your clipboard so you can paste it in yourself.
- **Design mode** (in the ✨ AI tools menu) — click an element and a live style panel opens: color, background, font size/weight, padding, margin, width, height, and border radius, all editable, plus buttons to move the element up or down among its siblings (button-stepped, not drag-and-drop). Changes preview instantly on the real page so you can see what you're doing — but nothing here is saved to the page itself. **Apply** needs a session bound (there's no clipboard fallback for Design mode); it packages only what changed — the before/after value of each property you touched, plus how many positions you reordered it — and sends that delta to the session with an explicit instruction to map it onto the component's source and edit the source, never the running page. Design mode tells the agent to edit your code; it does not hack the live page.

Both hand off to whichever session is currently selected in the browser's side chat pane (below) — neither spawns a new one.

### Spaces, tabs, and quick navigation

- **Spaces.** A space is a named group of tabs — not a separate profile or login, just a way to keep "work" tabs and "research" tabs apart. Create, rename, or delete one from the space switcher (next to the tab strip, or inline at the top of the vertical tab rail); switching spaces swaps which tabs are visible.
- **Vertical tab strip.** A toggle in the **⋯** overflow menu swaps the horizontal tab strip for a vertical one down the side — same tabs, same behavior, just a different layout. Your choice is remembered.
- **Tab housekeeping (Pin / Favorite / Today).** Right-click a tab for **Pin** (never auto-sleeps, discards, or archives), **Favorite** (exempt from auto-archiving only), or leave it **Today** (the default — eligible for Omniscio's normal inactive-tab cleanup). This is unrelated to the per-site _permission_ tiers above despite the similar name — it's about memory/tab-count housekeeping, not what the agent is allowed to do. The active tab and the agent's own tab are never auto-cleaned.
- **Command palette (Ctrl+K).** Fuzzy-search across every open tab (in any space), your bookmarks, and your history, plus quick actions: new tab, new space, "ask about all open tabs" (hands every open tab's URL and title to the active chat session as context), and toggling the side chat.
- **Most-recently-used switching (Ctrl+Tab).** Hold Ctrl and tap Tab to cycle open tabs most-recently-used-first, the same feel as Alt+Tab for windows; release Ctrl to land on the one you want.

## Related

This page is the companion to [AI-Drivable Browser (RETIRED — see My Real Chrome)](ai-browser.md), which carries the rest of the retired feature. The surfaces that replaced it are [My Real Chrome](real-chrome-bridge.md) and [Browser Logins](browser-logins.md); the plain single-page pane it built on is [Embedded Browser](embedded-browser.md); and the approval queue its plan preview posts through is [CLI Pending Actions](cli-pending-actions.md).
