AI-Drivable Browser (in development — enable it in Settings → Lab)
The AI-drivable browser — in development, and live again behind its Settings → Lab toggle (`aiBrowserEnabled`): it turns the embedded Browser pane multi-tab and lets an Omniscio agent drive it in its own isolated tab, with an autofill vault, Memories, Journeys, Design mode, workflow recording and CLI routes. For driving your own signed-in Chrome, use My Real Chrome instead.
What it is
In development — and live again. This feature was retired on 2026-08-02 and un-retired on 2026-08-31; it is gated behind the
ai-browserLab toggle (aiBrowserEnabled, Settings → Lab, off by default) and works when you turn it on. For driving your own signed-in Chrome, the road is My Real Chrome (a companion extension drives your real browser, tabs it opens itself) + Browser Logins — that is a different feature, not this one. This page describes the in-app AI-drivable browser.
This ships OFF. The AI-drivable browser is gated behind the
ai-browserLab flag (Settings → Lab, off by default). While it's off, the Browser pane is exactly the plain Embedded Browser — one page, no tabs, no agent access. Nothing on this page happens until you turn the flag on.
The AI-drivable browser turns Omniscio's embedded Browser pane into a real, multi-tab browser that an Omniscio agent can drive for you — open a page, read it, click, type, screenshot — all inside the pane, while you watch. It's Omniscio's answer to the "agent that can use a browser" idea, built so the agent works in its own isolated tab and can never touch your logged-in cookies.
Where to find it
Turning it on
- Settings → Lab → enable
ai-browser. (Or launch withAMC_SHOW_AI_BROWSER=1to reveal it.) - Make sure the Browser panel is available: Settings → Features → Enable Browser.
- Open Browser in the Omniscio sidebar. You'll now see a tab strip instead of a single page.
How it behaves
Autofill vault (saved logins for your own tabs)
The Passwords tab in History is a small password manager for the browser pane's own tabs — not the agent's. Log into a site in one of your own tabs and Omniscio offers to save it: you type the username/password into a save dialog yourself (it doesn't capture what you typed on the page). Come back to that exact site later and a toast offers to fill it in for you. The list shows the site, username, and last-used date; remove a saved login any time. The password itself is encrypted through your OS's own credential store (Keychain / Credential Manager / libsecret) — the management panel only ever shows metadata, never the password.
This is a different feature from Browser Logins, and the two don't share data:
| Autofill vault | Browser Logins | |
|---|---|---|
| Stores | An encrypted password, in your OS keyring | A cookie/session — never a password |
| Works in | Your own tabs, in the AI-drivable Browser panel | Any spawned agent (via the CLI) — and now the in-app agent tab / your own in-app tab (cookies-only, best-effort) |
| Reached from | Browser → History → Passwords | Settings → Browser Logins |
| Agent access | None — no tool, no CLI route reaches it | The whole point — amc-browser use <name>, or seed the in-app agent tab |
The agent's own isolated tab never sees your saved passwords, by design — there's no tool and no CLI route that reaches this vault, the same isolation that keeps your cookies out of the agent's tab keeps your passwords out of it too. A saved credential only fills on the exact site (scheme + host + port) it was saved on — it never carries over to a different subdomain.
Memories (ask a session about pages you've already seen)
Omniscio quietly indexes the pages you visit in your own tabs — never the agent's — chunking and embedding the text with the same local embedding model Omniscio uses elsewhere (nothing leaves your machine; a duplicate visit to the same content isn't re-indexed). A session can then search across everything you've browsed instead of you re-finding and re-opening it, with browser_search_memory — give it a query, get back the best-matching page chunks with URL, title, matched text, and when you visited, ranked by relevance. "What was that pricing page I looked at last week" becomes one tool call instead of digging through history.
Manage it from History → Memories: browse what's indexed, delete individual entries, clear everything, or turn indexing off entirely (it's on by default).
Journeys (your browsing, auto-grouped by topic)
History → Journeys groups your recent browsing into topic clusters automatically — nothing to save by hand. It's computed from the same page embeddings Memories uses, bucketing pages that are topically close together and labeling each group after its most-recently-visited page. It's a lightweight, on-the-fly grouping (computed when you open the tab, not a background job or a precise research-session tracker) — good for "what was I looking at when I was researching X."
Cursor chat & Compose (rewrite or draft text in any field)
The Write in field item in the ✨ AI tools menu captures whatever text field is focused in your own tab. Select text first and it opens in rewrite mode (tell it how to change the selection); focus an empty or unselected field and it opens in compose mode (tell it what to draft). Type your instruction in the prompt bar and submit — like Design mode, this needs a session bound in the side chat.
The bound session gets the field's context plus your instruction, then writes the result back into the page via browser_write_inline_text. The write-back is authorized by a single-use token, scoped to that exact tab and expiring in about two minutes, so nothing else can reach into your page and write to it. On plain <input>/<textarea> fields, the write-back is inserted the same way real typing is, so Ctrl+Z undoes it like anything you typed by hand. Rich text editors (Gmail, Notion, and similar contenteditable surfaces) are best-effort — the exact selection isn't always restored, so treat undo there as a bonus, not a guarantee. Password fields are never captured.
Workflow recording (turn what just happened into a skill)
Two ways to record a workflow into a skill — same result, different actor:
- You record. The Record workflow item in the ✨ AI tools menu captures your clicks, typed values, form submissions, and page navigations in your own tab while it's on. While recording, a Stop button surfaces in the toolbar itself (never buried in a menu) so you can always end it. Hit Stop and a review panel lists every typed value it saw; you decide which ones become
{{variables}}in the resulting skill — nothing is tagged automatically. A value that looks like a password shows up redacted and prompts you to tag it if you want it replaced with a variable rather than baked in literally. - The agent records. A session driving its own tab can call
browser_start_recording/browser_stop_recordingto capture its own navigate/click/type/submit sequence the same way, ending with the same review-and-tag step before saving.
Either path saves a draft skill — the same shape the Skills gallery (below) already understands, so it shows up there ready to run or edit. The honest ceiling: it's a literal, numbered replay of what happened, generalized only by the variables tagged at save time — it doesn't infer a broader workflow from a single recording.
Side chat pane
A collapsible panel on the right side of the Browser view holds a live chat with whichever session you pick — the same chat surface as the main Dashboard, just docked next to the browser so you're not switching views to talk to the agent that's driving (or co-piloting) the tab. This is the session Visual Edit, Design mode, and Cursor chat all hand off to. Toggle it from the command palette or its own collapse control.
While an agent drives the browser, a small activity ticker under the pane header narrates each step in plain English ("Opened example.com", "Clicked e12", "Couldn't click — element not found") — the last five actions, failures tinted red. The text an agent types or fills is never shown in these lines (it may be your passwords or personal data); only the target field and the outcome appear. Below the chat, quiet suggestion chips appear when the composer is empty — rule-based shortcuts for the current page ("Get the transcript" on a YouTube video, "Review this PR" on GitHub, "Summarize this page" anywhere). Clicking a chip only pre-fills the composer; nothing is ever sent until you press send, so a stray click never costs a paid call.
Skills (save a workflow as a one-click command)
A skill is a browser workflow you save in plain English and re-run in one click — "Summarize this page", "Find the cheapest flight and add it to my cart", "Pull the action items off this doc". It's the Omniscio take on Dia's Skills / Comet's Shortcuts.
Open it from the History toolbar button → Skills tab in the browser. There you can create a skill (a name, the plain-English steps, and an optional emoji for its card), edit or delete existing ones, and run any of them. Typing one out isn't the only way in — see Workflow recording above for recording one instead.
Under the hood a skill is a recipe — there's no separate "skill engine". Saving a skill writes a browser-scoped recipe (a normal Omniscio recipe flagged as a browser skill), so it shows up in the Skills gallery. Running one goes through the exact same gated run path as any recipe: it opens the standard Run dialog (pick a project, fill any inputs, confirm) and starts a recipe run — it never spawns anything directly. That means every safety rule above still applies while a skill runs: an acting skill routes its plan through request_plan_approval (the inbox approval described in "Plan preview / approval"), and each action it takes still obeys attended-only writes, category watch-mode, per-site tiers, and the Pause switch. A skill is a shortcut to a workflow, not a way around the guardrails.
v1 is your own skills only — there's no community gallery or marketplace yet.
Video pages (YouTube Q&A)
A "what does this video say" / "summarize this YouTube video" request from a session driving the agent tab: browser_navigate to the video URL, then browser_get_video_transcript — one call, no clicking through the page. It reads YouTube's own caption-track list off the page and fetches the chosen track's caption data directly, so it works even when the "Show transcript" panel is never opened. Only works on a YouTube watch page (youtube.com/watch?v=…) with at least one caption track, and only while ai-browser is enabled (Settings → Lab) — anything else (a non-YouTube site, a channel/Shorts page, a video with no captions) is a friendly error, not a best-effort scrape.
Two other paths do not cover this, for context:
- The
summarizeskill does not cover this. For a URL it does a plainWebFetchof the page HTML looking for "the main article text" — on a YouTube watch page that's just the title/description/JS shell; the transcript loads client-side and is never in the initial HTMLWebFetchsees. POST /download(yt-dlp) does not fetch captions either. ItsDownloadModeisvideo | audio | originalonly (src/main/services/conversion/download-types.ts) — it saves a media file, never subtitle/transcript text, and nothing pipes a downloaded video through the app's whisper transcriber (that pipeline only transcribes local screen recordings — a separate feature).
If browser_get_video_transcript ever fails on a page that visibly has a transcript (a YouTube shape change), the old manual workaround still applies as a fallback: browser_snapshot (or browser_get_page_content) to locate the "Show transcript" control under the description (may need expanding "...more" first), browser_click its ref, then browser_get_dom with a CSS selector on the transcript panel (e.g. ytd-transcript-segment-list-renderer) to read the rendered timestamp+text segments — or browser_network to spot the transcript/timedtext request the click just fired, then fetch that URL directly.
PDF / docs pages
A "what does this PDF say" / "ask about this PDF" request from a session driving the agent tab: browser_get_page_content (and browser_read_tabs) detect a PDF — by response content-type or a .pdf URL — and refuse with a PdfPageError instead of returning garbage, because a PDF has no DOM for Readability to read. The error names the fix: browser_save_pdf saves the active tab into Downloads as a real .pdf file (it calls CDP's Page.printToPDF directly — when the tab is already showing a PDF via Chromium's built-in viewer, that streams back the original PDF bytes, not a re-rendered copy, so selectable text survives), then the session runs the pdf-extract skill's Read tool on that saved path. Two calls, no new tool needed — save_pdf already existed for the Band-1 export pack; PDF pages are just its second use.
Two other paths do not cover this, for context:
POST /download(yt-dlp) does not fetch PDFs. Same as the video-transcript case above: itsDownloadModeisvideo | audio | originalonly, aimed at yt-dlp's media-site extractors — pointed at a plain PDF URL it has no matching extractor and fails. It's not a generic "download this URL to disk" endpoint.- Google Docs / Office docs opened as a live web page are not handled by
get_page_contenteither —docs.google.com/document/…is a JS app, not a PDF, so it won't hitPdfPageError, but its editor view often has little or no extractable DOM text (canvas-rendered), so Readability's plain-text fallback can come back thin or empty. Omniscio has a dedicated path for this instead:POST /gdoc/import(src/main/services/cli/cli-server-gdoc-routes.ts) pulls a Google Doc's real content as Markdown via the Drive API — use that for a Google Doc, not the browser tools.
Automations & co-work
Three common jobs are shipped as recipe templates in the Pattern Library (Recipes → New from pattern), not as new machinery. Each is a normal recipe built over the browser tools above plus the session's own tools; there is no new scheduler, no new engine, and no new file route. Pick one, fill in the blanks, and it becomes a runnable recipe.
Recurring browser jobs
"Every morning, open these dashboards, pull the numbers, and write me a summary." That is the Recurring browser report pattern (recurring-browser-report): a one-step recipe whose prompt tells the agent to browser_navigate + browser_get_page_content over a list of URLs, extract what you asked for, and save a dated markdown summary via its own Write tool. It is read-only by design (no clicking or submitting), which is what makes it safe to run unattended.
To make it recurring, wire the instantiated recipe to a cron job on the existing cron system — nothing new is built for scheduling:
- Instantiate the pattern into a recipe (it lands in your recipe list with an id).
- Create a recurring cron job of
type: "recipe"pointing at it. Over the CLI control server that is aPOST /cron/jobswithrunMode: "recurring", acronExpression(e.g.0 8 * * *for 8am daily), andjobConfig: { "recipeConfigId": "<the recipe id>" }. The cron recipe executor (cron-recipe-executor.ts) already dispatches it on schedule. - A cron job created via the API always registers as pending approval — you approve it once in the inbox before it starts firing.
Because a scheduled run is unattended, the Band-2 safety layer applies with no extra work: the moment the agent tries to act (not just read) on a finance / payments / social site while nobody is watching, the permission gate denies it with watch-mode-required (ai-browser-permission-evaluate.ts). So a recurring job is fine for reading dashboards, prices, or status pages, and cannot quietly buy or post anything. For a multi-step unattended run, have the recipe call browser_request_plan_approval up front so its plan is queued for you to approve.
Research to report (co-work)
The Browser research to report pattern (browser-research-to-report) is the co-work shape: one session researches a topic with browser_navigate + browser_get_page_content + browser_read_tabs, cross-checks across sources, then writes a cited markdown report with its own Write tool to its working directory. No file route or share pipeline is involved — the report is just a file the session wrote where you can open it. Run it as a normal session (you watch and steer) or on a schedule like above.
Guided shopping / booking (confirm-gated)
The Guided shopping pattern (guided-shopping) is a three-step, confirm-gated recipe for "compare these, then buy the best one":
- Compare — the agent reads listings/prices across the sites you named and produces a shortlist plus a proposed purchase plan. It changes nothing, and it calls
browser_request_plan_approvalso the plan is queued for you. - Approval gate — the recipe pauses at a human approval step. Nothing proceeds until you approve.
- Checkout — only after approval, the agent adds the item to the cart and drives to the checkout page, then stops at the payment step. It never enters card/address details and never places the order; you review and pay by hand.
This is the load-bearing invariant, enforced at two layers: the recipe's own approval gate sits before the checkout step, and underneath it the Band-2 permission gate never lets any agent complete a payment unattended (no unattended payments, ever). Booking flows follow the same shape — compare, approve, drive up to the confirm/pay step, hand off.
Privacy + safety notes
- The agent tab is isolated — it uses its own profile and never reads your
persist:browsercookies. - History and bookmarks are yours — the agent's navigation is never recorded to your history, and the agent can't clear it.
- Memories only indexes your own tabs — the agent's tab is never included in what gets indexed or what
search_memorycan find. - The autofill vault is invisible to the agent — no tool or CLI route reaches it; it's exclusively for your own tabs.
- The agent can only navigate to
http:/https:pages —navigateandtab.newrejectfile:,data:,about:, and similar schemes, so the agent can't be tricked into reading a local file or acting from a null-origin page. - Everything is local — no browsing data leaves your machine, and the
/ai-browser/toolroute is bound to127.0.0.1behind the CLI bearer token. - The embedded page is treated as untrusted and runs sandboxed, exactly like the plain Embedded Browser — the AI-drivable mode doesn't weaken that.
What's still coming
Design mode's element reordering is button-stepped (move up/down among siblings), not drag-and-drop. Cursor chat's rewrite/compose is best-effort inside rich text editors (Gmail, Notion, and similar contenteditable surfaces) — plain <input>/<textarea> fields are the fully-supported case. Workflow recording generalizes only the values explicitly tagged as variables at save time; it doesn't infer a broader workflow from one recording. Everything else on this page — multi-tab browsing, agent driving in an isolated tab, the full permission/safety layer, Design mode, Cursor chat, and workflow recording — is complete behind the flag.
For agents
Where it lives in the code
- UI:
src/renderer/src/features/browser/BrowserView.tsx(tab strip, toolbar, bars) +TabWebview(per-tab web view), withBookmarksManagerModal.tsx,HistoryModal.tsx(the six-tab History/Memories/Passwords/Journeys/Permissions/Skills modal),DownloadShelf.tsx,visual-edit.ts(the injected page inspector), andDesignModeStylePanel.tsx(the style panel). - Tool registry:
src/main/services/ai-browser/tools/index.ts— the single source of truth all 40 tools, the MCP server, and the/ai-browser/toolHTTP route derive from. - Agent driver:
src/main/services/ai-browser/ai-browser-driver.ts(the tool vocabulary over CDP), gated byai-browser-policy.ts(the auto/allowlist/off + pause gate) andai-browser-permission-evaluate.ts(per-site tiers + category watch-mode + attended check). - Isolated partition:
AI_BROWSER_AGENT_PARTITIONinsrc/shared/ai-browser.ts. - Autofill vault:
src/main/services/ai-browser/autofill-vault-service.ts(human tabs only — no tool or CLI route reaches it). - Memories / Journeys:
src/main/services/ai-browser/browser-index-service.ts(indexing +search_memory) +journeys-service.ts(topic clustering). - Spaces:
src/main/services/ai-browser/ai-browser-spaces-service.ts. - Cursor chat / inline write-back:
src/main/services/ai-browser/tools/write-inline-text.ts+inline-write-token.ts. - Workflow recording:
src/main/services/ai-browser/ai-browser-recording.ts+tools/start-recording.ts/stop-recording.ts(agent-tab path) andsrc/renderer/src/features/browser/use-browser-recording.ts(human-tab path). - CLI routes:
src/main/services/cli/cli-server-ai-browser-routes.ts(/ai-browser/tool) +cli-server-ai-browser-bookmarks-routes.ts. - Flag: the
ai-browserentry insrc/shared/unreleased-features.ts(status: 'in-development'). - Invariants + safe-change checklist:
.claude/memory/contracts/ai-browser-contract.md.
Related
The other road for driving a browser you are signed into yourself is My Real Chrome, which drives your own signed-in Chrome through a companion extension, and Browser Logins for reusing a captured session. The plain single-page pane it was built on is Embedded Browser, the approval queue its plan preview posts through is CLI Pending Actions, and the local HTTP surface its raw tool route lives on is cli-control.md.
Material that outgrew this page — the two point-at-the-page tools (Design mode and Visual Edit), spaces, tabs and quick navigation, and the agent's own isolated tab with its MCP tool vocabulary and permission tiers — continues in AI Browser (part 2).