Record for Agent
A mode on top of the recorder: record your screen while you talk, mark the moments that matter by hotkey, typed note, voice or a drawn box, then hit Send to agent. The agent reviews what you showed it, shows cropped screenshots to confirm it understood, and you talk it into a task list. It never files anything — the conversation is the whole output.
What it is
Record your screen while you talk. Click Send to agent. The agent watches what you showed it, shows you cropped screenshots to confirm it understood, and you talk it into a task list.
It is a mode on top of the existing Screen Recorder — not a second recorder and not a separate screen.
Where to find it
Turning it off
Settings → Screen Capture → Extras → Record for Agent. Off hides the marking tools and the Send button; ordinary screen recording is unaffected.
The video pass has its own switch directly beneath it — Let Gemini watch the video — so you can keep Record for Agent on while leaving the upload off.
How it behaves
The four steps
1. Record. Open Screen Recording, pick what to capture, and on Step 2 · Recording options flip Record for agent on before you press Start. The recording bar then shows a For agent indicator while it is on, and the marking tools appear on it.
The toggle is per-recording and always starts off — a for-agent take turns on pointer tracking, so it is a fresh choice each time rather than a setting that quietly follows you into the next recording. The quick-record hotkey deliberately skips it: quick-record is for grabbing something fast, and a marking take is a deliberate act. (If Record for Agent is switched off in Extras, the toggle is not shown at all.)
2. Mark the moments that matter. Four ways, and you can mix them freely:
| Way | How | Good for |
|---|---|---|
| Hotkey | A global shortcut (default Ctrl/Cmd+Shift+6), rebindable in Settings → Keyboard Shortcuts |
The fastest one. Works while you are in another app. |
| Typed note | The note box on the recording bar | When the exact words matter |
| Say it | Just say "flag this", "mark this", "note this", "this is the bug" | Hands-free, and the rest of the sentence becomes the label |
| Draw a box | The draw button, then drag on screen and optionally label it | Pointing at one specific thing on a busy screen |
3. Review before you send. Open the recording, hit Send to agent, and you get the list of every moment it found — including the spoken ones — numbered exactly as the agent will refer to them. Untick anything you do not want. Pick a project. Send.
4. Talk. One conversation covers everything in the recording. The agent sorts bugs from teaching by itself and shows you a cropped screenshot every time it refers to a moment.
The two rules that make it trustworthy
It never files anything. No tasks, no tickets, no notes. The conversation is the whole output — you decide what happens next. An agent that helpfully creates a task has misunderstood the feature.
That is the rule for the Send to agent button, which always opens a fresh conversation to think with. There is one other way in, and it is deliberately the opposite: an agent holding the command-line key can hand a recording to a session that is already running, and that one arrives as work — turn the moments into things to fix, and fix them. Same recording, same screenshots, same narration; only the closing instruction differs. It is not something the button can do and not something that happens by accident — you have to ask for it. See Sending a recording to an agent yourself.
It always shows you the picture. Every moment ships to the agent with its screenshot already cropped and highlighted, plus a ready-to-paste line — so showing you the picture is easier for it than not. Each reply then carries a badge: a green Grounded when it included a screenshot, an amber No screenshot in this reply when it did not. That is how you catch a misunderstanding on turn 2 instead of turn 20.
The agent can also open those screenshots, not just show them to you. The link it pastes is what draws the picture on your screen; the opening message additionally names the folder the image files sit in, so the agent can look at a frame itself before it says what is on it. Before that folder was named, an agent reviewing a recording had the links but no pixels — it could repeat your narration back to you, but it could not spot anything you had not already said out loud.
Honest limit: no design makes this impossible to skip — the agent's text is the deliverable. What it does is make compliance nearly free and failure impossible to miss.
Timing — why the screenshots line up
The video's first frame is not the instant you pressed record; on Windows the capture can take a second or two to warm up. Every marked moment is converted through the recorder's single media clock, so a moment's timestamp seeks to the frame you actually meant rather than the aftermath.
A moment you marked during that warm-up is flagged, and the agent is told its frame is not reliable rather than confidently describing the wrong thing.
The typed note is stamped when you open the box, not when you press Enter — you were looking at the thing then, and typing about it takes seconds.
Drawing a box, and editing it later
A box you drew live is saved with the moment, and it also appears in the video editor as a normal annotation you can drag, resize and retime. Whatever you leave it as is what the agent gets cropped — the editor is the later, more deliberate act, so it wins. Delete the annotation and the box you originally drew is used instead; the moment is never lost.
Letting a model watch the video (optional, off by default)
Everything above hands the agent your narration plus a handful of still screenshots. That cannot show motion — and motion is what recordings keep being about. Reviewing a real 19-minute take, an agent said so twice without being asked: "stills can't show motion", and of the tester's most-repeated bug, "a list that keeps re-refreshing — I cannot see this in a still frame."
Turn this on and Gemini watches the actual video first, then hands its reading to the agent that reviews the recording with you. Claude still runs the conversation.
- It goes in its own labelled section, above your narration, and says plainly that it is a machine's reading rather than your words — so the agent can weigh the two and tell you when they disagree.
- It is asked for what stills cannot carry: what moves, changes, flickers, reloads, hangs or repeats, and roughly when — not a second summary of what you already said.
- It flags things nobody mentioned. On a test clip it identified a one-per-second flicker and called it a probable render bug on its own.
- This is not the same as auto-detecting moments. It writes a description; it never adds marks to your timeline. The moments are still only the ones you made.
- If it cannot run, the send happens anyway and the message says why in plain words. No key, a refused upload, an error, a timeout — none of them can stop you sending a recording.
- Its answer is treated as untrusted text, exactly like text read off your screen, so it cannot smuggle instructions into the conversation.
The trade you are making: with this off, a review sends Anthropic a few cropped screenshots and a transcript made on this computer. With it on, the whole recording is uploaded to Google — everything visible for the entire take, not just the moments you marked. That is a materially different exposure, which is why it is off until you say otherwise.
It uses the same Gemini API key as everything else here; with no key set, it simply does not run.
Privacy
- Pointer tracking turns on for that recording only. Starting a Record-for-Agent take is the consent — the indicator shows while it is on, and the next ordinary recording has it off again.
- On Windows, the text of the field you are typing in is captured when you use the hotkey. It catches pasted text too. It runs only on the hotkey, because the note box and the drawing overlay are Omniscio's own windows — a check there would read your note back, not your app.
- Password fields are never captured. The refusal happens inside the Windows helper itself, so the characters never reach Omniscio at all. Anything the check cannot positively clear as "not a password" also contributes nothing.
- Nothing leaves this computer unless you send the recording. And with the optional video pass ON, sending uploads the WHOLE recording to Google — every second of it, not just the moments you marked. Off (the default) it never runs and nothing is uploaded.
- There is no keylogger. Omniscio's key hook deliberately throws away which key you pressed, and this feature does not change that. It reads a field's current contents at one instant you chose, never a stream of keystrokes.
What it deliberately does not do
- It does not guess at interesting moments for you. An automatic detector was built and measured, and it missed a small error badge appearing 100 times out of 100 — exactly the thing you would be recording to show us. It will not ship without measurements on real recordings.
- It does not split one recording into several conversations. A recording often contains several unrelated problems; that is expected, and they stay together.
- It does not send anything but the recording. Video, audio, cursor, and the notes you wrote. No application logs. The video file itself only ever leaves this computer if you turn the optional Gemini video pass on.
Sending a recording to an agent yourself
The Send to agent button is the normal way in. There is also a command-line door, which an agent can use on your behalf — and it can do one thing the button cannot.
POST /capture/<recordingId>/send-to-agent
{ "projectId": "…" } → opens a NEW review conversation, exactly like the button
{ "sessionId": "…" } → hands it to a session that is ALREADY RUNNING
Why the second one exists. A session that is already working on something has the context the recording is about. Opening a fresh conversation to tell it what you just showed it throws that away. So a recording delivered this way arrives as work — turn each moment into something concrete, then go fix it — rather than as a conversation to think in.
It waits its turn. A recording handed to a busy session lands at that session's next natural stopping point rather than interrupting it, and it survives the app restarting. If the session could never read it — archived, or wedged mid-turn — you are told the send did not happen rather than being told it worked while the message quietly goes nowhere.
What it costs. The projectId form starts a real agent session, so it costs money the
same way opening any conversation does. The sessionId form starts nothing — it is a
message into a session you are already paying for.
Why an agent cannot do this quietly. The route needs the full-trust command-line key, not the limited one handed to each session. A recording is not owned by any one conversation, so a lesser key that leaked would otherwise reach every recording you have ever made — the same reason an agent needs that key to look at any frame of one.
Related
This is a mode on top of the screen-recorder.md, and that page explains the recording, marking and editing basics this one builds on, so read it first if you have never used the recorder. Everything else here — the review conversation, the screenshot badges, the optional video pass — is specific to sending a recording to an agent.
Last verified 2026-09-23