---
title: Record for Agent
---

# Record for Agent

## What it is

Record your screen while you talk. Click **Send to agent**. The agent watches what you
showed it, shows you cropped screenshots to confirm it understood, and you talk it into a
task list.

It is a mode on top of the existing [Screen Recorder](screen-recorder.md) — not a second
recorder and not a separate screen.

## Where to find it

### Turning it off

Settings → Screen Capture → Extras → **Record for Agent**. Off hides the marking tools and
the Send button; ordinary screen recording is unaffected.

The video pass has its own switch directly beneath it — **Let Gemini watch the video** —
so you can keep Record for Agent on while leaving the upload off.

## How it behaves

### The four steps

**1. Record.** Open **Screen Recording**, pick what to capture, and on **Step 2 ·
Recording options** flip **Record for agent** on before you press Start. The recording
bar then shows a **For agent** indicator while it is on, and the marking tools appear on
it.

The toggle is per-recording and always starts off — a for-agent take turns on pointer
tracking, so it is a fresh choice each time rather than a setting that quietly follows you
into the next recording. The quick-record hotkey deliberately skips it: quick-record is
for grabbing something fast, and a marking take is a deliberate act. (If Record for Agent
is switched off in Extras, the toggle is not shown at all.)

**2. Mark the moments that matter.** Four ways, and you can mix them freely:

| Way | How | Good for |
| --- | --- | --- |
| **Hotkey** | A global shortcut (default `Ctrl/Cmd+Shift+6`), rebindable in Settings → Keyboard Shortcuts | The fastest one. Works while you are in another app. |
| **Typed note** | The note box on the recording bar | When the exact words matter |
| **Say it** | Just say "flag this", "mark this", "note this", "this is the bug" | Hands-free, and the rest of the sentence becomes the label |
| **Draw a box** | The draw button, then drag on screen and optionally label it | Pointing at one specific thing on a busy screen |

**3. Review before you send.** Open the recording, hit **Send to agent**, and you get the
list of every moment it found — including the spoken ones — numbered exactly as the agent
will refer to them. Untick anything you do not want. Pick a project. Send.

**4. Talk.** One conversation covers everything in the recording. The agent sorts bugs
from teaching by itself and shows you a cropped screenshot every time it refers to a
moment.

### The two rules that make it trustworthy

**It never files anything.** No tasks, no tickets, no notes. The conversation is the whole
output — you decide what happens next. An agent that helpfully creates a task has
misunderstood the feature.

That is the rule for the **Send to agent** button, which always opens a fresh conversation
to think with. There is one other way in, and it is deliberately the opposite: an agent
holding the command-line key can hand a recording to a session that is **already running**,
and that one arrives as work — turn the moments into things to fix, and fix them. Same
recording, same screenshots, same narration; only the closing instruction differs. It is
not something the button can do and not something that happens by accident — you have to
ask for it. See [Sending a recording to an agent yourself](#sending-a-recording-to-an-agent-yourself).

**It always shows you the picture.** Every moment ships to the agent with its screenshot
already cropped and highlighted, plus a ready-to-paste line — so showing you the picture is
easier for it than not. Each reply then carries a badge: a green **Grounded** when it
included a screenshot, an amber **No screenshot in this reply** when it did not. That is
how you catch a misunderstanding on turn 2 instead of turn 20.

The agent can also **open** those screenshots, not just show them to you. The link it pastes is what draws the picture on your screen; the opening message additionally names the folder the image files sit in, so the agent can look at a frame itself before it says what is on it. Before that folder was named, an agent reviewing a recording had the links but no pixels — it could repeat your narration back to you, but it could not spot anything you had not already said out loud.

Honest limit: no design makes this impossible to skip — the agent's text is the
deliverable. What it does is make compliance nearly free and failure impossible to miss.

### Timing — why the screenshots line up

The video's first frame is not the instant you pressed record; on Windows the capture can
take a second or two to warm up. Every marked moment is converted through the recorder's
single media clock, so a moment's timestamp seeks to the frame you actually meant rather
than the aftermath.

A moment you marked during that warm-up is flagged, and the agent is told its frame is not
reliable rather than confidently describing the wrong thing.

The typed note is stamped when you **open** the box, not when you press Enter — you were
looking at the thing then, and typing about it takes seconds.

### Drawing a box, and editing it later

A box you drew live is saved with the moment, and it also appears in the video editor as a
normal annotation you can drag, resize and retime. Whatever you leave it as is what the
agent gets cropped — the editor is the later, more deliberate act, so it wins. Delete the
annotation and the box you originally drew is used instead; the moment is never lost.

### Letting a model watch the video (optional, off by default)

Everything above hands the agent your narration plus a handful of still screenshots. That
cannot show **motion** — and motion is what recordings keep being about. Reviewing a real
19-minute take, an agent said so twice without being asked: *"stills can't show motion"*,
and of the tester's most-repeated bug, *"a list that keeps re-refreshing — I cannot see
this in a still frame."*

Turn this on and Gemini watches the actual video first, then hands its reading to the agent
that reviews the recording with you. Claude still runs the conversation.

- **It goes in its own labelled section, above your narration**, and says plainly that it
  is a machine's reading rather than your words — so the agent can weigh the two and tell
  you when they disagree.
- **It is asked for what stills cannot carry**: what moves, changes, flickers, reloads,
  hangs or repeats, and roughly when — not a second summary of what you already said.
- **It flags things nobody mentioned.** On a test clip it identified a one-per-second
  flicker and called it a probable render bug on its own.
- **This is not the same as auto-detecting moments.** It writes a description; it never
  adds marks to your timeline. The moments are still only the ones you made.
- **If it cannot run, the send happens anyway** and the message says why in plain words.
  No key, a refused upload, an error, a timeout — none of them can stop you sending a
  recording.
- **Its answer is treated as untrusted text**, exactly like text read off your screen, so
  it cannot smuggle instructions into the conversation.

**The trade you are making:** with this off, a review sends Anthropic a few cropped
screenshots and a transcript made on this computer. With it on, the **whole recording is
uploaded to Google** — everything visible for the entire take, not just the moments you
marked. That is a materially different exposure, which is why it is off until you say
otherwise.

It uses the same Gemini API key as everything else here; with no key set, it simply does
not run.

### Privacy

- **Pointer tracking turns on for that recording only.** Starting a Record-for-Agent take
  is the consent — the indicator shows while it is on, and the next ordinary recording has
  it off again.
- **On Windows, the text of the field you are typing in is captured when you use the
  hotkey.** It catches pasted text too. It runs only on the hotkey, because the note box
  and the drawing overlay are Omniscio's own windows — a check there would read your note
  back, not your app.
- **Password fields are never captured.** The refusal happens inside the Windows helper
  itself, so the characters never reach Omniscio at all. Anything the check cannot
  positively clear as "not a password" also contributes nothing.
- **Nothing leaves this computer unless you send the recording.** And with the optional
  video pass ON, sending uploads the WHOLE recording to Google — every second of it, not
  just the moments you marked. Off (the default) it never runs and nothing is uploaded.
- **There is no keylogger.** Omniscio's key hook deliberately throws away which key you
  pressed, and this feature does not change that. It reads a field's current contents at
  one instant you chose, never a stream of keystrokes.

### What it deliberately does not do

- **It does not guess at interesting moments for you.** An automatic detector was built and
  measured, and it missed a small error badge appearing 100 times out of 100 — exactly the
  thing you would be recording to show us. It will not ship without measurements on real
  recordings.
- **It does not split one recording into several conversations.** A recording often
  contains several unrelated problems; that is expected, and they stay together.
- **It does not send anything but the recording.** Video, audio, cursor, and the notes you
  wrote. No application logs. The video file itself only ever leaves this computer if you
  turn the optional Gemini video pass on.

### Sending a recording to an agent yourself

The **Send to agent** button is the normal way in. There is also a command-line door, which
an agent can use on your behalf — and it can do one thing the button cannot.

```
POST /capture/<recordingId>/send-to-agent
{ "projectId": "…" }   → opens a NEW review conversation, exactly like the button
{ "sessionId": "…" }   → hands it to a session that is ALREADY RUNNING
```

**Why the second one exists.** A session that is already working on something has the
context the recording is about. Opening a fresh conversation to tell it what you just
showed it throws that away. So a recording delivered this way arrives as work — turn each
moment into something concrete, then go fix it — rather than as a conversation to think in.

**It waits its turn.** A recording handed to a busy session lands at that session's next
natural stopping point rather than interrupting it, and it survives the app restarting. If
the session could never read it — archived, or wedged mid-turn — you are told the send did
not happen rather than being told it worked while the message quietly goes nowhere.

**What it costs.** The `projectId` form starts a real agent session, so it costs money the
same way opening any conversation does. The `sessionId` form starts nothing — it is a
message into a session you are already paying for.

**Why an agent cannot do this quietly.** The route needs the full-trust command-line key,
not the limited one handed to each session. A recording is not owned by any one
conversation, so a lesser key that leaked would otherwise reach every recording you have
ever made — the same reason an agent needs that key to look at any frame of one.

## Related

This is a mode on top of the [screen-recorder.md](screen-recorder.md), and that page explains the recording, marking and editing basics this one builds on, so read it first if you have never used the recorder. Everything else here — the review conversation, the screenshot badges, the optional video pass — is specific to sending a recording to an agent.
