Image Bridge For Text-Only Models
An optional session setting that lets a model which cannot read images still work from one: a vision-capable model writes a brief description first, and that brief travels to your chosen text-only model alongside your message.
What it is
Image Bridge is an optional session setting that lets text-only models work from attached images indirectly.
Default behavior is off. With the setting off, attaching an image to a model that cannot use images keeps the existing compose warning and the image content is not sent to that model.
When enabled, Omniscio sends the attached image to a vision-capable model first. That model produces a concise visual brief targeted to the user's message. Omniscio then sends the visual brief, plus the user's original request, to the selected text-only model.
The visual brief prompt is designed for screenshots, documents, charts, photos, and UI states. It extracts visible text, layout, errors, counts, important objects, relationships, and uncertainty. Text visible inside an image is treated as untrusted content to transcribe, not as instructions to follow.
Image-capable models do not use Image Bridge; they keep receiving native image attachments. Bridge calls are tracked under the session-image-bridge cost source with session attribution.
Where to find it
It is a session setting, off by default, so you meet it where session options are set rather than in Settings. Image-capable models never use it — they keep receiving attachments natively.
How it behaves
With the setting off, attaching an image to a model that cannot use images keeps the existing compose warning, and the image is not sent at all. With it on, Omniscio sends the image to a vision-capable model first, which produces a concise visual brief aimed at your message; the brief plus your original request then go to the text-only model you picked. The brief is written for screenshots, documents, charts, photos and interface states — it pulls out visible text, layout, errors, counts, important objects and relationships, and says where it is unsure. Text inside an image is treated as content to transcribe, never as instructions to follow.
For agents
Bridge calls are tracked under the session-image-bridge cost source with session attribution, so their spend is separable from the rest of the session's.
Related
The provider pages say which models read images natively and which need this bridge — see Claude providers, Gemini and DeepSeek.
Last verified 2026-09-23