Omniscio documentation
Browse all documentation
  1. Getting Started13
  2. Sessions & Agents115
  3. Inbox & Notifications59
  4. Projects & Tasks95
  5. Automation & Scheduling75
  6. Knowledge & Memory26
  7. AI Features60
  8. Integrations100
  9. Plugins & Marketplace33
  10. Cloud & Teams56
  11. Settings & Customization58
  12. Account & Billing28
  13. Troubleshooting84
  14. CLI & API Reference22
  15. Legal & Policies4
  16. Uncategorised22

Local Chat (chat with a local model via Ollama)

A chat area for talking to a model that runs on your own computer — free, private and offline. It is deliberately not a coding session: no tools, no file access, nothing that edits your project. Replies stream in, conversations are saved and renameable, and nothing you type leaves the machine.

What it is

Status: in-development (gated). Hidden by default; revealed by the Settings → Lab → "Local Chat" toggle or by launching with AMC_SHOW_LOCAL_CHAT=1. Until a developer ships it, most users won't see it.

Local Chat is a dedicated Chat area in Omniscio where you talk to an AI model that runs on your own computer — free, private, and offline. It's separate from Omniscio's coding sessions: there are no tools, no file access, and no agent that edits your project. It's just a chat box, like talking to ChatGPT, but the model lives on your machine and nothing you type leaves it.

You open it from the toolbar (the Chat button, a speech-bubble icon). Replies stream in word-by-word. Your conversations are saved in a list on the left — you can rename them or delete them (with an undo).

Where to find it

From the toolbar — the Chat button, a speech-bubble icon. It is hidden by default and revealed by Settings → Lab → "Local Chat"; until it is shipped, most users will not see it.

How it behaves

First-time setup (even if you've never installed anything)

The first time you open Chat, a short setup wizard runs:

  1. Engine check. Omniscio looks for Ollama. If it isn't installed/running, you get a friendly screen with one Download Ollama button (it opens the official download page). You install it the normal way; Omniscio notices automatically the moment Ollama starts running. (Omniscio only opens the page — it never silently installs software.)
  2. Machine scan. Omniscio checks how much memory your PC has (and your graphics card, best-effort) and recommends models that will actually run well. Each suggestion is labelled: ✅ Runs great on your PC, ⚠️ Will be slow, or 🚫 Too big for this PC.
  3. Pick & download. Click a recommended model and Omniscio downloads it through Ollama with a progress bar.
  4. Chat. You're dropped into a conversation. Type and hit Enter for a newline, Ctrl+Enter to send (or click Send). Press Stop to cut a reply short.

Returning later with Ollama and a model already present skips straight to chatting.

Recommended models (the curated list)

Omniscio suggests a small, hand-picked set sized to your RAM (the exact list evolves):

Model Best for Approx download Needs about
Llama 3.2 (1B) Older / low-memory PCs ~1.3 GB 4 GB RAM
Llama 3.2 (3B) Most laptops ~2 GB 8 GB RAM
Llama 3.1 (8B) Best all-rounder ~4.7 GB 16 GB RAM
Llama 3.3 (70B) Powerful PCs only ~40 GB 64 GB RAM

You can also download any other model Ollama supports; the chat model dropdown lists everything you've installed.

What it can and can't do (v1)

  • Can: plain text conversation, multiple saved chats, streaming replies, hardware-aware model recommendations, one-click model download.
  • Can't (yet): edit files or use tools (that's what Omniscio's coding sessions are for), see images, do voice, or use multiple models in one chat. Replies render as plain text (no markdown formatting yet).
  • Cloud models are designed to be addable later (the same Chat area would gain a provider switch) but aren't built in this version.

Privacy & cost

Everything runs on your machine: no API key, no account, no internet needed (after the one-time model download), and no cost. Quality is well below Claude — local models are smaller — so Local Chat is best for quick questions, brainstorming, and offline use, not heavy coding.

Mobile

You can chat from Omniscio's mobile web view (it talks to the same Ollama running on your desktop). The setup wizard and model downloads are desktop-only.

For agents

How it works under the hood

Local Chat uses Ollama — a free, separately-installed app that runs open models like Meta's Llama on your computer. Omniscio talks to Ollama over a local connection (http://127.0.0.1:11434); it never sends your messages to any cloud service. If the local model fails for some reason, Omniscio shows you an error — it will never silently fall back to a paid cloud model behind your back (that would cost money and leak your data).

Related

  • agent-conversations.md — the coding sessions Local Chat is the lightweight, conversation-only alternative to.
  • support-chat.md — the other chat surface, which connects you to the team rather than to a local model.

Last verified 2026-09-23