---
title: Local Chat (chat with a local model via Ollama)
---

# Local Chat (chat with a local model via Ollama)

## What it is

> **Status: in-development (gated).** Hidden by default; revealed by the
> **Settings → Lab → "Local Chat"** toggle or by launching with
> `AMC_SHOW_LOCAL_CHAT=1`. Until a developer ships it, most users won't see it.

Local Chat is a dedicated **Chat** area in Omniscio where you talk to
an AI model that runs **on your own computer** — free, private, and offline. It's
separate from Omniscio's coding sessions: there are no tools, no file access, and no
agent that edits your project. It's just a chat box, like talking to ChatGPT, but
the model lives on your machine and nothing you type leaves it.

You open it from the toolbar (the **Chat** button, a speech-bubble icon). Replies
stream in word-by-word. Your conversations are saved in a list on the left — you
can rename them or delete them (with an undo).

## Where to find it

From the **toolbar** — the **Chat** button, a speech-bubble icon. It is hidden by default and revealed by **Settings → Lab → "Local Chat"**; until it is shipped, most users will not see it.

## How it behaves

### First-time setup (even if you've never installed anything)

The first time you open Chat, a short setup wizard runs:

1. **Engine check.** Omniscio looks for Ollama. If it isn't installed/running, you get
   a friendly screen with one **Download Ollama** button (it opens the official
   download page). You install it the normal way; Omniscio notices automatically the
   moment Ollama starts running. _(Omniscio only opens the page — it never silently
   installs software.)_
2. **Machine scan.** Omniscio checks how much memory your PC has (and your graphics
   card, best-effort) and recommends models that will actually run well. Each
   suggestion is labelled: ✅ _Runs great on your PC_, ⚠️ _Will be slow_, or
   🚫 _Too big for this PC_.
3. **Pick & download.** Click a recommended model and Omniscio downloads it through
   Ollama with a progress bar.
4. **Chat.** You're dropped into a conversation. Type and hit Enter for a newline,
   **Ctrl+Enter to send** (or click Send). Press **Stop** to cut a reply short.

Returning later with Ollama and a model already present skips straight to chatting.

### Recommended models (the curated list)

Omniscio suggests a small, hand-picked set sized to your RAM (the exact list evolves):

| Model           | Best for               | Approx download | Needs about |
| --------------- | ---------------------- | --------------- | ----------- |
| Llama 3.2 (1B)  | Older / low-memory PCs | ~1.3 GB         | 4 GB RAM    |
| Llama 3.2 (3B)  | Most laptops           | ~2 GB           | 8 GB RAM    |
| Llama 3.1 (8B)  | Best all-rounder       | ~4.7 GB         | 16 GB RAM   |
| Llama 3.3 (70B) | Powerful PCs only      | ~40 GB          | 64 GB RAM   |

You can also download any other model Ollama supports; the chat model dropdown
lists everything you've installed.

### What it can and can't do (v1)

- **Can:** plain text conversation, multiple saved chats, streaming replies,
  hardware-aware model recommendations, one-click model download.
- **Can't (yet):** edit files or use tools (that's what Omniscio's coding sessions are
  for), see images, do voice, or use multiple models in one chat. Replies render
  as plain text (no markdown formatting yet).
- **Cloud models** are _designed to be addable later_ (the same Chat area would
  gain a provider switch) but aren't built in this version.

### Privacy & cost

Everything runs on your machine: **no API key, no account, no internet needed**
(after the one-time model download), and **no cost**. Quality is well below
Claude — local models are smaller — so Local Chat is best for quick questions,
brainstorming, and offline use, not heavy coding.

### Mobile

You can _chat_ from Omniscio's mobile web view (it talks to the same Ollama running on
your desktop). The setup wizard and model downloads are desktop-only.

## For agents

### How it works under the hood

Local Chat uses **[Ollama](https://ollama.com)** — a free, separately-installed
app that runs open models like Meta's Llama on your computer. Omniscio talks to Ollama
over a local connection (`http://127.0.0.1:11434`); it never sends your messages
to any cloud service. If the local model fails for some reason, Omniscio shows you an
error — it will **never** silently fall back to a paid cloud model behind your
back (that would cost money and leak your data).

## Related

- The coding sessions (real Claude agents that edit your files): see the session
  docs — Local Chat is the lightweight, local alternative for conversation only.
- Developer reference for the invariants:
  `.claude/memory/contracts/local-chat-contract.md`.
