bicameral: Claude does the thinking, a local model does the typing
Most of what a coding agent spends tokens on isn’t hard. Renaming a function across a package, writing the obvious boilerplate, running the tests and fixing the typo that broke them. A frontier model is overqualified for that work, and every file it reads on the way goes to the cloud.
The hard parts still need a frontier model: deciding what to change, untangling a bug nobody understands yet, and checking the result. So I split the work. bicameral is a small open-source tool that runs Claude Code with two kinds of models. Claude plans and reviews, and an open model on your own machine does the edits.
The name comes from Julian Jaynes’s bicameral mind: one half of the brain gives instructions, the other carries them out.
How it works
claude ──► bicameral ──┬─ worker requests ──► open model on your machine
└─ everything else ──► Claude (api.anthropic.com)
You run bicameral wherever you’d run claude. It starts a small proxy, points Claude Code at it, and adds two things to the session:
- a
local-workersubagent whose model is a placeholder name the proxy recognizes, and - a
delegate-localskill that tells Claude which jobs to hand to the worker and how to brief it.
When Claude delegates, the proxy sees the worker’s model name, rewrites it to your local model, strips your Anthropic credentials from the request, and sends it to the local server. Everything else passes through to Anthropic unchanged. The worker gets its own model name on purpose. Hijacking an existing one like Haiku would also send Claude Code’s internal housekeeping calls to the local model.
bicameral finds Splash, Ollama and LM Studio by itself, and works with anything else that speaks the Anthropic /v1/messages API.
What gets delegated
The split only works if the briefs are good. The worker knows only what Claude writes down, so the skill pushes Claude to give it:
- absolute file paths and the exact change,
- the command that verifies it,
- and what “done” looks like.
Good jobs for the worker are targeted edits, renames, boilerplate, run-the-tests-and-fix loops, and collecting file contents or command output. Claude keeps design decisions, unclear bugs and security-sensitive changes, and it reads the worker’s diff before telling you anything worked. A small model will make mistakes. The point is that it makes them where Claude can see them.
Two ways to run it
As a wrapper. bicameral launches claude behind the proxy. The worker is a normal subagent, so you see its full transcript in the session.
As an MCP tool. Some Claude Code features, notably Remote Control (driving a session from your phone), refuse to run with a custom API endpoint. For those, bicameral also runs as an MCP server:
claude mcp add --scope user bicameral -- bicameral mcp
Claude then gets a delegate tool. Each call starts a separate, stripped-down claude -p in the task’s folder with the local model as its only model. Its shell access is limited to the commands Claude lists in allow_commands, and you approve each call, so you see the brief before anything runs. There are also check and stop tools, so Claude can run several workers in the background and cancel one that loses its way.
What I learned building it
Prompt size matters more than model speed. My first attempt ran a full Claude Code session on the local model. The first two turns took 334 and 416 seconds, almost all of it the local GPU reading Claude Code’s large system prompt. With a slim worker (short system prompt, six tools, no plugins), the first turn took 6.8 seconds. A whole bug-fix task, including running the tests, finished in 35 seconds.
A 27B model is a capable worker. On my test task, fixing a bug in an invoice module with three tests, a Qwen 27B model found the fix in five turns. Its exact-match edit applied on the first try and the tests passed.
Agents will poll if you let them. My first check tool returned whenever the worker did anything, so Claude called it 47 times on one task. Changing it to wait until the worker finishes brought that down to a single check.
Try it
curl -fsSL https://raw.githubusercontent.com/tudalex/bicameral/master/install.sh | sh
Start a local model server, then run bicameral in any project. It’s MIT-licensed, and the code, docs and releases for macOS and Linux are on GitHub at github.com/tudalex/bicameral.
If you’re working out which parts of your own AI workflows could run on local hardware, get in touch.