Blog
Notes on private, local AI
Practical write-ups from real deployments — serving, on-device inference, hardware and cost. RSS feed.
bicameral: Claude does the thinking, a local model does the typing
An open-source tool that lets Claude Code hand mechanical work to an open model on your own machine. Free tokens, files that stay local, and Claude still checks the result.
When does it make sense to run LLMs on your own hardware?
A practical checklist for deciding between a hosted API and self-hosted open-weight models — and the hidden costs on both sides.