Self-hosted LLMs
Open-weight models served on your own GPUs or Apple Silicon, behind an OpenAI-compatible API your team already knows how to use.
I help companies deploy private, self-hosted AI — open-weight LLMs on your own servers, on-device models inside your apps, and ML pipelines that keep sensitive data in-house.
Hands-on engineering, not just strategy. I design, build and deploy — then make sure your team can run it.
Open-weight models served on your own GPUs or Apple Silicon, behind an OpenAI-compatible API your team already knows how to use.
Quantized models running directly on iPhone, iPad, Mac and Android — summarization, chat and vision with no server round-trip.
Question-answering and semantic search over internal documents, mail and scans — indexed, embedded and served inside your network.
Tool-using agents that run on local models, with routing that escalates to a frontier API only when a task genuinely needs it.
Detection, classification and tagging for images and video, from dataset labelling through training to batch or real-time inference.
Which model, which quantization, which box. Benchmarks on your workload so you buy the hardware you need — and not more.
Open-weight models now handle a large share of real business workloads. When privacy, cost or latency matter, running them yourself is often the better trade — and I'll tell you honestly when it isn't.
Prompts, documents and outputs never leave infrastructure you control — simpler GDPR and client-confidentiality conversations.
A fixed hardware budget instead of per-token bills that grow with every new user and every longer context.
Inference next to the user or on the device itself. No network dependency, no third-party outage taking you down.
Open-weight models and standard APIs. Swap models as better ones ship, without rewriting your product.
Every engagement is built around measurable results on your own data, early.
We map the use case, data sensitivity, latency targets and budget, and decide whether local is the right call at all.
A working proof of concept on your data, with measured quality, speed and cost — not a slide deck.
Production serving, monitoring and evaluation set up on your hardware, cloud account or devices.
Documentation, runbooks and a walkthrough so your team owns and operates it from day one.
Fixed-scope where possible, so you know what you're getting. Ongoing retainers available after delivery.
A focused session to pressure-test a plan.
Two to four weeks to prove it works on your data.
From prototype to something people rely on.
I'm a software engineer who has spent years building systems end to end — from backend services and infrastructure to native mobile apps. These days I focus on making modern AI practical to run privately: serving open models on workstations and servers, shipping on-device inference in iOS apps, and wiring up the pipelines around them.
You work directly with me, from the first call to the handover. No account managers, no junior hand-offs.
An open-source tool that lets Claude Code hand mechanical work to an open model on your own machine. Free tokens, files that stay local, and Claude still checks the result.
Read →A practical checklist for deciding between a hosted API and self-hosted open-weight models — and the hidden costs on both sides.
Read →Send a few lines about what you're trying to do, what data is involved and any constraints. I'll reply within two business days, usually with a few questions and a suggested next step.