Three posts in, you might think I'm about to tell you a local AI box is magic. I'm not. This is the post where I argue against my own product for a few minutes — because the fastest way to lose a professional's trust is to oversell, and because the limits are the most important thing to understand before you buy anything, from me or anyone else.
So here's what running the model on your own hardware does not fix.
It doesn't make the model smarter. A box in your office runs open models — very good ones, but not the frontier. On the standard capability indexes, the small models that run in real time on local hardware land well below GPT-5-class or Opus-class systems on raw knowledge and hard reasoning. If you ask a local model a trivia question about the world, it will be wrong more often than the big cloud one. That gap is real, and anyone who tells you "it's ChatGPT on a box" is lying to you.
Here's the nuance that makes the box work anyway: the gap is huge on knowledge and small on grounded tasks. When you make a model summarize a document you gave it, or draft from a template you provided, the open models are competitive with — sometimes better than — the frontier on faithfulness (staying true to the source instead of inventing). The lesson writes itself: the box should never answer from its own memory. It should answer from your files, your calendar, your call log. Ground everything. A grounded 30-billion-parameter model beats an ungrounded frontier model for the jobs a professional actually needs, because the failure that hurts you isn't "it didn't know" — it's "it made something up."
It doesn't make agents reliable. The dream of "an AI that just runs my practice" is, today, a dream. On realistic office-task benchmarks the best agents finish roughly 30% of the work. Independent testing shows near-perfect success on tasks under a few minutes and single-digit success on tasks that take hours. Gartner expects more than 40% of agentic-AI projects to be cancelled by 2027. This is a property of the models, not of where they run — and local hardware runs smaller models, which are actually more prone to being fooled by a malicious instruction hidden in an email. So the honest design isn't "an autonomous agent." It's deterministic pipelines with the AI confined to the parts it's good at — classifying, extracting, drafting — and a human clicking the button on anything that matters.
It doesn't tell you what to ask. Half of non-users say they simply don't know how to use these tools. Most small firms give no AI training and have no plans to. A blank chat box is intimidating and easy to abandon — which is a big part of why usage is a mile wide and an inch deep. Local doesn't change that. The fix is product design: ship named jobs with buttons — "Morning Docket Brief," "Unbilled Time Sweep" — not a cursor blinking in an empty box. (That's Part 5.)
It doesn't remove the verification burden. Ninety-six percent of people do follow-up work on AI output; nearly half say fixing it takes as long as doing it themselves. Running the model locally doesn't make its output correct. What does help is showing your work — every answer citing the exact file, email, or call it came from, so checking takes ten seconds instead of ten minutes. (That's Part 6.)
And it doesn't make the hardware maintenance-free. A powerful little AI appliance is still a computer doing hard things. Drivers have bugs. Models need tuning. This is real work — it's just work that should land on me, the vendor, not on you. But I'm not going to pretend the box is an appliance you plug in and forget on day one. Anyone who does is setting you up to be their beta tester.
Why am I telling you all this? Two reasons.
One: because it's true, and I'd rather you buy a box understanding exactly what it is than buy a fantasy and feel burned. The whole premise of this product is that you have confidential, high-stakes work and you need a tool you can trust. Trust starts with me not lying to you about the limits.
Two: because the limits are the same everywhere. The model-quality gap, the agent unreliability, the verification burden — those aren't cloud-vs-local tradeoffs. The frontier tools have them too; they just market harder. So the real question isn't "is local as smart as the cloud?" It's: given that every option has these same limits, which one also protects your confidentiality, keeps your costs flat, and can't change the deal on you? Asked that way, the honest weaknesses of the box stop being a reason not to buy it — and the cloud's structural problems (Parts 2 and 3) become the deciding factor.
I trust this box precisely because I know where its edges are. That's the opposite of how most AI gets sold to you, and it's on purpose.
Next up, Part 5: enough theory — the actual jobs a box like this does today, with names, for a real office.
What's the AI limitation that's bitten you hardest? I'll tell you honestly whether a box fixes it.
— Banksy AI
Practical AI, done for you. Runs on your hardware. Your data never leaves.

