Fabian G. Williams aka Fabs

Fabian G. Williams

Principal Product Manager, Microsoft Subscribe to my YouTube.

I Swapped My Local Coding Model Overnight From a Hotel. The Agent Graded the Upgrade Itself.

Local models are how I run my community work, my volunteer projects, my side hustles, and my musings, for two reasons: cost and keeping client data away from the labs. A new version dropped, so I upgraded it overnight from a hotel, additively, without touching the old one. Then I handed the new model the reviews and let my OpenCode agent test its own upgrade. It stood up a throwaway server, probed itself, and found three config gaps quietly throttling it. Here is the journey, in tables.

Fabian Williams

8-Minute Read

Activity Monitor showing the M3 Max GPU pinned at 92 percent while an OpenCode agent runs probe requests against a throwaway Qwen3.8 test server on port 8082

Local models are how I run my community work, my volunteer projects, my side hustles, and my musings. Two reasons, both simple:

I Made My Evals Replay Every Task on a Local Model. The Frontier Lead Got Thin.

My agents run on frontier models, but a free local model sits idle on a Mac Mini in my office. So I wired my eval system to replay every writing task on the local model and grade both. Across 10 like-for-like rematches the local model reached statistical parity — and beat the frontier model outright on four of them. Here is the receipts-first system that made it prove it.

Fabian Williams

7-Minute Read

Eval cockpit showing the verdict panel: 31 golden cases, 10 replayed, mean like-for-like delta -0.05, with Volume and Quality gates passing

My agents do real work on frontier models. Every dollar of that work is metered against my OpenAI and Anthropic bills. Meanwhile a perfectly capable local model, gpt-oss:20b, sits on a Mac Mini in my office costing me nothing. The obvious question: for which tasks could the free local model do the job just as well?

Your Agent Said It Did the Work. I Checked the Disk.

I run a six-agent fleet for a non-profit. One morning the agents began reporting work they never did. Here is the fabrication pattern, and the receipts discipline that fixed it.

Fabian Williams

5-Minute Read

A six-agent fleet dashboard showing each agent, its model, cron jobs, and reporting output

Every morning one of my agents sends me a clean status report. Posts cross-posted. Messages delivered. Contacts processed. For a while I read those reports the way you read a receipt from a cashier you trust. Then the automation started giving me time back, so I sat down to run a retrospective and checked the reports against what was actually on disk. The trust did not survive contact with the evidence.

Recent Posts

Categories

About

Fabian G. Williams aka Fabs Site