Fabian G. Williams aka Fabs

Fabian G. Williams

Principal Product Manager, Microsoft Subscribe to my YouTube.

I Swapped My Local Coding Model Overnight From a Hotel. The Agent Graded the Upgrade Itself.

Local models are how I run my community work, my volunteer projects, my side hustles, and my musings, for two reasons: cost and keeping client data away from the labs. A new version dropped, so I upgraded it overnight from a hotel, additively, without touching the old one. Then I handed the new model the reviews and let my OpenCode agent test its own upgrade. It stood up a throwaway server, probed itself, and found three config gaps quietly throttling it. Here is the journey, in tables.

Fabian Williams

8-Minute Read

Activity Monitor showing the M3 Max GPU pinned at 92 percent while an OpenCode agent runs probe requests against a throwaway Qwen3.8 test server on port 8082

Local models are how I run my community work, my volunteer projects, my side hustles, and my musings. Two reasons, both simple:

I Made My Evals Replay Every Task on a Local Model. The Frontier Lead Got Thin.

My agents run on frontier models, but a free local model sits idle on a Mac Mini in my office. So I wired my eval system to replay every writing task on the local model and grade both. Across 10 like-for-like rematches the local model reached statistical parity — and beat the frontier model outright on four of them. Here is the receipts-first system that made it prove it.

Fabian Williams

7-Minute Read

Eval cockpit showing the verdict panel: 31 golden cases, 10 replayed, mean like-for-like delta -0.05, with Volume and Quality gates passing

My agents do real work on frontier models. Every dollar of that work is metered against my OpenAI and Anthropic bills. Meanwhile a perfectly capable local model, gpt-oss:20b, sits on a Mac Mini in my office costing me nothing. The obvious question: for which tasks could the free local model do the job just as well?

How Do You Trust an Autonomous AI Agent? Evals Are the Answer.

I run an autonomous AI agent at home — 16 cron jobs daily. It says 'done' but did it actually do anything? I built an eval framework to find out. Here's what broke, what I learned, and why agent evals are fundamentally different from LLM evals.

Fabian Williams

10-Minute Read

OpenClaw Eval Dashboard showing mixed results across 9 dimensions — the honest picture after adding freshness, failure rate, and delivery gap scoring

I run an autonomous AI agent on a Mac Mini in my house. She handles 16 daily cron jobs — finances, email triage, outreach campaigns, device monitoring, morning briefings. The agent says “done.” But did it actually do anything? I built a 9-dimension eval rubric to find out. Along the way I discovered that my evals were broken, my agent was better than I thought, and the most important metric isn’t pass/fail — it’s whether a failure is your fault or the agent’s fault.

Recent Posts

Categories

About

Fabian G. Williams aka Fabs Site