Fabian G. Williams aka Fabs

Fabian G. Williams

Principal Product Manager, Microsoft Subscribe to my YouTube.

I Made My Evals Replay Every Task on a Local Model. The Frontier Lead Got Thin.

My agents run on frontier models, but a free local model sits idle on a Mac Mini in my office. So I wired my eval system to replay every writing task on the local model and grade both. Across 10 like-for-like rematches the local model reached statistical parity — and beat the frontier model outright on four of them. Here is the receipts-first system that made it prove it.

Fabian Williams

7-Minute Read

Eval cockpit showing the verdict panel: 31 golden cases, 10 replayed, mean like-for-like delta -0.05, with Volume and Quality gates passing

My agents do real work on frontier models. Every dollar of that work is metered against my OpenAI and Anthropic bills. Meanwhile a perfectly capable local model, gpt-oss:20b, sits on a Mac Mini in my office costing me nothing. The obvious question: for which tasks could the free local model do the job just as well?

Sense Before Act: Four Artifacts Every Agent Iteration Must Produce Before It Decides

The matched observation-side discipline to announce-intent-before-action. Four small artifacts (Sensor Roll Call, Conflict Receipt, Gap Map, Cross-Source Dedup) the agent must emit before any decision. Currently running on a real fleet at MACONA; the receipts are public.

Fabian Williams

12-Minute Read

Paying Down Supervision Debt: Why the Five Control Points That Decide Whether Your Agent Ships Have Nothing to Do With Your Model

Five infrastructure control points decide whether an agent reaches production. The Agent Reliability Kit sits on one of them: observability. A public, consumer-readable audit-trail receipt that satisfies both the security review and the finance review with the same URL.

Fabian Williams

9-Minute Read

Five-box diagram of the agent infrastructure control points (Runtime, Identity, Data, Tool/Write, Observability) with the Agent Reliability Kit highlighted on the Observability box, plus a multilayer kill-switch strip across the bottom

By the end of this post you will know which of the five infrastructure control points your agent stack is shipping without, and you will have a concrete pattern for paying down the observability piece of that debt: a public, consumer-readable audit-trail receipt that satisfies both your security team and your finance team with the same document. The receipt was already in production when the broader practitioner conversation started naming the gap.

Recent Posts

Categories

About

Fabian G. Williams aka Fabs Site