Fabian G. Williams aka Fabs

Fabian G. Williams

Principal Product Manager, Microsoft Subscribe to my YouTube.

I Added a Second Local Agent This Week. Here Is the Receipt for Every Human Decision Behind It.

I trialed a second local coding agent, Hermes from Nous Research, on my own MacBook, pointed at the same local Qwen 3.8 model I already run. Installed additively so nothing already working could break, governed with manual approvals, then tested until it proved it behaves. Here is the trial, in tables and screenshots, plus a receipt for the human hours behind writing it up.

Fabian Williams

7-Minute Read

The Hermes agent running against a local Qwen3.8-27B MLX server, with the Apple M3 Max GPU pinned at 97 percent on the first turn

This week I trialed a second local coding agent, Hermes from Nous Research, on my own MacBook Pro M3 Max, pointed at the same local Qwen 3.8 model I already run. I installed it additively, so nothing already working could break, governed it with manual approvals, and did not stop until it proved it behaves. Here is the trial, in tables and screenshots.

Two Agents, One Local Model: Do They Run in Parallel, or Take Turns? I Measured It.

A reader asked what happens if I run two local coding agents against the same Qwen 3.8 model on one Mac at the same time. I thought I knew the answer. I was wrong. So I read the server source, wrote a barrier-synchronized load driver to remove the human-ordering bias, and measured it. Batching is real up to 32 wide, but it is not free, and three innocent-looking choices collapse it back to a single lane.

Fabian Williams

9-Minute Read

A line chart showing aggregate throughput rising with concurrent agents while per-agent decode rate falls, on one local MLX model

A reader on Reddit asked me a sharp question about my last post. I had trialed a second local coding agent, Hermes, pointed at the same local Qwen 3.8 model my other agent already uses. His question was simple. Did I ever run both agents at the same time, two separate harnesses hammering one model on one Mac at once. And what about Hermes spawning its own sub-agents against that same endpoint. Is any of that predictable.

I Made My Evals Replay Every Task on a Local Model. The Frontier Lead Got Thin.

My agents run on frontier models, but a free local model sits idle on a Mac Mini in my office. So I wired my eval system to replay every writing task on the local model and grade both. Across 10 like-for-like rematches the local model reached statistical parity — and beat the frontier model outright on four of them. Here is the receipts-first system that made it prove it.

Fabian Williams

7-Minute Read

Eval cockpit showing the verdict panel: 31 golden cases, 10 replayed, mean like-for-like delta -0.05, with Volume and Quality gates passing

My agents do real work on frontier models. Every dollar of that work is metered against my OpenAI and Anthropic bills. Meanwhile a perfectly capable local model, gpt-oss:20b, sits on a Mac Mini in my office costing me nothing. The obvious question: for which tasks could the free local model do the job just as well?

Recent Posts

Categories

About

Fabian G. Williams aka Fabs Site