<?xml version="1.0" encoding="utf-8" standalone="yes" ?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
  <channel>
    <title>Build In Public on Fabian G. Williams</title>
    <link>https://www.fabswill.com/tags/build-in-public/</link>
    <description>Recent content in Build In Public on Fabian G. Williams</description>
    <generator>Hugo -- gohugo.io</generator>
    <language>en</language>
    <lastBuildDate>Mon, 17 Aug 2026 00:00:00 +0000</lastBuildDate>
    
	<atom:link href="https://www.fabswill.com/tags/build-in-public/index.xml" rel="self" type="application/rss+xml" />
    
    
    <item>
      <title>I Added a Second Local Agent This Week. Here Is the Receipt for Every Human Decision Behind It.</title>
      <link>https://www.fabswill.com/blog/i-added-a-second-local-agent-and-kept-the-receipt/</link>
      <pubDate>Mon, 17 Aug 2026 00:00:00 +0000</pubDate>
      
      <guid>https://www.fabswill.com/blog/i-added-a-second-local-agent-and-kept-the-receipt/</guid>
      <description>This week I trialed a second local coding agent, Hermes from Nous Research, on my own MacBook Pro M3 Max, pointed at the same local Qwen 3.8 model I already run. I installed it additively, so nothing already working could break, governed it with manual approvals, and did not stop until it proved it behaves. Here is the trial, in tables and screenshots.
One housekeeping note before the story. Last week&amp;rsquo;s post got called AI slop on Reddit.</description>
    </item>
    
    <item>
      <title>I Swapped My Local Coding Model Overnight From a Hotel. The Agent Graded the Upgrade Itself.</title>
      <link>https://www.fabswill.com/blog/local-model-upgrade-qwen-38-agent-self-tested/</link>
      <pubDate>Sat, 15 Aug 2026 00:00:00 +0000</pubDate>
      
      <guid>https://www.fabswill.com/blog/local-model-upgrade-qwen-38-agent-self-tested/</guid>
      <description>TL;DR Local models are how I run my community work, my volunteer projects, my side hustles, and my musings. Two reasons, both simple:
 Cost. None of this work should bill against a frontier account. Trust. Client and customer data stays on my own disk, away from the labs.  Bottom line up front: I am running Qwen3.8-27B (4-bit) on my MacBook Pro M3 Max, 128 GB of unified memory and a 40-core GPU, my personal dev rig, driven by an OpenCode harness.</description>
    </item>
    
    <item>
      <title>I Made My Evals Replay Every Task on a Local Model. The Frontier Lead Got Thin.</title>
      <link>https://www.fabswill.com/blog/local-model-frontier-rematch-auto-replay-evals/</link>
      <pubDate>Sun, 02 Aug 2026 00:00:00 +0000</pubDate>
      
      <guid>https://www.fabswill.com/blog/local-model-frontier-rematch-auto-replay-evals/</guid>
      <description>TL;DR My agents do real work on frontier models. Every dollar of that work is metered against my OpenAI and Anthropic bills. Meanwhile a perfectly capable local model, gpt-oss:20b, sits on a Mac Mini in my office costing me nothing. The obvious question: for which tasks could the free local model do the job just as well?
Today I answered it with a system instead of a guess. I taught my eval framework to automatically replay every writing task on the local model right after the frontier model runs it, then grade both with the same judge and record the gap.</description>
    </item>
    
    <item>
      <title>How Do You Trust an Autonomous AI Agent? Evals Are the Answer.</title>
      <link>https://www.fabswill.com/blog/how-do-you-trust-an-autonomous-ai-agent/</link>
      <pubDate>Sat, 28 Mar 2026 00:00:00 +0000</pubDate>
      
      <guid>https://www.fabswill.com/blog/how-do-you-trust-an-autonomous-ai-agent/</guid>
      <description>TL;DR I run an autonomous AI agent on a Mac Mini in my house. She handles 16 daily cron jobs — finances, email triage, outreach campaigns, device monitoring, morning briefings. The agent says &amp;ldquo;done.&amp;rdquo; But did it actually do anything? I built a 9-dimension eval rubric to find out. Along the way I discovered that my evals were broken, my agent was better than I thought, and the most important metric isn&amp;rsquo;t pass/fail — it&amp;rsquo;s whether a failure is your fault or the agent&amp;rsquo;s fault.</description>
    </item>
    
  </channel>
</rss>