,

Is the software factory working?

When we replatformed Twill, we built a software factory: agents write most of the code, a merge queue lands it, CI decides what gets through. Which means there’s a lot of activity, and every release I cut (several a week) includes a list of changes. But I need to know if the factory is working. Are we generating impact or just endless lines of code?

I have a collection of scripts to pull and analyse our output, which I initially put together for a talk. They’ve been useful enough that I keep looking at them, so I added them to the AI dev starter kit. They cover the output and guardrail numbers below; hotfixes and CI spend come from elsewhere.

The things I think are worth paying attention to: output, reliability, and efficiency.

Output

What did we make, and what kind of work was it?

This chart shows features went from 66% of PRs to 16% while volume went up 24x, and guardrails from 12% to 36%. The scripts classify by PR title, which gave 22% guardrails for September; for this chart I ran a more detailed classification on top. Guardrails should ship as part of the work as well: in September 89% of feature PRs carried their own tests, up from 62% in April.

With just two engineers, the feature work is where we’re bottlenecked, whereas the guardrail work is more mechanical and can scale much further with orchestration.

None of this tells you what we actually shipped. That’s what the headline features are for, the same ones that go in the monthly board report. September’s feature share looks low, but actually it was fewer, bigger features – something we changed our process (adding more detailed specs) to support.

September’s included a member referral programme, marketing pages rebuilt in-app, a Chrome extension for LinkedIn referrals, and a state-of-the-business dashboard. (The extension is in another repo, so it’s not in these numbers.) For things that are internal only, like the dashboard, we shipped early and iterated aggressively with actual users, which means more fixes but also earlier feedback and usage.

Reliability

Hotfixes

Hotfixes are my primary reliability metric: how many issues were severe enough that we shipped an additional release the same day. August had about 10, around half of them in launch week. This has dropped since then, in part due to the increased number of guardrails. September had one. October also had one (already!), but that was a big change (the marketing page rebuild – a bunch of DNS and redirect updates), and I’d planned for that release to need extra attention.

Guardrails

The chart above shows the evolution of guardrails. Rules, then checks, then guards, then guards on the agent itself. They reinforce each other: “changes must have tests” was in CLAUDE.md from day one. It wasn’t reliably true until coverage was a required check. That required check doesn’t mean the tests are good, so the agent dispatch instructions now require mutation testing.

These are the kind of things that used to come out of post-incident reviews, but now they need to evolve continuously. If you’re iterating on your product with AI, you should be iterating on your guardrails too.

Efficiency

Median time a PR was open: 1.2 hours in August, 3.5 in September, partly merge-queue contention, partly DB changes (which get more attention and human review) increasing in volume and complexity. PRs landing in two commits or fewer: 67% down to 53%.

CI spend belongs here too. I have another skill (not public, too tied to our stack) that I use to check the spend is adding value. It should be increasing linearly with output, and where it’s increasing above linearly, the value should be clear.

More guardrails and more checks mean work is harder to land. That’s a trade, but with two developers incidents are extremely expensive, so keeping them low is the priority. It also means this system can’t scale indefinitely; with one or two more developers, it would need rethinking.

go deeper

Navigating the AI Shift

From anxious and reactive to genuinely capable. Build real fluency with AI — without the hype.

Comments

Leave a Reply

This site uses Akismet to reduce spam. Learn how your comment data is processed.