Our agent rulebook grew to 442 lines in six weeks. Cutting it to 90 is what made it a harness

The serial harness did not start as a plugin. It started as a 32-line instruction file in one project and grew, over six weeks, into a 442-line rulebook that the agent could not follow and the team could not maintain. Cutting it to 90 lines and moving the rules into hooks is what made it a harness. This post walks the strain points in the order they hit, what each change fixed, what it cost, and the point where adding more stopped paying.

A serial harness strains where any queue strains

The rulebook the agent reads at the start of every session was 32 lines on 8 June and 442 lines on 15 July, and by then the agent had stopped following it. Everything below follows from how we got that number back down.

A serial harness is a queue with checks. It strains where a queue strains: when the thing being built gets too big to fit in one head, and when the checks fall behind the code.

The first sign in our project was the rulebook itself. The top-level instruction file was 32 lines on 8 June. Two days later it was 70. A week later, 84. By 22 June, 222 lines. By 15 July, 442. Every line was added for a reason, usually a reason that had just cost a day: do not edit generated files; run the tests before you stop; do not patch findings into open documents; cite the decision record. Each rule was true on its own.

Together they were a document the agent read once at the start of a session and had mostly forgotten by the tenth tool call.

The second sign was the ticket queue. In week 28 the project opened 64 tickets, and in week 29, 139. Serial work still produces findings, and findings still need filing. Without a fixed shape for a finding, each one became a small essay, median 462 words, most citing the tickets that led to it. The queue was still a queue, but its length stopped being a number anyone could read.

The third sign was drift between the documents and the code. Tickets cited 546 file paths over the summer, and by the end 209 of those paths did not exist. The code moved on and the prose about it stayed where it was.

Iteration one: the two most-broken rules became hooks and stopped being broken

The change: take the rules that could be checked by a program and make them programs. A hook that fires before every edit refuses writes to generated files and names the command that regenerates them. A hook that fires when the session ends its turn runs the tests if there are uncommitted changes, remembers the result for that exact set of files, and never runs twice on an unchanged tree.

What it fixed: the two most-repeated rules disappeared from the rulebook and stopped being broken, because a hook holds whether or not the agent remembers it.

What it cost: each hook is a shell script that has to run on a stock machine with no dependencies. We wrote a small JSON parser in awk rather than require a tool that is not on a fresh Mac. That is a real cost, paid once, and the seven hook scripts total 350 lines.

Iteration two: every finding gets one of four shapes

The change: an intake step. Anything new, whether a bug, an idea, a review, or a handoff from another session, goes through one command that files it as exactly one of: a new gate, an open question, a dated decision, or a regression of something shipped. Long inputs are kept verbatim as sources and never summarised into the standing documents. The active plan is never edited by intake, so findings queue behind it.

What it fixed: the essays stopped. A gate has a fixed shape: what it builds, the demo, what an attacker must try, the sentence the owner says. A finding that cannot be written in that shape is not ready to be a gate, and becomes a question instead. The queue got a length again.

What it cost: friction. Filing a finding takes a minute longer than typing it into whatever document is open, and that minute is what forces the finding into a shape someone can read later.

Iteration three: a script, not the builder, writes the done file

The change: the done file is written by one script and no one else. The script refuses to write unless the plan is that gate's, the adversarial record is clear, and every proof in the done file passes right now. A hook refuses the agent's edits to the file, and each entry carries a checksum the docs check verifies.

What it fixed: "done" meant something. Before, a gate was done when the builder said its tests passed, and the person who has to run the product on a Tuesday afternoon had not touched it. Now that person runs the demo, tries to break it, and says the sentence, or does not.

What it cost: the owner's time, one demo per gate. There is no way around this and we stopped looking for one.

Iteration four: a second agent that did not build the gate tries to break it

The change: after a gate is built and before the owner sees it, a fresh agent that did not build it tries to make it fail, from its own isolated copy of the repository. It attacks neighbours of the happy path, boundaries, order and repetition, the environment, and the acceptance sentence taken literally. Only failing tests come back. A gate with a reproduced failure cannot ship until that exact test passes.

What it fixed: the class of bug where the builder's tests prove the builder's theory. The first draft of one gate looked complete and had four build stoppers a reviewer caught, and a second review of the fixed draft found five more.

What it cost: a second agent's run per gate, and the machinery to isolate it: a snapshot before, a check after that voids the run if anything but tests changed, and a fence that refuses the agent's writes outside the test folder. This is the most code in the harness for a single feature, 567 lines across five scripts.

Iteration five: the rulebook goes from 442 lines to 90

By 22 July the rulebook was 90 lines, down from 442 a week earlier. Nothing was lost. The rules that mattered had become hooks and scripts, and the rest had been restating them.

This is the iteration that makes the others count. A harness that only adds will end up as the first harness we described in part one: a marketplace of guidance nobody reads. The cut is the test of whether the previous iterations worked. If a rule cannot be deleted because no check enforces it, what is missing is the check, not the rule.

We stopped adding when the rulebook fit on a screen

We stopped when three things were true.

The rulebook fit on a screen. Ninety lines, and most of them say where things live.

Every rule that could be a check was a check. What remained as prose was judgment: what makes a good gate, how to write the acceptance sentence, when a finding is a question rather than a gate.

The next proposed addition was about the harness rather than the product. When the conversation turns to improving the process rather than shipping the next gate, the process is big enough, and the right move is to go and build the gate.

We then pulled the process out of the project, stripped the domain, and published it as a plugin, so the next project starts with iteration five and not iteration zero.

What this account cannot show

The order above is a reconstruction. The published plugin has two commits, both from the day it was extracted, so the dates in this post belong to the rulebook's history and not to the hooks themselves. We know the rulebook was 442 lines on 15 July and 90 on 22 July. We do not have a dated commit for the day each hook arrived, and the reasons given for each change are what we remember, not what a log says.

It is also one project. The strain points hit in this order for one small team building one product over one summer, and a larger team or a longer build might hit them in a different order, or hit ones we did not.

Why this matters for a system that has to outlive its author

Every system we deploy has to hold up past its first version. The first version is the easy one: one problem, one owner, the person who built it still in the room. The second version arrives when the business changes, the owner is on holiday, and the person maintaining it did not write it.

The harness is how we build for that day from the first day. Every gate ships with the commands that prove it still works, so a change six months later can be checked against them. The done file is a record only a script can write, so "it used to work" is a fact with a checksum, not a memory. The adversarial run means the tests were written by someone trying to break the thing, which is the only kind of test that matters when nobody remembers why a line is there.

We do not use the harness for its elegance. A system nobody can maintain is a cost, and we would rather build one thing that holds than five that look good in a demo.

If your own rulebook is longer than a screen, the cheapest first step is the one we took in iteration one: find the rule the agent broke most often last week and write the hook that refuses it. Then delete the rule. If you cannot delete it, you have found the check that is missing.

We build this kind of thing for companies. How we work.