A harness with 122 guidance documents produced 703 tickets in ten weeks. Seven hooks replaced it

We built two harnesses for the same job: keeping a coding agent on course while it builds a product over months. The first was a marketplace of seven plugins, parallel by design, with 122 documents telling the agent how to think. It produced 703 tickets in ten weeks on one project, a third of them still open, and a rulebook that grew to 442 lines before we cut it to 90. The second is one plugin, serial by design, with six commands and seven hooks. It replaces most of the rules with checks that refuse to let bad things happen. The first harness optimised for how much the agent could do at once. The second optimises for how much a person can still understand.

The harness decides whether the people keep the thread

A coding agent's output is only as usable as the harness around it, and ours stopped being usable in July. The numbers below are how we noticed, and what we replaced it with.

The harness is everything around the coding agent that is not the code: the rules it reads at the start of a session, the commands it runs, the checks that fire when it edits a file or ends its turn, and the documents it keeps about what is done and what is next. A good harness makes the agent's next move obvious and its mistakes loud. A bad one lets the agent feel productive while the people responsible for the product lose track of what it did.

We ran the same kind of work through both. The work was a greenfield product with a small team, built over one summer, ticket by ticket, with the agent doing most of the typing. The bad harness came first, because it is the one that looks right on a whiteboard. The good one came out of the wreckage.

The first harness was seven plugins and 13,470 lines of advice

It was a plugin marketplace. Seven plugins, each with its own commands, agents, and skills: one for code quality sensors, one for auditing delivery pipelines, one for reviewing changes against precedent, one for running parallel worktrees, one for decision records, one with twenty-five skills on how to design agents, one for tickets and epics. The pitch was reuse: install once, get the same standards in every repository, and let the agent fan out across several worktrees at a time.

Here is what it measured like by the end of the summer.

  • Files: 216
  • Directories: 121
  • Lines, all files: 27,629
  • Markdown documents: 122
  • Lines of markdown the agent might read: 13,470
  • Largest single skill document: 484 lines
  • Commits, May to September: 64
  • Commits after July: 9

Half the repository was prose. The largest documents were guidance on how to prompt, how to keep context small, and how to improve one's own loops, written for an agent and read by an agent.

After July hardly anyone touched them. Of 216 files, 179 were last changed in July and one in September. The harness did not evolve with the project it served.

The product it served opened 703 tickets and closed 468

The product repository had 1,554 files and 1,380 commits. Over ten weeks it accumulated 703 tickets. At the peak, week 29, the team opened 139 tickets and closed 87. The next week, 137 opened and 100 closed. Opening outran closing in nine of the ten weeks; the one exception was week 28, when 65 closed against 64 opened. When we stopped, 235 tickets were open, the youngest 44 days old, the median 61.

The tickets were not small. The median ticket ran 462 words, and the total ticket text came to 363,655 words, about four novels, most of it written by agents for agents.

They also fed on each other. Of 703 tickets, 633 cited another ticket, and 323 cited three or more. A finding produced a review, the review produced findings, and each finding became a ticket that cited the chain that spawned it. Tickets cited 546 distinct file paths, and by the time we measured, 209 of those paths no longer existed. Sixty-five of the open tickets pointed at files that were gone.

One paraphrased example, names and product removed. A review flagged that one interface adapter reached straight into storage instead of going through the layer the architecture said it should. That is a real finding, and it fits in one line. The ticket it became ran to a "grill", a structured debate with a register of related decisions, a question of whether the exception should be ruled or gated, and references to four earlier tickets and two decision records. Deciding it meant reading all of them, and nobody did. It closed by merge ten days later, as part of something else.

Another, also closed by merge: an evaluation ticket that adds one anti-case to a test suite, so the agent is shown restraining itself where it should. That is useful work. Its body documents twelve paid runs, the cost, and the reasoning, and cites the decision record that justifies the restraint, so a good piece of work ended up stapled to a chain nobody can hold in their head.

The rulebook grew from 32 lines to 442 before we cut it to 90

The project's top-level instructions to the agent went from 32 lines in early June to 442 lines by mid-July, in 78 commits. Then, within one week in late July, they were cut to 90. That cut is the moment the second harness started, and part two of this post walks through it change by change.

None of this was the agent going rogue. It did exactly what the harness asked: review everything, file everything, reason in writing, cite precedent. The harness confused activity with progress, and parallel lanes multiplied the activity. Three lanes each filing findings against each other's work produce three times the reading, and nothing like three times the progress.

The second harness is one plugin with a queue and seven hooks

The second harness is one plugin. Six commands, two agents, seven hooks, eleven shell scripts, and templates for a project's documents. Fifty-seven files and 4,227 lines, templates included.

The idea is serial. Work moves through one ledger of gates, and a gate is one slice the product owner can run and judge. It names what it builds, the demo that shows it, what an attacker must try in order to break it, and the sentence the owner can say once it works. Gates wait in a roadmap. One is in flight, in a plan. Shipped ones sit in a done file with the commands that prove they still work.

/scope  ->  /gate  ->  build  ->  /red  ->  owner runs the demo  ->  /handoff
                ^                                                       |
                +--------------------- next gate ----------------------+

/intake any time something new turns up

The rules are short because the hooks do the work. Where the first harness told the agent not to edit generated files, this one refuses the edit and names the command that regenerates the file. Where the first harness asked the agent to run tests before stopping, this one runs them when the session tries to end its turn on uncommitted changes, and keeps it working once if they fail. Where the first harness had a ticket process, this one has an intake step that files a finding as exactly one thing, a gate or an open question or a dated decision, and never edits the active plan.

The one document only a script may write is the done file. A hook refuses the agent's edits to it, and each entry carries a checksum. A gate moves to done when the owner has run the demo, tried to break it, and said the sentence. Green tests do not move it, and neither does a report that says it is finished.

{
  "generated": ["src/generated/**"],
  "checks": {
    "on_edit": [{ "paths": ["src/**/*.py"], "run": "ruff check $SB_FILE" }],
    "on_stop": [{ "paths": ["src/**", "tests/**"], "run": "pytest -q" }]
  },
  "red": { "test_paths": ["tests/red/**"] }
}

That block is a project's whole configuration for the hooks: which paths are generated and therefore refused, which check runs on each edit, which runs when the session tries to stop, and where the adversarial tests live. It is one committed file, so every clone gets the same guards, and the hooks read it on every call, so a change takes effect on the next tool call.

Planning still happens, and it still happens big. The gate command uses the agent's plan mode, modified so a separate reviewer attacks every draft before the owner sees it, at least twice. The plan can be ambitious, but what ships is one gate.

The parallel part is the one place we kept a second agent, and it is adversarial rather than cooperative. A fresh agent that did not build the gate tries to break it, from its own isolated copy of the repository. Only failing tests come back, and the builder fixes the code until they pass. Two agents, one of them hostile, is the most parallelism we found we could stay on top of.

What the second harness costs

Serial work has a ceiling. One gate is in flight at a time, so a team that wants three features built this week gets one, and the other two wait in the roadmap where their order is visible. We think that trade is right for a small team, and it is still a trade.

The owner pays in time. A gate is not done until the person who has to live with the product has run the demo and tried to break it, which is one sitting per gate that the first harness never asked for.

And the comparison above is one project, one team, one summer. The numbers are real, but they are a before and after on the same product, not two teams running side by side, so anyone measuring their own harness should expect the shape of the result rather than the figures.

The five rules the numbers taught us

  1. Refuse, don't advise. A rule the agent reads is a suggestion it has mostly forgotten by the tenth tool call, whereas a hook that blocks the edit holds whether or not anyone remembers the rule. Move every rule you can into a check.
  2. One thing in flight. Parallel lanes multiply findings faster than anyone can read them. Serial work has a queue, and a queue has a length you can see.
  3. Only the owner says done. Green tests prove the builder's own theory, and the gate is only proven when the person who has to live with it runs the demo.
  4. File findings, never patch them in. Every new thing becomes exactly one of a gate, a question, or a decision. Nothing gets edited because it happened to be open.
  5. Keep the rulebook shorter than the code. When the instructions to the agent grow faster than the product, the harness has become the product, and it is time to cut.

Try it on an empty repository, or count your dead tickets

The good harness is public: serial-build-harness. Install it as a plugin, run the scope interview in an empty repository, and you have the loop above on day one. The bad harness stays private, and its numbers are above, which is the useful part of it.

If you already have a harness, the cheapest check is the one we should have run in June: count the open tickets, then count the ones that cite a file that no longer exists. The second number is how much of your queue is about a product that has already moved on.

We build this kind of thing for companies. How we work.