Our MCP server had 32 files of documentation. A headless agent could read none of them
We built a tool that measures whether an agent with no context can use an MCP server correctly, by running one against it and reading what it did. We ran it on two of our own servers on the same day, with the same model and the same harness. On the first, the agent needed 11 turns and got the right answer by a method it admitted was luck. On the second, it needed 5 turns and told us exactly how far to trust it. The only difference was what crossed the wire. Everything the first server "documented" was sitting in files no remote caller can open.
A cold caller gets the tool list and nothing else
The documentation you wrote for your MCP server does not reach the agents that call it, and you cannot see that from where you sit. MCP, the Model Context Protocol, is the standard by which an agent connects to a server and asks it for a list of tools; what the agent receives is that list, and whatever else the server chooses to send over the connection.
When you publish a server, you are handing those tools to callers you will never meet. Some are people. More and more are agents running headless, with no chat window, no repository access, and no one watching. The question that decides whether a run goes well is not whether the tools are documented but whether an agent that knows nothing about your system can use it correctly on the first try.
An author cannot answer that by looking. You wrote the docstring, so it reads fine to you, and you cannot un-know the domain.
A linter cannot answer it either. It can confirm that a tool has a description. It cannot tell you that the description's first line is the entire onboarding a remote agent will ever receive, that the pool it ranks against is three test fixtures, or that one boolean parameter quietly makes a write unreadable. Those failures are invisible from inside, and they are the ones that burn a run.
Thirty-two files of guidance, none of it on the wire
An MCP server can carry a lot of guidance: skills, references, a contract, a README with the workflow spelled out in order. In one of our servers, that came to five skills and 32 files, including a skill written specifically for headless agents driving the service.
None of it reaches a headless agent driving the service. A remote client sees the protocol, which means an initialize payload, a list of tools, and whatever resources the server chooses to serve. Files in a repository do not cross the wire. The better the documentation gets, the more confident the author becomes, and the further the surprise is from where anyone is looking.
We call this the reachability trap. The best explanation of the system lives in a 28-line module docstring at the top of the server file. The workflow order lives in prose. The warning that the data is fixtures lives in a README. A cold caller gets nine tool descriptions in registration order and a server name, and that is all.
The tool runs a real agent with no context and reads the transcript
It runs three steps. The first two are static and free, and the third is the one that produces evidence.
The connect snapshot reads the server source and reports what a client receives before it acts: the instructions block or its absence, the tools in the order they were registered, resources, providers, and how many lines of explanation never leave the file. The repo inventory lists the guidance the repository holds, the ordering rules stated in prose, and how much of it a remote caller can reach.
The headless trial is the experiment. It spawns a real agent with only the target server attached, file and shell tools disallowed, and a strict config so it cannot inherit anything from the session. It cannot read the repository. That isolation is the point: a remote caller cannot read your skills either, and the difference between those two situations is what we are measuring.
Then we read the transcript for five tells: the agent guessed a value the server never offered; stalled on something it could not discover; succeeded by accident; invented a fact and reported it confidently; hit an error and could not recover.
The score that comes out is called Prior, for how much the agent had to already know. Zero is the optimum. Four means it completed the task wrongly while reporting success, which is the worst outcome on the scale, worse than an outright failure, because nothing downstream knows it happened.
Here is the connect snapshot of the first server we audited, as the tool prints it:
initialize: serverInfo.name = 'recruiting-match' version = None instructions = None
tools/list: 9 tools. resources: 0 prompts: 0 providers: none
not on the wire: a 28-line module docstring
The instructions field is empty, so the caller's whole briefing is nine tool names and their first lines. The 28-line docstring that explains the system is counted on the last line because it is the thing a caller would most want and cannot get.
Eleven turns and luck against five turns and a caveat
We ran the trial on that server with a plain task: work out who is on the bench, find the best-matching role, record the result, and say how much to trust it. The agent took 11 turns. It found the candidate by ranking a throwaway job description and reading the id out of the result, because no tool lists records. It ranked on a placeholder because no tool returns a record's text. It got the right answer, and then it said this about its own work, unprompted:
I got there via a noise-level exploratory ranking that happened to agree with the real answer. That's a lucky corroboration on a 3-role pool, not a demonstrated method.
That is a Prior of 3 rather than 4, because it caught itself.
What let it catch itself is the interesting part: every save came from something in the response body. A warnings list in the payload said the pool was fixtures, and the agent passed that on. A calibration tool refused to let a weak ranking stand, and the agent downgraded its claim. None of the saves came from documentation, because all 32 files sat unread.
The same day we ran the same trial shape against the audit tool's own server, which was built to pass its own test. It has an instructions block that names an entry point and states what the server does not do, and it serves its skill over the wire as resources. The agent took 5 turns, called the entry point it was told to, answered correctly, and separated what it had inspected from what it had judged without being asked.
Then it quoted a caveat back to us that it could only have read over the wire: that a report without a trial is a lint, not evidence. That sentence is in the instructions block. It crossed the wire, the agent read it, and it changed what the agent claimed.
That is a Prior of 0, on the same model, the same harness, and the same day. The legible surface took less than half the turns because the agent did not spend six of them working out what it was looking at.
Three fixes are constructor parameters, and one needs new code
The fixes are cheap, and they are all things the framework already provides. An instructions block is one constructor parameter, and it rides the initialize response, so every client gets it before calling anything. The shape that works is a sentence of purpose, an entry point, and the thing the server does not do. Serving the skills as resources is one provider, and the default mode preserves the front-door-then-depth shape over the wire instead of flattening it. Typed errors and an error-handling middleware turn a dropped stream into a sentence an agent can act on.
The one finding no instructions block fixes is a missing capability. On the first server, the primary tool needed a record and nothing returned one. That is not a documentation problem. The audit found it because the agent could not work around it, and an author would not have, because the author knows how to get the record another way.
What we chose not to build, and what the score cannot tell you
Scoring is a judgement against described levels, on purpose. We considered a script that averages the five tells into a number and did not write it, because that would manufacture precision the evidence does not have. Someone reads the transcript and picks the level.
The trial is never run silently. It spends money and takes minutes, so the server that wraps the audit returns the command for a person to run rather than running it. A tool that spends your money without telling you fails the test it is supposed to administer.
And the numbers above are one trial per server. Eleven turns against five is a real gap on the day, but a single run of a language model is a sample, not a distribution, and we would expect the turn counts to move between runs while the Prior levels stay where they are.
Run the snapshot on your own server
The code is published under our Labs: mcp-legibility, with the two audits above in it as they were run. The connect snapshot is static and costs nothing. Point it at your server and read the first line. If it says instructions = None, a cold caller is in the position our recruiting-match server put its callers in, and everything you wrote for them is on the wrong side of the wire.
We build this kind of thing for companies. How we work.