A CLAUDE.md changes what a coding agent does. The studies I read say it doesn’t change how many tasks the agent solves.
I keep one repo for the setup my agents run in. This week I wanted to add a line to its CLAUDE.md saying the project’s goal is to implement the most robust and recent techniques for building AI harnesses. First I went looking for what has been measured about these files.
| Study | What they ran | What they found |
|---|---|---|
| Gloaguen et al., ETH Zurich (preprint, latest version 29 September 2026) | 300 SWE-bench Lite tasks, plus 138 tasks from 12 smaller Python repos whose developers wrote their own context files. Four agents, Claude Code with Sonnet 4.5 among them | No general improvement in task success, and over 20% more inference cost |
| Khatri (preprint, July 2026) | Claude Code and Codex on 17 tasks from 3 repos whose AGENTS.md files rated Good or Excellent, 288 runs scored against hidden tests | No measurable change in correctness, bounded to within 10–15 points |
| Lulla et al. (workshop paper, January 2026) | Codex only, 124 merged pull requests from 10 repos, with and without the repo’s AGENTS.md | Median runtime 28.6% lower, output tokens 16.6% lower, similar completion |
They disagree on cost, from 20% more to 17% fewer tokens. They agree that the file doesn’t decide whether the task gets solved. Khatri’s look at the failures explains why: “agents fail on implementation skill — feature design, pattern selection, exact wiring — not missing repository knowledge that a context file could supply.”
Two details from the ETH paper. Files written by the repo’s own developers beat LLM-generated ones by 7 points, and they helped three of the four agents. The fourth was Claude Code. The paper also softened between versions. In February the abstract said context files “tend to reduce task success rates”. By June it said they do “not generally improve” them. If you have seen “AGENTS.md makes agents worse”, that is the first version talking.
The flat results don’t come from agents ignoring the file. When a context file mentioned uv, agents used it 1.6 times per task. When it didn’t, fewer than 0.01 times. With a developer-written file they took 3.34 more steps on average. With a generated file, GPT-5.2 spent 22% more reasoning tokens.
A line in a CLAUDE.md isn’t free even when it’s harmless, because the agent acts on it. Anthropic’s test for each line, “Would removing this cause Claude to make mistakes?”, reads differently once you know that.
The same paper concludes that the files are “useful for specifying non-standard coding practices”. That matches the one big effect I found. Vercel tested agents on Next.js 16 APIs that aren’t in the models’ training data:
It’s a vendor’s own eval with no statistics, but the direction makes sense: a file helps with what the agent can’t get from the repo or from training.
When the ETH team deleted every Markdown file and docs/ folder from the repos, the generated context files started to help, by 2.7% on average. A generated file only pays where nothing else documents the repo.
Eight of the 12 developer-written files in the ETH benchmark had a codebase overview. The verdict: “context files, even developer-provided ones, are not effective at providing a repository overview.” The agent finds its way around the repo anyway. A one-line pointer to the README does the job of a tour of every folder.
IFScale, a 2025 benchmark, gave 20 models a report to write under as many as 500 keyword instructions at once. Accuracy, by number of instructions, with the Claude models run without thinking:
| Model | 50 | 100 | 250 | 500 |
|---|---|---|---|---|
| o3 (high) | 99.6 | 98.2 | 97.8 | 62.8 |
| gemini-2.5-pro-preview | 99.6 | 98.4 | 84.8 | 68.9 |
| claude-sonnet-4 | 98.0 | 94.4 | 77.2 | 42.9 |
| claude-opus-4 | 100 | 94.6 | 67.9 | 44.6 |
The reasoning models held near-perfect through 150 or more instructions, then fell. Claude Sonnet 4 declined steadily from the start. Models mostly dropped instructions outright rather than half-following them, and favoured the ones that came first.
These are 2025 models, nothing from the Claude 4.5 or 5 generation, on a synthetic writing task rather than agents in a codebase, and the paper gives no line count for a CLAUDE.md. The “150–200 instructions” and “under 300 lines” you will see quoted come from a practitioner blog’s reading of this paper.
Wording matters too. In the Vercel eval, “You MUST invoke the skill” did worse than “Explore project first, then invoke skill”.
No study I found tests a purpose or mission statement: whether “use the most recent techniques” or “write high-quality code” helps, does nothing, or sends the agent off doing more than it was asked. There is only opinion, and it points one way. Anthropic’s context-engineering post names “vague, high-level guidance that fails to give the LLM concrete signals for desired outputs or falsely assumes shared context” as one of two failure modes.
So I didn’t write the mission line. “Most recent” and “most robust” pull against each other, since a new technique is the least proven one, and an agent left to weigh the two will lean toward whatever suits the task in front of it. The studies say it would follow the line. Here is what went in instead:
It is the harness Stan’s agents run in: every Claude Code session on the Mac and, through hermes/, the agents on the box. Keep it current with the field, and adopt a technique once it has been measured to work here. Until then a new technique is undecided work, so it goes in docs/roadmap.md with a line on how we would test it.
The second half is the kind of line the evidence supports: concrete, specific to this repo, checkable. The first half is untested. I kept it because it gives the reason for the rule, which Anthropic’s prompting guide recommends.
My own global CLAUDE.md, the one that loads into every session, is 368 lines. It’s next.