Study of 33,097 Agentic Pull Requests Finds AI Coding Agents Favor AGENTS.md Over README and API Docs
A new empirical study finds coding agents overwhelmingly read and write their own instruction files, rarely touch classical documentation, and almost never use docs to recover from failures.
Overview
A new empirical study finds that when autonomous coding agents interact with documentation, they overwhelmingly read and write agent-specific instruction files rather than the READMEs and API references written for humans, and that documentation plays almost no role when an agent gets stuck and needs to recover from a failure. The paper, “From Agent Behaviour to Agent-Friendly Documentation”, posted to arXiv on August 20, 2026 by researchers Zhijun Gao and Jing Chen, analyzed 557 real agentic coding sessions from the SWE-chat dataset alongside 33,097 pull requests from the AIDev dataset, according to the full text of the paper.
What We Know
Instruction files and working notes dominate; classical docs barely register. Across 3,033 coded documentation interactions in the SWE-chat sessions, “instruction files and working notes account for 60.5% of all documentation interactions, versus 10.6% for classical technical documentation and 1.3% for API references,” according to the paper. Broken down further, agent instruction files alone made up 35.4% of all interactions and agent working notes another 25.1%, according to the full-text paper.
The same pattern shows up in pull requests. Looking at the AIDev dataset, the researchers found that AGENTS.md was changed in 692 pull requests, CLAUDE.md in 362, and copilot-instructions.md in 287 — the most-changed individual documentation files in the corpus, according to the full-text paper.
Reading documentation rarely leads directly to a code edit. The researchers measured the probability that an agent edits code in the step immediately after reading documentation and found it to be 0.002, with an adjusted odds ratio of 1.33 once other factors were controlled for; the authors write that “these analyses provide no consistent behavioural evidence for the coupling” between documentation and implementation, according to the full-text paper.
Code changes tend to come first, documentation follows. In multi-commit pull requests that touched both code and documentation, code was touched first about 4.7 times more often than documentation, according to the full-text paper.
Documentation almost never helps when an agent is stuck. Studying 2,034 failure episodes, the researchers found that reading documentation was the first recovery action in only 109 cases — 5.4% of episodes — compared with 631 cases (31.0%) where the agent reread code instead, according to the full-text paper. More broadly, “consultation is self-initiated (70.2%) far more often than it is failure-driven (7.5%),” the paper found.
Consulting docs correlates with less immediate testing, not more. The researchers state: “No explicit documentation-based validation sequence was observed, and consultation is associated with less immediate testing.” The associated statistics show a lift of 0.23 and an adjusted odds ratio of 0.39 for testing activity in the steps immediately following a documentation read, according to the full-text paper.
The sample leans heavily on one tool. The SWE-chat sessions covered six agent families — Gemini CLI, Agent, Claude Code, OpenCode, Codex, and Cursor — but “87% of the corpus comes from a single agent family,” Claude Code, according to the full-text paper.
The authors describe a “two-lobed cycle” instead of a linear pipeline. Rather than a straightforward path from reading documentation to writing validated code, the paper concludes that “agents’ interaction with documentation is a recurrent consultation process that produces reasoning and further documentation and is only loosely coupled to a largely independent code-modification process,” according to the full-text paper.
What We Don’t Know
The paper’s own limitations section tempers several of its findings. The fine-grained labels assigned to ambiguous documentation file types were produced by a language-model classifier rather than checked by human coders, and the authors write plainly: “No human validation of these labels has been performed,” according to the full-text paper. The study also measured documentation strictly by file path, and the authors note that “docstrings, inline comments, and prose embedded in source files are invisible to our instrument,” meaning the reported documentation-interaction rates are likely undercounts, according to the full-text paper. And because the SWE-chat sessions are opt-in telemetry dominated by one agent, and the AIDev pull requests come from public, early-adopter repositories, the authors caution: “Neither dataset necessarily generalises to private codebases,” according to the full-text paper. As of this writing, the paper has not yet drawn coverage from other outlets; it was posted to arXiv on August 20, 2026.
Analysis
The findings carry a concrete implication for teams building or maintaining developer tooling and technical documentation in an era when a growing share of pull requests are agent-authored. The study’s authors argue that if a project wants documentation to actually change agent behavior — rather than simply be present in the repository — prose alone may not be enough. They point out that documentation reads are “frequently followed by further reads,” while jumping between linked documents was “entirely unattested” in the data, which the authors say “motivates studying self-contained documents with locally retrievable structure” instead of documentation that assumes a reader will follow a chain of links, according to the full-text paper. On the finding that documentation consultation coincides with less testing rather than more, the authors suggest that closing that gap “plausibly requires artefacts an agent can execute — runnable examples, doctests, schema contracts — rather than prose,” according to the full-text paper.
The study also flags a housekeeping problem specific to the agent era: the working notes, plans, and verification logs that agents themselves generate and consult are accumulating as a new category of repository artifact that existing tooling doesn’t account for. “Plans, thoughts/ directories, and verification logs accumulate in repositories as durable artefacts. Repository hygiene tooling, code review checklists, and documentation quality metrics currently have no category for them,” the authors write, according to the full-text paper. Combined with the finding that files like AGENTS.md, CLAUDE.md, and copilot-instructions.md rank among the most frequently changed documentation in the AIDev pull-request corpus, the study offers a specific, evidence-based answer to a question many engineering teams are currently guessing at: which documentation surface is worth investing in if the goal is influencing how an AI coding agent behaves.