Section 15 of 45 10 min read

AGENTS.md and the Instruction File Problem

Where the AGENTS.md standard landed, the interop reality, what belongs in it, the per-turn arithmetic, and the audit.

Objective

Resolve the file-format situation as it actually stands, and give a concrete allocation rule because this is the most confused area of agent configuration.

15.1 Where the Standard Landed #

AGENTS.md originated as a cross-vendor convention among teams working on several agent products, each of whom needed the same file and none of whom wanted five competing names for it. It is now stewarded by the Agentic AI Foundation, a Linux Foundation project launched in December 2025 with founding contributions of MCP, AGENTS.md, and the goose framework.39

Neutral governance changes two things in practice. The filename is stable, so adopting it is not a bet on one vendor’s roadmap. And the specification stays deliberately thin—it is hard to bolt vendor-specific fields onto a spec nobody owns, which is why AGENTS.md still has no required frontmatter while every tool-specific format keeps accumulating options.

The file is plain Markdown at the repository root with no required fields. Nesting is supported and the specification’s rule is that the file nearest the code being edited wins—which the hosts implement unevenly, so §14.4 is where to check before relying on it.

15.2 The Interop Reality #

As of September 2026, here is where things stand:

FormatRead by
AGENTS.mdCodex, Cursor, Copilot, Jules, Aider, Zed, Windsurf, Devin, and others; Gemini CLI only once context.fileName names it
CLAUDE.mdClaude Code, in preference to AGENTS.md wherever one is present
.cursor/rules/*.mdcCursor (in addition to AGENTS.md)
.github/copilot-instructions.mdGitHub Copilot (in addition to AGENTS.md)
.cursorrulesCursor, legacy and slated for deprecation; migrate to an Always Apply rule

Claude Code was the notable exception until recently, and the exception has narrowed rather than disappeared. Native AGENTS.md reading shipped in v2.1.277 on September 18, 2026, and it is conditional: by default Claude reads AGENTS.md only when no CLAUDE.md, .claude/CLAUDE.md or CLAUDE.local.md sits in the working directory or above it—a user-level or managed CLAUDE.md does not count, and neither do .claude/rules/ files. Where one of those three does exist, the CLAUDE.md files win and AGENTS.md is not read at all. A Project instructions setting in /config overrides the default, including a mode that loads both. The provider restriction that older write-ups report has since lapsed: before v2.1.281, published September 23, 2026, some sessions—Bedrock among them, and any session with telemetry disabled—read CLAUDE.md only, and the documented remedy now is to update rather than to work around the platform.48 What still disables native reading is an installation older than v2.1.277, a disabled agents-md plugin, and in some cases the first session after an upgrade.48

So the import is no longer universally required, and it remains the right answer in three situations: a repository that already carries a CLAUDE.md for Claude-specific additions, a team on a provider or version without native support, and anyone who wants one loading path rather than one that depends on which files happen to be present. Written as a one-line import at the top of CLAUDE.md, both files exist without duplication:

@AGENTS.md

<!-- Claude Code specific additions below, if any. -->

That single line establishes AGENTS.md as the source of truth and reduces CLAUDE.md to a pointer. Two cheap CI checks protect that arrangement—that the pointer is still a pointer, and that the canonical file has not grown into a manual:

#!/usr/bin/env bash
# ci/check-agent-files.sh — the import line is intact and the files are small.
# This is NOT a drift check: see the note below the script.
set -euo pipefail

expected="$(printf '@AGENTS.md\n')"
if [[ "$(head -n1 CLAUDE.md)" != "$expected" ]]; then
  echo "CLAUDE.md must begin with '@AGENTS.md'. Do not duplicate content." >&2
  exit 1
fi

# Instruction files are cached prefix on every turn. Keep them small.
limit=2000
for f in AGENTS.md .github/copilot-instructions.md; do
  [[ -f "$f" ]] || continue
  words=$(wc -w < "$f")
  if (( words > limit )); then
    echo "$f is $words words (limit $limit). Move procedures into skills." >&2
    exit 1
  fi
done

Be exact about what that buys, because it is easy to file it as a drift gate and stop thinking. It checks the import line and the word counts. It does not compare derived content against the canonical file, so it passes an AGENTS.md telling the agent to run the tests alongside a copilot-instructions.md telling it never to—which is the failure the section is about. Real drift detection regenerates each derived file from the source and compares it to what is committed, and needs the generator to exist first; §43.2 builds that script and gives it the --check mode this one lacks. Its scope is narrow: that generator is written for scoped rules, reading .agent/scoped/*.md and checking the two directories it emits into, so its --check will correctly exit 0 over synchronized scoped outputs while the root AGENTS.md and copilot-instructions.md contradict each other. It is the pattern to copy rather than the finished gate. Point the same regenerate-and-compare discipline at the root files—derive them from the canonical file and check them the same way—and the contradiction above is caught. Where a surface does not support imports, that is the tool to reach for.

The position that resolves the format question: pick AGENTS.md as the single source of truth, derive everything else, and let CI compare the derived files against it. Adopting it costs one file and one import line, and because the specification is under neutral governance (§15.1) it is not a bet on any vendor’s roadmap. The failure this prevents has nothing to do with confusion over filenames. It is four files that started identical, diverged over a quarter, and now give different developers different behavior from the same repository—with no error anywhere to tell them.

15.3 What Belongs in It #

The evidence in §8.6 establishes the standard: content the agent could not infer, and nothing else. Concretely, six section types qualify.

# AGENTS.md

## Commands
- Build: `make build`
- Test: `make test` (never `pytest` directly — it skips the fixtures setup)
- Lint: `make lint` — must pass before any commit
- Single test: `make test T=tests/path/to/test_file.py::test_name`

## Constraints
- Target Python 3.11. Do not use 3.12+ syntax.
- No new runtime dependencies without a comment justifying the addition.
- `src/generated/` is machine-generated. Never edit by hand.
- Migrations are forward-only. No `DROP COLUMN` in any migration.

## Conventions the linter does not catch
- Errors cross module boundaries as domain exceptions, never as HTTP status codes.
- Every public function that touches the database takes an explicit session
  parameter. There is no ambient session.

## Testing
- Integration tests need `docker compose up -d postgres` first.
- Tests must not hit the network. If one does, it is a bug in the test.

## Security
- Never log request bodies on `/v1/auth/*` — they contain credentials.
- Secrets come from the environment. Never read `.env` directly in application code.

## Pull requests
- Title format: `[area] imperative summary`
- Every PR touching `src/api/` needs a note about backward compatibility.

Every line there fails the “could the model discover this?” test in the right direction. The security section matters in particular: a study that collected 2,303 context files across 1,925 repositories hand-labeled a 332-file Claude Code subset of them and found testing content in 75 percent, implementation detail in 69.9 percent, and architecture in 67.7 percent, with security and performance each appearing in only 14.5 percent.40 The artifact steering every agent almost never mentions the two quality attributes that degrade fastest under agent-generated code. Three security lines is a very cheap correction.

15.4 What Does Not Belong #

ContentWhy notWhere instead
Project descriptionThe README existsNowhere
Architecture overviewDiscoverable from structureNowhere, or a diagram in docs
Style rules the linter enforcesThe linter is the oracleLinter config
Multi-step proceduresLoaded every turn for rare useA skill (§16)
Historical rationaleRule without reason is actionableAn ADR
Anything under **/*.tsx onlyCosts on every non-React turnPath-scoped instructions

15.5 The Per-Turn Arithmetic #

Instruction files are billed on every turn regardless of relevance, as tool schemas and skill metadata are (§14.1), and that deserves explicit arithmetic at least once.

A 200-word line of guidance is roughly 260 tokens. On a team of fifty developers averaging fifty turns a day:

50 devs × 50 turns × 22 days = 55,000 turns/month
260 tokens × 55,000          = 14.3M tokens/month

At Sonnet 5 base input ($2.00/MTok), uncached:  $28.60/month
At Sonnet 5 cache read ($0.20/MTok), cached:     $2.86/month

Three dollars a month for one paragraph, if it caches. Twenty-nine if it does not. The number is small; the discipline it should produce is not, because the file usually contains thirty such paragraphs, and most of them buy nothing.

The subtler cost is easy to overstate, and being exact about it matters. A cache is invalidated by a change to the submitted prefix, not by a change on disk. Claude Code reads the project-root and user-level file once at session start and holds it in memory: editing it mid-session neither invalidates the cache nor applies the edit, and the new content loads on the next /clear, /compact or restart.48 Nested and path-scoped files behave differently—they load when a matching file is first read, so an edit before that point does take effect. The file also arrives as a user message after the system prompt rather than inside the system block,48 so on a harness that does re-render it mid-turn, what it invalidates is everything from that message onward rather than very nearly the whole request. Section 23 covers the mechanics. The operational rule survives either way, and on Claude Code for a sharper reason than cost—a mid-session edit that silently does not apply is its own failure mode: edit instruction files at task boundaries, not during work.

15.6 Path Scoping #

A glob in frontmatter keeps a rule out of the request until something matches it, which is the most underused mechanism on every platform. Be precise about what that buys, because it is a claim about loading and not about residency: on Claude Code a paths: rule loads when a matching file is first read, and from then on its content is part of the conversation history, so later turns—including unrelated ones—still carry it, at cached-read rates rather than free.48 Compaction bounds that: afterward the rule reloads only when a matching file is read again. Retention past the first match is host-specific, so check yours rather than assuming it matches either extreme. The saving is front-loaded: a rule that never matches never costs anything, and a rule that matches once in a long session is closer to always-on than to absent. Two examples follow, the first scoping React conventions to component files and the second scoping migration rules to the two directories where migrations live.

---
applyTo: "**/*.tsx"
---
- Components are function components. No classes.
- Server components by default; add `"use client"` only when a hook requires it.
- Never import from `@/lib/server/*` in a client component.
---
description: Database and migration conventions
globs: ["migrations/**/*.sql", "src/db/**/*.py"]
alwaysApply: false
---
- Every migration has both `up` and `down`.
- Never alter a column type in place. Add, backfill, drop in a later migration.
- Index creation on tables over 1M rows uses `CONCURRENTLY`.

The frontmatter key differs by platform—paths in .claude/rules/ on Claude Code, applyTo on Copilot, globs on Cursor—and the matching mechanism is the same on each. What is not established as identical is what happens after a match: the retention behavior above is documented for Claude Code, and this document does not have equivalent evidence for the others, so do not assume a scoped rule unloads itself on the next non-matching turn anywhere. Section 43 gives the full crosswalk.

15.7 The Audit #

The audit takes forty-five minutes, requires no permission, and returns more for the effort than anything else here.

1

Read the file end to end. Most owners never have.

2

Delete anything the model could learn from the repository.

3

Move every procedure into a skill.

4

Move every file-type-specific rule into a path-scoped file.

5

Move every rationale into an ADR and keep only the rule.

6

Add three security lines if there are none.

7

Count words. If the always-on file is over ~500 words, go back to step 2.

8

Commit, then run the same task on the old and new file in two fresh sessions and compare.

Step eight is the only real experiment available, and doing it once grounds the change in something other than theory.

References cited in this section

3 of 81 · numbering matches the PDF

  1. 39Linux Foundation, "Linux Foundation Announces the Formation of the Agentic AI Foundation (AAIF), Anchored by New Project Contributions Including Model Context Protocol (MCP), goose and AGENTS.md," December 2025 and the AGENTS.md specification at https://agents.md/. Primary announcement plus specification. Establishes neutral governance for both MCP and AGENTS.md, which is the durable signal for an organization standardizing on either.www.linuxfoundation.org/press/linux-foundation-announces-the-formation-of-the-agentic-ai-foundation ↗
  2. 48Claude Code documentation (settings, hooks, sub-agents, and skills references), verified September 11, 2026, cross-checked against an independently compiled feature and settings snapshot at https://hidekazu-konishi.com/entry/claude_code_features_settings_reference_2026.html. Vendor documentation plus a third-party catalog that links each row back to the official docs. Cited for the settings precedence tree (user, project, project-local, CLI flags, enterprise managed, in ascending precedence, with the managed layer a floor that CLI flags cannot relax for scalar values and deny rules—list-valued keys such as permissions.allow and the sandbox allow and exclusion arrays merge across scopes instead, so lower scopes can add entries and widen access, which allowManagedPermissionRulesOnly exists to prevent for permission rules), the hook event catalog including PostCompact and its auto/manual matcher, the hook exit-code semantics, subagent frontmatter fields and isolation: "worktree", and the documented routing of subagent permission prompts—foreground subagents pass prompts through to the user, background subagents surface them in the main session naming the asking subagent, and auto-denial is a permission-mode behavior rather than a property of delegation. An earlier revision of this entry asserted that subagents cannot raise interactive prompts at all, so approval-required calls always resolve as denials; that was wrong, and §18.5 was corrected before this entry was. The same revision compressed the exit-code semantics to "0 allow, 1 allow with warning, 2 deny," which conflates the handler's process status with the event's decision, and the hook printed in §19.2 is the counterexample: it emits a permissionDecision of deny and exits 0. Exit 0 means the handler succeeded and Claude Code reads the decision from stdout JSON—silence is not approval, it is merely no decision, and the call continues through the normal permission flow. Exit 1 is a non-blocking error that Claude Code proceeds past, not a warning-flavored allow. Exit 2 blocks, but which events can block is event-specific: PreToolUse and UserPromptSubmit block, while PermissionRequest, PostToolUse, Notification, SessionStart and others do not honor it. Read the per-event table rather than a three-value mapping. Also cited, against the memory page and the v2.1.277 release notes of September 18, 2026, for native AGENTS.md loading and its conditions: by default Claude reads AGENTS.md only where no CLAUDE.md, .claude/CLAUDE.md or CLAUDE.local.md sits in the working directory or above it, while a user-level CLAUDE.md, a managed one and .claude/rules/ files do not count against it; a Project instructions setting in /config selects other modes, including loading both; and nested and subdirectory files load on access. The provider limitation this entry previously recorded as current is now version-scoped: the memory page places it before v2.1.281, published September 23, 2026, and directs affected Bedrock users to update rather than describing an ongoing platform gap. The verification date in this entry was accurate when made; this is a product change after it, not a correction to it. The conditions that remain current are an installation before v2.1.277, a disabled agents-md plugin, and in some cases the first session after an upgrade. An earlier revision of this document said Claude Code simply does not read AGENTS.md and presented the import line as a universal requirement; §15.2, §38.6, §43.1 and Appendix D were corrected together. The third-party snapshot is dated May 2026 and its model-name rows are consequently stale against the lineup in reference 12; the mechanism rows cited here were re-checked against the current official pages.docs.claude.com/en/docs/claude-code ↗
  3. 40Study collecting 2,303 context files across 1,925 repositories and three agentic coding tools. Separate the corpus from the analysis population, which an earlier revision of this document did not: the content percentages come from a manually labeled subset of 332 Claude Code files, labeled by two inspectors with a third resolving 438 disagreements, and not from a content analysis of all 2,303. On that subset, testing content appears in 75 percent, implementation detail in 69.9 percent, and architecture in 67.7 percent, with security and performance each appearing in only 14.5 percent. The source carries some internal count inconsistencies around label totals; do not infer a replacement denominator from the rounded percentages, and do not substitute a later revision for the pinned v1, since a restatement there would not be the figures quoted here. The artifact steering every agent almost never mentions the two quality attributes that degrade fastest. Cited via reference 1; the percentages quoted here match version 1 of the preprint which is the version to pin, since a later revision may restate them.arxiv.org/html/2511.12884v1 ↗
PDF↓