Subagents and Delegation
What delegation protects, what it costs, why the evidence is weak, the defensible position, and the tools contract a subagent should get.
Cover what delegation actually does to context and cost, when it helps, and the evidence, which is weaker than the marketing.
18.1 The Mechanism #
A subagent is a separate model execution possessing its own context window. The parent invokes it with a task; the subagent works; only its final message returns to the parent. Intermediate tool calls and reasoning stay inside the subagent’s window and never enter the parent’s.43
That isolation is the legitimate reason to use delegation at all. A search task that reads twenty files and produces a three-paragraph answer costs the parent three paragraphs instead of twenty files.
The mechanics: a non-fork subagent starts with a fresh context window; a tool omitted from its definition is not in its session at all; and governance controls exist for spawn depth, concurrency, and budget caps that terminate background subagents.
18.2 It Protects the Window and Raises the Bill #
Both halves are true, but any material presenting only the first half of that sentence is selling something.
Delegation adds a full model invocation carrying its own system prompt and its own tool definitions, and the parent still pays for the returned summary. Whether the prefix is cache-cold depends on which kind of subagent it is: a fresh subagent starts from nothing and pays a cold prefix, while a Claude Code fork inherits the parent conversation and reuses its cache, so the write is one the parent already paid for.48 For a task the parent could have done in three turns, a fresh subagent is usually more expensive—but not as a matter of arithmetic that holds everywhere. A cheaper model on the delegated leg, parent work the delegation avoids, and a large parent context the subagent does not carry can each reverse the ordering, so establish it for your own workload rather than assuming it.
The break-even runs roughly as follows: delegate when the subagent’s context consumption would have exceeded the parent’s window pressure budget, and when the returned summary is genuinely much smaller than the work. Reading twenty files to answer one question qualifies. Reading one file does not.
18.3 The Evidence Is Weak #
Multi-agent orchestration is heavily marketed, while the published evidence does not support it as a default for software engineering.
The most-cited positive result reports a 90.2 percent improvement over a single agent on an internal research evaluation, and in the same publication notes that multi-agent systems use about fifteen times more tokens than chat, that single agents use about four times more than chat, that token usage alone explained 80 percent of performance variance on the relevant benchmark, and that most coding tasks involve fewer truly parallelizable subtasks than research.44 Keep both denominators, because the fifteen is against chat and not against an agent: the matched comparison is roughly fifteen against four. The gain is substantially a token-spend gain, and the vendor with the strongest published result explicitly carves coding out of it.
The counterevidence, moreover, is stronger than the case:
- A peer-reviewed study of over 1,600 annotated execution traces across seven multi-agent frameworks produced a fourteen-mode failure taxonomy in three categories—system design, inter-agent misalignment, and task verification—with inter-annotator agreement at κ = 0.88, and found that obvious interventions such as better prompts and added structure produce partial gains only.45 Multi-agent failure sits in the architecture, beyond the reach of prompt changes.
- Under matched thinking-token budgets, single-agent systems were the best performer or statistically indistinguishable from it at every budget except the lowest, across five multi-agent variants and three models. Sequential multi-agent closes the gap as the context degrades: it is already level at 50 percent corruption in the study’s masking and deletion arms—0.223 against 0.223, and 0.216 against 0.219, on overlapping intervals—and clearly ahead in some 70 percent settings, though not all, since single agents still lead at 70 percent deletion.46 The regime where the switch pays is the one good context engineering exists to prevent. That study covers multi-hop question answering rather than software engineering, which should be stated.
- The useful architectural discriminator from a third vendor: read actions parallelize and write actions do not. Multi-agent fits breadth-first read-heavy exploration and misfits shared context and coordinated writes.47
And there is no published measurement that sub-agent context isolation improves task success on software engineering work. The mechanism is documented; the benefit is asserted.
18.4 The Defensible Position #
Use parallelism for read-heavy exploration, candidate generation, and independent work streams carrying partitioned scope. Concentrate writing and synthesis in one place. Treat multi-agent orchestration for a single coherent change as unproven and expensive.
---
name: dependency-auditor
description: >
Use to audit third-party dependencies for a specific package or lockfile.
Reads manifests, checks advisories, and returns a ranked risk summary.
Read-only — cannot modify files.
tools:
- read_file
- grep
- run_command:npm audit
model: claude-haiku-4-5
---
You audit dependencies. You do not fix them.
For the target given to you:
1. Read the manifest and lockfile.
2. Run `npm audit --json` and parse the result.
3. For each advisory, determine whether the vulnerable path is actually
reachable from application code. Grep for the importing module.
4. Return at most 10 findings, ranked by reachable severity.
Return format: a numbered list. For each finding — package, advisory ID,
whether reachable (yes/no/unknown), and one sentence of evidence.
Do not return the raw audit output. Do not propose upgrades.
Four design decisions appear in that definition, and each reduces risk:
- The
toolslist is explicit and minimal. Three entries, one of them a single named command rather than a general shell. Omittingtoolson most surfaces grants everything, including every MCP tool. This is the line that matters. - A cheap model is specified. Delegated work is often mechanical and does not need the frontier tier.
- The output is bounded. “At most 10 findings” and “one sentence of evidence” cap what returns to the parent, which is the entire point of delegation.
- Scope is stated negatively twice. “You do not fix them” and “Do not propose upgrades” prevent scope creep in a context the parent cannot see.
The tool names above are host-neutral, and translating the shape matters more than copying the strings. Claude Code’s built-ins are Read, Grep, Glob and Bash, and a single-command restriction is written as a Bash permission pattern rather than as a run_command: entry.48 Note also what this definition is and is not: it grants one named command because its job requires running one, so it is a narrow grant rather than a read-only one.
18.5 The Tools Contract #
The persona prose in a subagent definition is advisory. The rest of the frontmatter is not: alongside tools and disallowedTools, Claude Code honors maxTurns as a hard execution limit, permissionMode as the mode the subagent runs under, hooks scoped to that subagent, isolation: "worktree" for checkout separation, and omitClaudeMd to withhold the instruction files.48 Those are runtime controls the model cannot talk its way past, and the tool list is the widest of them rather than the only one.
# Grants everything the parent has, including MCP servers. Almost never correct.
---
name: helper
description: General purpose helper
---
# Grants exactly three capabilities. Correct.
---
name: helper
description: General purpose helper
tools: [read_file, grep, glob]
---
The first definition above grants every tool the parent holds, including every MCP tool connected to the session; the second grants exactly three. Omission is no neutral default. It is a maximal grant. Treat a committed subagent definition as a standing permission grant in the repository because that is what it is.
There is a second reason the tool list matters more for a subagent than for the parent, and it is easy to miss. A subagent has no interactive surface of its own, so what happens at an approval-required call is a property of the host and the execution mode rather than of delegation as such. Claude Code passes a foreground subagent’s permission prompts through to the user, and surfaces a background subagent’s prompts in the main session, naming which subagent is asking.48 Under an auto-deny permission mode—or any genuinely unattended run—the same call resolves as a denial instead, and a permission tier reading as “check with the user” at the parent reads as “fail” wherever nobody is present to answer. Establish which mode your delegation actually runs in before relying on either behavior.
The pattern that follows holds under either mode: bound the subagent’s tool list to what its job actually needs, and keep actions you want a human to see at the parent, where the prompt is unambiguous.48 The dependency auditor above is bounded that way—read, grep, and one named command, with no write tool and no general shell—which is a scope decision first and a prompt-handling decision second. Delegation interacts with verification in a way that has to be stated carefully (§18.6), because the distinction is what supplies the evidence, not which agent runs. A subagent asked to review the work and give an opinion is self-critique by another name, and §10.3’s critic measurements apply to it whatever the permission model does. A subagent that executes the existing test suite, the compiler or the type checker and returns what actually came back is an oracle check (§10.4), and moving that execution into a subagent does not convert it into the model’s own judgment—isolating verbose test output is one of the documented reasons to delegate at all.48 What you must not accept is the middle case: a subagent’s claim that the tests passed is model output, not an oracle result, and it sits in the same position as any other summary crossing back to the parent. Require the run’s actual evidence—the exit status, the failing-test names, the diagnostics—rather than the assurance, and remember that an oracle is only as good as its coverage.
18.6 Failure Modes #
- Delegating trivial work. Strictly more expensive with no isolation benefit.
- Omitting the tools list. Grants everything, silently.
- Unbounded return. A subagent that returns its full transcript has defeated its own purpose.
- Multi-agent for a single coherent change. Roughly fifteen times a chat interaction’s tokens against a single agent’s four (§18.3), an architectural failure taxonomy, no software-engineering evidence, and, at repository scale, a merge-conflict rate from concurrent work that reasoning alone says is higher.
- Treating a subagent’s judgment as verification. Self-critique by another name (§10.3), and the critic measurements apply.30 Delegating an oracle—running the suite, returning its actual output—is a different thing and a documented use of subagents; accept the evidence, not the assurance (§18.5).
References cited in this section
7 of 81 · numbering matches the PDF
- 43Anthropic, "Subagents," Claude Code documentation Vendor documentation; cited as product fact for the isolation mechanism—fresh context window, only the final message returning to the parent, and tool grants bounded by the definition. No published measurement exists of whether sub-agent isolation improves task success on software engineering work; the mechanism is documented and the benefit is asserted.code.claude.com/docs/en/sub-agents ↗
- 48Claude Code documentation (settings, hooks, sub-agents, and skills references), verified September 11, 2026, cross-checked against an independently compiled feature and settings snapshot at https://hidekazu-konishi.com/entry/claude_code_features_settings_reference_2026.html. Vendor documentation plus a third-party catalog that links each row back to the official docs. Cited for the settings precedence tree (user, project, project-local, CLI flags, enterprise managed, in ascending precedence, with the managed layer a floor that CLI flags cannot relax for scalar values and deny rules—list-valued keys such as permissions.allow and the sandbox allow and exclusion arrays merge across scopes instead, so lower scopes can add entries and widen access, which allowManagedPermissionRulesOnly exists to prevent for permission rules), the hook event catalog including PostCompact and its auto/manual matcher, the hook exit-code semantics, subagent frontmatter fields and isolation: "worktree", and the documented routing of subagent permission prompts—foreground subagents pass prompts through to the user, background subagents surface them in the main session naming the asking subagent, and auto-denial is a permission-mode behavior rather than a property of delegation. An earlier revision of this entry asserted that subagents cannot raise interactive prompts at all, so approval-required calls always resolve as denials; that was wrong, and §18.5 was corrected before this entry was. The same revision compressed the exit-code semantics to "0 allow, 1 allow with warning, 2 deny," which conflates the handler's process status with the event's decision, and the hook printed in §19.2 is the counterexample: it emits a permissionDecision of deny and exits 0. Exit 0 means the handler succeeded and Claude Code reads the decision from stdout JSON—silence is not approval, it is merely no decision, and the call continues through the normal permission flow. Exit 1 is a non-blocking error that Claude Code proceeds past, not a warning-flavored allow. Exit 2 blocks, but which events can block is event-specific: PreToolUse and UserPromptSubmit block, while PermissionRequest, PostToolUse, Notification, SessionStart and others do not honor it. Read the per-event table rather than a three-value mapping. Also cited, against the memory page and the v2.1.277 release notes of September 18, 2026, for native AGENTS.md loading and its conditions: by default Claude reads AGENTS.md only where no CLAUDE.md, .claude/CLAUDE.md or CLAUDE.local.md sits in the working directory or above it, while a user-level CLAUDE.md, a managed one and .claude/rules/ files do not count against it; a Project instructions setting in /config selects other modes, including loading both; and nested and subdirectory files load on access. The provider limitation this entry previously recorded as current is now version-scoped: the memory page places it before v2.1.281, published September 23, 2026, and directs affected Bedrock users to update rather than describing an ongoing platform gap. The verification date in this entry was accurate when made; this is a product change after it, not a correction to it. The conditions that remain current are an installation before v2.1.277, a disabled agents-md plugin, and in some cases the first session after an upgrade. An earlier revision of this document said Claude Code simply does not read AGENTS.md and presented the import line as a universal requirement; §15.2, §38.6, §43.1 and Appendix D were corrected together. The third-party snapshot is dated May 2026 and its model-name rows are consequently stale against the lineup in reference 12; the mechanism rows cited here were re-checked against the current official pages.docs.claude.com/en/docs/claude-code ↗
- 44Anthropic, "How we built our multi-agent research system," Anthropic Engineering, 2025. Vendor-published. Reports a 90.2 percent improvement over a single agent on an internal research evaluation, and in the same publication notes roughly fifteen times the token usage of chat for multi-agent systems and about four times chat for single agents, that token usage alone explained 80 percent of performance variance on the relevant benchmark, and that most coding tasks involve fewer truly parallelizable subtasks than research. Both denominators are chat; neither is a matched comparison against one coding agent, and an earlier revision of this document quoted the fifteen as though it were. The vendor with the strongest published multi-agent result explicitly carves coding out of it.
- 45Peer-reviewed study of over 1,600 annotated execution traces across seven multi-agent frameworks, producing a fourteen-mode failure taxonomy in three categories (system design, inter-agent misalignment, task verification) with inter-annotator agreement at κ = 0.88, and finding that better prompts and added structure produce partial gains only. Establishes that multi-agent failure is architectural rather than promptable. Cited via reference 1.
- 46Study comparing single-agent and multi-agent systems under matched thinking-token budgets across five multi-agent variants and three models, finding single-agent best or statistically indistinguishable at every budget except the lowest. The word "only" was wrong in an earlier revision of this document: in the paper's Table 13, sequential agents are already level with single agents at 50 percent corruption—0.223 versus 0.223 under masking, 0.219 versus 0.216 under deletion with overlapping bootstrap intervals—and the 70 percent results are mixed rather than a clean handover, with sequential ahead under masking and substitution but behind under deletion. The aggregate single-agent result stands and individual table cells do not overturn it; what they overturn is a clean corruption threshold. Covers multi-hop question answering rather than software engineering, which should be stated. Cited via reference 1.
- 47Vendor guidance establishing the read/write discriminator for parallelism: read actions parallelize and write actions do not. Vendor-published; the reasoning is mechanical rather than measured, and it is the most useful architectural heuristic available on this question. Cited via reference 1.
- 30CR-Bench, arXiv:2603.11078. Preprint. Measurement of critic subagents on CR-Bench-verified, the 174-case verified subset of a 584-case corpus of real pull-request defects; the verified subset, not the full corpus, is the population every figure here is drawn from. Be precise about the context the reviewers got, which an earlier revision of this document described as full repository context: the prompts in Appendix B supply the repository name, PR number, title, description and diff, and nothing else from the tree. The benchmark's own comparison table claims "Full PR Context," meaning the whole pull request rather than isolated diff hunks, which is a different thing. The paper is explicit about the consequence—it attributes weak recall on usability and functional-suitability defects to context "not fully contained within the PR diff," calling these "closed-context code review agents, lacking access to the broader system state." That makes the measurement a floor for diff-scoped review rather than a verdict on what a repository-aware reviewer could do. Single-shot reviewer at 27.0 percent recall and 3.6 percent precision; iterative self-critique raising recall to 32.8 percent while collapsing signal-to-noise from 5.11 to 1.95 on the large model and 2.89 to 0.91 on the small one. Signal-to-noise is bug hits plus valid suggestions over noise, so below 1.0 the reviewer emits more noise than useful output of any kind, and precision counts confirmed bugs alone against a usefulness rate of 83.6 percent on the same row. Also cited via reference 1.arxiv.org/abs/2603.11078 ↗