Recycling Context
The case for discarding context, four operations, what compaction silently costs, when to reset, and how to reset well.
Cover the full lifecycle of a context window: what to keep, what to summarize, what to discard, and when to throw it all away. This is the pivotal quality lever in agentic work.
24.1 The Case for Discarding #
Section 4.3 established the mechanism: models become more likely to err when the context contains their prior errors,8 multi-turn performance degrades roughly 39 percent driven by over-reliance on early assumptions,9 and the largest real failure category is a wrong belief formed early and never revisited.10 Section 12.1 established that performance degrades with input length independent of relevance.35
Taken together, these say something uncomfortable: a long session accumulates commitment rather than understanding.
The corollary is that clearing context is no mere cost-saving measure trading away quality. On a session that has gone wrong, clearing context is the quality measure.
24.2 Four Operations #
| Operation | What it does | Cost | When |
|---|---|---|---|
| Append | Add to the end | Cheap; preserves cache | Default |
| Compact | Replace history with a summary | Rebuilds from the replacement point; prefix ahead of it can still hit | Approaching window pressure |
| Prune | Remove specific items | Invalidates from the removal point | Known-irrelevant bulk |
| Reset | Discard everything, start fresh | Discards history; a matching startup prefix can still hit | Task boundary, or after two failures |
Most harnesses perform the first automatically, and the second automatically at a threshold. The third and fourth remain manual, and that is where the judgment lives.
24.3 Compaction, and What It Silently Costs #
Compaction replaces older messages with a summary once the window fills. It is necessary, and it carries three costs that are routinely underestimated.
It breaks the cache, from the compaction point forward. The summary is different content in the same position, so everything after it is a miss on the next turn. Be precise about the scope, because “one full-price turn” overstates it in both directions. Tool definitions and the system block sit ahead of the replacement and can still be read from cache—Claude Code documents reusing the system-prompt layer across a compaction, and reloading project context from disk, which hits when the instruction files have not changed since session start.48 And the replacement is a summary, so the rebuilt suffix is much shorter than what it replaced. The expensive request is usually the summarization call itself, not the turn after it—and that call reads your warm prefix too, unless the session has been idle past the cache lifetime, which is why /compact costs most on a long-resumed session. Expect a rebuilt suffix, not a cold request.
It loses conversation-only instructions. Vendor documentation warns that specific instructions given early in a conversation may not survive summarization.53 Be precise about which of four things you mean because they behave differently: something said once in chat is summarized and can vanish; a request-level system field is re-sent on every call and is not part of the summarized history at all; a root instruction file may be re-read from disk and re-injected, which Claude Code documents doing after /compact; and a path-scoped rule reloads only when a matching file is next read.48 Anything that must hold for the whole task belongs in a file the agent re-reads, not in a chat message—which is what the root-file behavior automates, and what the pattern below does by hand on hosts that do not.
It loses the wrong things. A summarizer optimizing for information density will drop the constraint you stated once in turn three because it appeared once and looked incidental.
The mitigation for the second and third is the practical heart of this section:
<!-- CONSTRAINTS.md — re-read by the agent at every task boundary -->
# Standing constraints for this task
These apply for the entire task and survive compaction because this file
is re-read, not remembered.
- Target Python 3.11. No 3.12+ syntax.
- Do not modify anything under `src/generated/`.
- Do not add dependencies.
- Every change must keep `make test` green. Run it before claiming done.
- If you cannot satisfy a constraint, stop and say so. Do not work around it.
Re-read CONSTRAINTS.md, then continue with step 4.
That one-line prompt issued after a compaction event is worth more than any amount of prompt engineering applied to the original message. It costs about 125 tokens (with the CONSTRAINTS.md), and it restores the constraints the compaction dropped.
Better still, the re-injection can be made automatic. Claude Code fires a SessionStart hook with a compact matcher after compaction completes, and plain text the hook writes to stdout is added to the context—the documented use for it is precisely this: reload primer context and re-inject project rules.48 Note which event that is. A PostCompact event also exists, and carries the matcher distinguishing an automatic trigger from a manual one, but its stdout reaches only the debug log; the events whose output reaches the model are SessionStart, UserPromptSubmit, UserPromptExpansion, and PostModelSwitch. Registering this on PostCompact gets you a hook that runs, exits clean, and injects nothing. A hook fires on the event rather than on the model’s cooperation (§19.1), so moving constraint re-injection from a habit into a hook converts it from something you must remember into something that happens whether or not you do.
#!/usr/bin/env bash
# .claude/hooks/post-compact.sh - re-inject standing constraints after compaction.
# Registered under hooks.SessionStart, matcher "compact", in .claude/settings.json.
set -euo pipefail
[[ -f CONSTRAINTS.md ]] || exit 0
printf 'Standing constraints for this task, re-read after compaction:\n\n'
cat CONSTRAINTS.md
exit 0
The wider pattern generalizes past compaction. Any constraint you find yourself restating by hand after a predictable event is a candidate for a hook on that event, and the hook catalog covers session start, session end, prompt submission, and both sides of compaction.48 A habit that depends on you remembering it is a control only while you remember it.
24.4 Deciding When to Reset #
The signals follow, in rough order of how strongly each argues for a clean start:
| Signal | Action |
|---|---|
| The agent has failed the same way twice | Reset. The failures are now in context. |
| Moving to an unrelated task | Reset. Prior context is pure noise. |
| The agent is arguing with you about something settled | Reset. A false premise has locked in. |
| Context pressure past your budgeted range (§29.5) | Compact, or reset if at a boundary. |
| The agent is repeating itself | Reset. Classic self-conditioning. |
| A tool returned 40K tokens you do not need | Prune if possible; otherwise reset. |
| Everything is going well | Append. Do nothing. |
The second row is the underrated case. Finishing the migration work and moving on to the API refactor in the same session carries forty thousand tokens of migration context into a task where every one of them is a distractor.
24.5 Resetting Well #
A reset starts over with everything learned and none of the mess. The prompt below extracts that state to a file before the context is cleared.
<!-- The handoff pattern: write state to disk, then clear. -->
Before I clear the context, write a handoff to HANDOFF.md containing:
1. What we were trying to accomplish (2 sentences).
2. What is now done, with file paths.
3. What is left, as a numbered list.
4. Any constraint or decision we established that is not obvious from
the code — especially anything we ruled out and why.
5. The exact command to verify current state.
Do not include narrative. Do not include anything I can read in the diff.
Then, in a fresh session:
Read HANDOFF.md and CONSTRAINTS.md. Then continue from item 1 of the
remaining work. Do not re-do anything listed as done.
This is a deliberate externalization of state from the context window into the filesystem, and it is the pattern behind every published account of long-horizon agentic work: specification, plan, implementation notes, and a running decision log live in repository files rather than in the window, with verification after each milestone.54
Item four is what makes the handoff significant. “We ruled out approach X because it breaks the streaming path” is information that exists nowhere else and would cost the fresh session an hour to rediscover.
24.6 Pruning Specific Content #
Where the harness supports it, removing a single large tool result costs less than a full reset. The function below walks the history backward, keeps the most recent tool results intact, and replaces older oversized ones with a stub.
def prune_large_tool_results(messages, keep_bytes=2_000, keep_recent=3):
"""Replace oversized tool results with a stub, preserving the most
recent ones. Invalidates cache from the first replacement onward."""
out, tool_seen = [], 0
for msg in reversed(messages):
content = msg.get("content")
if isinstance(content, list):
new_blocks = []
# Reversed here too. Parallel tool use puts several results in one
# message (§4.2), so walking blocks forward inside a backward walk
# over messages would preserve the oldest of a batch.
for block in reversed(content):
if isinstance(block, dict) and block.get("type") == "tool_result":
tool_seen += 1
size = len(str(block.get("content", "")).encode("utf-8"))
if tool_seen > keep_recent and size > keep_bytes:
block = {**block,
"content": f"[pruned: {size} bytes, "
f"re-run the tool if needed]"}
new_blocks.append(block)
msg = {**msg, "content": list(reversed(new_blocks))}
out.append(msg)
return list(reversed(out))
Both reversals are required, and the inner one is easy to omit. A turn that made three parallel tool calls comes back as one user message carrying three tool_result blocks, so “most recent” is an ordering within a message as well as across them; iterating blocks forward inside a backward walk preserves the earliest of a batch and prunes the newest, which is exactly backward. Restoring the original block order before returning keeps the rendered history readable. The threshold is measured on encoded bytes because len() on a Python string counts code points—for non-Latin text that is two to four times off the number the stub reports.
The stub text is doing that on purpose. Replacing content with an explicit marker that says what was there and how to get it back is materially better than deleting silently because the model can decide to re-fetch rather than reasoning as though the data never existed.
Two caveats apply. Pruning invalidates cache from the first modified position onward, so the next turn is expensive. And there is no published measurement that selective pruning outperforms full reset on software engineering tasks. The mechanism is sound and the benefit remains unmeasured. [unmeasured]
24.7 The Working Rhythm #
As a working habit, it goes like this:
When the task changes, the session ends.
Re-read after every compaction.
Decide at the top of the range you budgeted in §29.5—compact or finish. That is earlier than the sixty percent people reach for by habit, and on a 1M-token window it is earlier still.
, not a third explanation. The count is an operational heuristic rather than a measured optimum—chosen because an unnecessary reset is cheap and a third explanation compounding into the context is not.
Every reset produces a file.
, so the cache stays warm (§22.6).
That rhythm costs nothing to adopt, and it addresses the largest measured failure categories in agentic work directly.
References cited in this section
7 of 81 · numbering matches the PDF
- 8Peer-reviewed work identifying self-conditioning: models become more likely to err when the context contains their own prior errors. Not a long-context artifact—injecting artificial error histories reproduces the effect, and larger models are more susceptible despite better long-context handling. The same work finds that reasoning-trained models eliminate self-conditioning entirely, which is a meaningful qualification on the practices in §24. Cited via reference 1.
- 9Multi-turn conversational degradation of approximately 39 percent, decomposed into a minor aptitude loss and a large increase in unreliability, driven by models making early assumptions and over-relying on them. The study tested conversational generation rather than agentic coding trajectories, so transfer to coding is plausible and unproven. Cited via reference 1.
- 10Failure taxonomy across 1,794 complete agent trajectories and more than 63,000 execution steps, seven models and three scaffolds. Source of the false-premise rate (30.7 percent), the epistemic/competence/environment breakdown (57.9 / 32.8 / 9.4 percent), the finding that 82 percent of failed trajectories continue executing after the failure is empirically unrecoverable, that the first observable signal surfaces roughly ten steps after the decisive error, and that 71 percent of successful trajectories recover from at least one error. The paper separates three events, and an earlier revision of this document collapsed two of them: the decisive error, t_lock (the point after which no correct recovery is observed), and the first observable signal. The 82 percent continued-execution figure is measured from t_lock; the ten-step lag is measured from the decisive error. Also the source of the prefix-monitor results cited in §25.4: roughly 2 to 3 percent false positives and about 82 percent precision at recognizing a locked-in failure, against recall under thirty percent and a median lead time of zero relative to t_lock, with only 3.7 to 8.7 percent of failures flagged before lock-in. That is failure confirmation rather than loop detection, and §25.4 is scoped accordingly. The strongest published failure taxonomy for coding agents. Cited via reference 1.
- 35Kelly Hong, Anton Troynikov, and Jeff Huber, "Context Rot: How Increasing Input Tokens Impacts LLM Performance," Chroma Research, 2025 Vendor-published research; the publisher sells vector databases and has an interest in the conclusion, which should be stated. Tests 18 frontier models and finds that none use context uniformly, that reliability degrades with input length, and that focused prompts substantially outperform full prompts containing the same relevant material plus distractors. The finding is consistent with independent work on position effects and on effective context length (references 20 and 36) and is the empirical basis for compaction, sub-agent isolation, and aggressive curation.www.trychroma.com/research/context-rot ↗
- 48Claude Code documentation (settings, hooks, sub-agents, and skills references), verified September 11, 2026, cross-checked against an independently compiled feature and settings snapshot at https://hidekazu-konishi.com/entry/claude_code_features_settings_reference_2026.html. Vendor documentation plus a third-party catalog that links each row back to the official docs. Cited for the settings precedence tree (user, project, project-local, CLI flags, enterprise managed, in ascending precedence, with the managed layer a floor that CLI flags cannot relax for scalar values and deny rules—list-valued keys such as permissions.allow and the sandbox allow and exclusion arrays merge across scopes instead, so lower scopes can add entries and widen access, which allowManagedPermissionRulesOnly exists to prevent for permission rules), the hook event catalog including PostCompact and its auto/manual matcher, the hook exit-code semantics, subagent frontmatter fields and isolation: "worktree", and the documented routing of subagent permission prompts—foreground subagents pass prompts through to the user, background subagents surface them in the main session naming the asking subagent, and auto-denial is a permission-mode behavior rather than a property of delegation. An earlier revision of this entry asserted that subagents cannot raise interactive prompts at all, so approval-required calls always resolve as denials; that was wrong, and §18.5 was corrected before this entry was. The same revision compressed the exit-code semantics to "0 allow, 1 allow with warning, 2 deny," which conflates the handler's process status with the event's decision, and the hook printed in §19.2 is the counterexample: it emits a permissionDecision of deny and exits 0. Exit 0 means the handler succeeded and Claude Code reads the decision from stdout JSON—silence is not approval, it is merely no decision, and the call continues through the normal permission flow. Exit 1 is a non-blocking error that Claude Code proceeds past, not a warning-flavored allow. Exit 2 blocks, but which events can block is event-specific: PreToolUse and UserPromptSubmit block, while PermissionRequest, PostToolUse, Notification, SessionStart and others do not honor it. Read the per-event table rather than a three-value mapping. Also cited, against the memory page and the v2.1.277 release notes of September 18, 2026, for native AGENTS.md loading and its conditions: by default Claude reads AGENTS.md only where no CLAUDE.md, .claude/CLAUDE.md or CLAUDE.local.md sits in the working directory or above it, while a user-level CLAUDE.md, a managed one and .claude/rules/ files do not count against it; a Project instructions setting in /config selects other modes, including loading both; and nested and subdirectory files load on access. The provider limitation this entry previously recorded as current is now version-scoped: the memory page places it before v2.1.281, published September 23, 2026, and directs affected Bedrock users to update rather than describing an ongoing platform gap. The verification date in this entry was accurate when made; this is a product change after it, not a correction to it. The conditions that remain current are an installation before v2.1.277, a disabled agents-md plugin, and in some cases the first session after an upgrade. An earlier revision of this document said Claude Code simply does not read AGENTS.md and presented the import line as a universal requirement; §15.2, §38.6, §43.1 and Appendix D were corrected together. The third-party snapshot is dated May 2026 and its model-name rows are consequently stale against the lineup in reference 12; the mechanism rows cited here were re-checked against the current official pages.docs.claude.com/en/docs/claude-code ↗
- 53Vendor documentation warning that as the context window fills, older messages are replaced with a summary and that specific instructions from early in a conversation may not be preserved. Unresolved attribution: the product and page behind this exact wording could not be re-identified, and it is retained only as the basis for the constraint re-injection pattern in §24.3, not as support for any broader claim. Where a host documents its own behavior, that documentation governs: Claude Code states that the project-root instruction file is re-read from disk and re-injected after compaction, that nested files and path-scoped rules reload on access, and that an instruction which disappears was given only in conversation.<sup>48</sup> §24.3 is scoped to conversation-only instructions accordingly.
- 54Nicholas Carlini, "Building a C Compiler with a Team of Parallel Claudes," Anthropic Engineering, February 5, 2026. Vendor-published, n = 1. Describes externalizing specification, plan, implementation notes, and a running decision log into repository markdown, with verification commands and repair after each milestone. An anecdote rather than a measurement—sixteen agents across nearly two thousand Claude Code sessions over two weeks, consuming 2 billion input tokens and generating 140 million output tokens at a total just under $20,000, but the clearest published account of the long-horizon file-state pattern, and the architectural shape is independently justified by references 6, 7, and 8.