Section 23 of 45 7 min read

Invalidation: Everything That Breaks a Cache

The invalidation hierarchy, the six habits that break a cache, the silent ones nobody instruments, and a diagnostic procedure.

Objective

Give the complete inventory of what invalidates a cached prefix because most cache problems are one of these and people rarely have the list.

23.1 The Hierarchy #

Caching follows the rendering order—tools → system → messages—and a change at any level invalidates that level and everything after it.12 The table is request-level: it describes what changes between two API calls. A harness sitting in front of the API can route the same user-facing action differently, and §23.1’s closing paragraphs cover the two cases where it does.

What changesToolsSystemMessagesNote
Tool definitions (name, description, schema, order)✘✘✘Total invalidation
Web search or citations toggle✓✘✘Modifies the system prompt
speed setting✓✘✘Per request; a CLI toggle differs
Thinking configurationmodel-specificmodel-specific✘Rendered into the prompt
Effort setting, per requestmodel-specificmodel-specific✘Same mechanism; two exceptions below
tool_choice✓✓✘
Adding or removing an image anywhere✓✓✘
Editing an earlier message✓✓✘Rewrites history from that point
Compaction / summarization✓✓✘Replaces the content it summarizes
Model change✘✘✘Different weights, different cache

The top row is the expensive case, and it is a claim about the definitions in the submitted prefix rather than about servers. Changing the set of tool definitions in the request rewrites the block that renders first, which invalidates everything after it. Whether an MCP server change reaches that row depends on how its tools load (§17.1). Where tool search defers them—the default on supported models—a server connecting, disconnecting or changing its tool list appends content and leaves the cached prefix intact.48 Where definitions load into the prefix instead, adding one invalidates the cache and so does deliberately removing one: that is the case wherever tool search is unavailable or disabled, which the vendor documents for pre-4.5-generation models on Google Cloud’s Agent Platform, for a custom ANTHROPIC_BASE_URL gateway, and for an Azure-hosted Foundry deployment that rejects tool search.48

Even on that eager path the trigger is narrower than “a server changed.” A stdio server whose process simply exits leaves the definitions and the cache alone; a call to one of its tools returns an error rather than running. A remote server that reconnects on its own normally does the same, with one documented exception that proves the rule above: a request sent while it is reconnecting can add the WaitForMcpServers tool if the conversation has not listed it yet, and that addition invalidates the cache once before the tool stays listed for the remainder of the session.48 The server’s own definitions never changed—the set did, which is the thing the cache is keyed on. Otherwise it is connecting, adding tools, or removing one on purpose that costs you. And editing the MCP config file changes nothing by itself; the cost arrives at the restart that connects or disconnects something. This is why “settle your tool set before you start” is a practice rather than fussiness: it costs nothing, and it is what keeps the eager case from surprising you.

Two rows carry exceptions, and both are the same lesson as the reconnection case: the cache is keyed on what the request actually contains, so how a control reaches the request decides what it costs.

Effort. Changing the request-level output_config.effort value always invalidates message blocks, and the model-specific columns above apply to tools and system. Two paths avoid that. Setting effort explicitly to the model’s default is equivalent to omitting it and does not invalidate at all. And on models supporting per-message effort, an effort change carried in a role: "system" message inside messages leaves the cached prefix intact.12 Claude Code uses that second path where it can: on Opus 5.5 and Fable 5.1 with an API key or a Claude subscription, /effort keeps the cache and applies without a confirmation prompt; on other models it warns you first because the next request reads the whole history uncached. The exemption does not extend to Bedrock, Google Cloud’s Agent Platform or a Claude apps gateway, nor where experimental betas are disabled or a HIPAA configuration applies—and on Fable 5.1 it arrived in v2.1.260, so an older installation still pays.48

Fast mode. At the API, switching between speed: "fast" and standard speed invalidates system and message caches.12 A CLI toggle is not the same event. Claude Code’s fast-mode header is part of the cache key, so the first request sent with fast mode on reads the whole history uncached, at fast-mode rates—which is why enabling it early in a session costs less than deep into a long one. That cost lands once per conversation: afterward the header stays and only the speed value varies, which is not part of the key, so turning fast mode off, falling back to standard speed after a rate limit, and turning it on again all keep the cache. /clear and /compact reset that state, since they rebuild the cache anyway.48

Neither exception makes the rows above wrong, and neither generalizes: a thinking-mode or budget change is not covered by either, and a direct API speed change is not cache-safe because a CLI toggle can be.

23.2 The Six Habits #

All of them follow from one idea: caching is a prefix match, so what matters is what changes and when. They are listed separately because each has a distinct trigger.

Do not touch the prefix mid-task

System instructions, tool schemas, and instruction files all sit at the front. Edit them between tasks.

Settle the tool set first

Connect what you need at session start. Where schemas load into the prefix, adding a server on turn thirty re-bills everything before it; where tool search defers them, the same change only appends (§23.1). Settling early is free on both paths, so do it rather than working out which one you are on.

Work in bursts

Caches lapse on inactivity. Batch related questions.

Append, do not rewrite

Editing an earlier message rewrites history from that point onward. A follow-up question is cheaper than re-asking the original question better.

Time compaction deliberately

Compaction replaces the history it summarizes, so the conversation layer is rebuilt from the replacement point on. It does not follow that the whole request reprices: an unchanged tool and system prefix ahead of that point can still be read from cache, which is what the table above says. Pay the suffix once, at a task boundary, on purpose. Section 24 develops this.

Trim the instruction file, but not during work

Shortening it saves on every turn forever. What a mid-session edit costs depends on the harness: where the file is re-rendered into the request it costs the cache from that point once, and on Claude Code it costs nothing and does nothing, because the file was read at session start and the edit does not apply until /clear, /compact or a restart (§15.5).48 Same file, different economics depending on when you touch it and on what is reading it.

23.3 The Silent Ones #

There are four causes that produce neither an error nor an obvious symptom.

Non-deterministic JSON key ordering. Some languages randomize map iteration order during serialization. If your tool arguments or tool results serialize with different key order between requests, the bytes differ and the cache misses. Anthropic’s troubleshooting guidance names Swift and Go specifically.12

# Force stable ordering anywhere you serialize into a prompt.
import json
payload = json.dumps(data, sort_keys=True, separators=(",", ":"))
// C#: System.Text.Json preserves property declaration order for POCOs,
// but Dictionary<string,T> ordering is not guaranteed across runtimes.
// Sort explicitly when the result enters a prompt.
var ordered = new SortedDictionary<string, object>(data);
var payload = JsonSerializer.Serialize(ordered);

A timestamp in the prefix. Any now() above the breakpoint guarantees a miss on every request. Move it below.

A gateway stripping cache markers. If you route through a proxy or gateway, it may drop cache_control markers while returning success—in which case your entire conversation bills as uncached input on every turn with no error anywhere. Verify by checking that cache_read_input_tokens is nonzero after the second request.

Falling below the minimum. Shortening a prefix past the model’s floor silently disables caching (§22.5).

23.4 A Diagnostic Procedure #

When hit rate drops and the cause is unclear, work through the following:

1
Confirm the reads are actually zero.

Check cache_read_input_tokens on a second identical request. If it is nonzero, caching works and the problem lies in prefix variance rather than configuration.

2
Check the breakpoint position.

Is it on a block that changes? This is the most common cause (§22.3).

3
Diff two consecutive rendered requests byte for byte.

Whatever differs first is your answer. Anthropic’s cache diagnostics beta does this for you.12

4
Check for a tool change.

Did anything connect or disconnect?

5
Check for a settings change.

Thinking, effort, tool_choice, images, speed.

6
Check timing.

Are requests more than one TTL apart?

7
Check the minimum.

Is the prefix above the floor for this specific model?

8
Check the gateway.

Is anything between you and the provider rewriting the request?

References cited in this section

2 of 81 · numbering matches the PDF

  1. 12Anthropic, "Prompt Caching," Claude Platform documentation verified September 8, 2026. Vendor documentation; cited as product fact for mechanism, pricing multipliers, minimum cacheable lengths, invalidation behavior, the 20-block lookback window, pre-warming, and data retention. The pricing table in this reference is the primary source for all Anthropic rates quoted in this document.platform.claude.com/docs/en/build-with-claude/prompt-caching ↗
  2. 48Claude Code documentation (settings, hooks, sub-agents, and skills references), verified September 11, 2026, cross-checked against an independently compiled feature and settings snapshot at https://hidekazu-konishi.com/entry/claude_code_features_settings_reference_2026.html. Vendor documentation plus a third-party catalog that links each row back to the official docs. Cited for the settings precedence tree (user, project, project-local, CLI flags, enterprise managed, in ascending precedence, with the managed layer a floor that CLI flags cannot relax for scalar values and deny rules—list-valued keys such as permissions.allow and the sandbox allow and exclusion arrays merge across scopes instead, so lower scopes can add entries and widen access, which allowManagedPermissionRulesOnly exists to prevent for permission rules), the hook event catalog including PostCompact and its auto/manual matcher, the hook exit-code semantics, subagent frontmatter fields and isolation: "worktree", and the documented routing of subagent permission prompts—foreground subagents pass prompts through to the user, background subagents surface them in the main session naming the asking subagent, and auto-denial is a permission-mode behavior rather than a property of delegation. An earlier revision of this entry asserted that subagents cannot raise interactive prompts at all, so approval-required calls always resolve as denials; that was wrong, and §18.5 was corrected before this entry was. The same revision compressed the exit-code semantics to "0 allow, 1 allow with warning, 2 deny," which conflates the handler's process status with the event's decision, and the hook printed in §19.2 is the counterexample: it emits a permissionDecision of deny and exits 0. Exit 0 means the handler succeeded and Claude Code reads the decision from stdout JSON—silence is not approval, it is merely no decision, and the call continues through the normal permission flow. Exit 1 is a non-blocking error that Claude Code proceeds past, not a warning-flavored allow. Exit 2 blocks, but which events can block is event-specific: PreToolUse and UserPromptSubmit block, while PermissionRequest, PostToolUse, Notification, SessionStart and others do not honor it. Read the per-event table rather than a three-value mapping. Also cited, against the memory page and the v2.1.277 release notes of September 18, 2026, for native AGENTS.md loading and its conditions: by default Claude reads AGENTS.md only where no CLAUDE.md, .claude/CLAUDE.md or CLAUDE.local.md sits in the working directory or above it, while a user-level CLAUDE.md, a managed one and .claude/rules/ files do not count against it; a Project instructions setting in /config selects other modes, including loading both; and nested and subdirectory files load on access. The provider limitation this entry previously recorded as current is now version-scoped: the memory page places it before v2.1.281, published September 23, 2026, and directs affected Bedrock users to update rather than describing an ongoing platform gap. The verification date in this entry was accurate when made; this is a product change after it, not a correction to it. The conditions that remain current are an installation before v2.1.277, a disabled agents-md plugin, and in some cases the first session after an upgrade. An earlier revision of this document said Claude Code simply does not read AGENTS.md and presented the import line as a universal requirement; §15.2, §38.6, §43.1 and Appendix D were corrected together. The third-party snapshot is dated May 2026 and its model-name rows are consequently stale against the lineup in reference 12; the mechanism rows cited here were re-checked against the current official pages.docs.claude.com/en/docs/claude-code ↗
PDF↓