Prefix Caching, Mechanically
Caching in one sentence, what is stored, where the breakpoint goes, minimum cacheable length, TTLs, and how to measure hit rate.
Explain caching precisely enough that you can predict a hit or a miss before making the request, which is the difference between caching working and caching being configured.
22.1 One Sentence #
Caching is a prefix match from the first token forward, and it holds up to the first thing that changed.
Everything else follows from that single fact. Change something at position 100 of a 50,000-token prefix, and 49,900 tokens of cache have been invalidated. The cost is not proportional to the size of the change; it is proportional to everything after it.
22.2 What Is Stored #
What is stored is key-value tensors, the model’s internal attention state for the prefix, rather than a readable copy of your prompt. Both major providers state this explicitly.12,49 Anthropic further specifies that KV representations and cryptographic hashes are held in memory only, are not stored at rest, and that caching is eligible for zero-data-retention configurations.12
This is invariably the first question a security-minded reviewer asks, and the answer should be ready.
Isolation operates per organization on every provider, and per workspace within an organization on the Claude API, Claude Platform on AWS, and Microsoft Foundry; Bedrock and Google Cloud isolate at the organization level only.12 If you run multiple workspaces, that difference affects your hit rate.
22.3 Where the Breakpoint Goes #
One rule resolves ninety percent of caching problems:
Place the breakpoint on the last block whose content is identical across the requests you want to share a cache.
The failure this prevents is the most common caching bug, and it fails silently.
Suppose blocks 1 through 5 are a large static system context and block 6 contains a timestamp plus the user message. You set the breakpoint on block 6:
- Request 1: cache write at block 6. The stored hash includes the timestamp.
- Request 2: the timestamp differs, so the prefix hash at block 6 differs. The lookback walks back through blocks 5, 4, 3, 2, 1, but no request ever wrote an entry at any of those positions, so there is nothing to find. No hit.
You pay a fresh cache write on every single request and never receive a read. Your usage fields will show cache_creation_input_tokens climbing and cache_read_input_tokens at zero, which is the signature to watch for.
The mechanism behind that, stated precisely by Anthropic’s documentation:12
Marking a block writes exactly one entry: a hash of the prefix ending at that block. No entries are written for earlier positions.
The system computes the hash at your breakpoint and, failing a match, walks back one block at a time looking for a prior write—not for stable content.
Beyond that, checking stops.
The fix is to move the breakpoint to block 5, which is the last block that stays the same.
# WRONG — breakpoint on content that changes every request
system = [
{"type": "text", "text": LARGE_STATIC_CONTEXT},
{"type": "text", "text": f"Current time: {now()}. User: {user_msg}",
"cache_control": {"type": "ephemeral"}}, # ← hash differs every time
]
# RIGHT — breakpoint at the end of the stable prefix
system = [
{"type": "text", "text": LARGE_STATIC_CONTEXT,
"cache_control": {"type": "ephemeral"}}, # ← stable hash
]
messages = [
{"role": "user", "content": f"Current time: {now()}\n\n{user_msg}"},
]
Automatic caching falls into the same trap, since it places the breakpoint on the last cacheable block, which in the structure above is the one that changes. Use an explicit breakpoint when your suffix varies.
22.4 Multiple Breakpoints #
Up to four on Anthropic; up to four cache writes per request on OpenAI GPT-5.6+.12,49 Breakpoints themselves are free. You pay for content cached and read, not for markers, so the only reason to use fewer than four is that you do not need them.
They pay off when sections change at different rates:
response = client.messages.create(
model="claude-opus-5",
max_tokens=2048,
tools=[
*OTHER_TOOLS,
{**LAST_TOOL, "cache_control": {"type": "ephemeral"}}, # 1: tools, ~never change
],
system=[
{"type": "text", "text": SYSTEM_INSTRUCTIONS,
"cache_control": {"type": "ephemeral"}}, # 2: instructions, weekly
{"type": "text", "text": knowledge_base_snapshot,
"cache_control": {"type": "ephemeral"}}, # 3: RAG context, daily
],
messages=[
*history,
{"role": "user", "content": [
{"type": "text", "text": user_message,
"cache_control": {"type": "ephemeral"}}, # 4: conversation, per turn
]},
],
)
The payoff structure: update the knowledge base and segments 1 and 2 survive. Change the conversation and 1, 2, and 3 survive. Change a tool and everything is invalidated because tools render first.
There is a second and less obvious reason to use multiple breakpoints. The lookback window is 20 blocks, and a growing conversation can push your breakpoint 20 or more blocks past the last write—at which point the lookback misses an entry that exists. Placing a second breakpoint closer to that position from the start accumulates a write there before you need it.12
22.5 Minimum Cacheable Length, and the Trap Below It #
Prefixes shorter than a model-specific minimum are not cached at all, and no error is returned to signal it. On Anthropic the minimum ranges from 512 tokens (Opus 5, Fable 5.1) to 4,096 (Opus 4.6, Haiku 4.5), with 1,024 for Sonnet 5 and Opus 4.8.12 On OpenAI it is 1,024 for GPT-5.6 and later, and on earlier models it varies with the request settings.49
The consequence is counterintuitive: if your reusable prefix is just below the minimum, lengthening it with useful stable material can be cheaper than leaving it short.
OpenAI publishes the break-even formula. With minimum cacheable length M, original prefix length L < M, cache-read multiplier r, cache-write multiplier w, and N total requests, expansion is cheaper when:
L > M × ( r + (w − r) ÷ N )
With M = 1,024, r = 0.1, w = 1.25, the crossover is 102.4 + 1177.6/N tokens. Across ten requests, expanding a prefix of at least 221 tokens to 1,024 is cheaper. As reuse grows the crossover approaches 102.4 tokens; a 103-token prefix needs at least 1,963 requests to benefit, and a 102-token prefix never does.49
def should_expand(current_len, minimum, n_requests, r=0.1, w=1.25):
"""Is it cheaper to pad a sub-minimum prefix up to the cacheable floor?"""
breakeven = minimum * (r + (w - r) / n_requests)
return current_len > breakeven, breakeven
print(should_expand(current_len=800, minimum=1024, n_requests=50))
# (True, 125.952) — pad it; 800 is well above the 126-token crossover
Pad with material that helps—additional examples, more house rules, a fuller rubric—rather than with filler. And measure that behavior stays stable because you have changed the prompt.
22.6 Time, and the Batching Consequence #
Caches expire on inactivity. The lifetime is measured from the start of the request that writes or reads the entry, not from the end of its response, so generation time counts against it.12 A response that takes four minutes to stream leaves roughly one minute for the follow-up to land within a five-minute TTL.
| Provider | Default TTL | Extended option |
|---|---|---|
| Anthropic | 5 minutes, refreshed on every hit | 1 hour at 2× write |
| OpenAI GPT-5.6+ | 30 minutes, refreshed | none (30m is the only value) |
| OpenAI earlier | 24h: ≈30 min typical, 24 h maximum; in_memory (5–10 min typical, 1 h maximum) under ZDR | the other of the two |
| Google explicit | 1 hour default TTL | configurable, with storage charge |
One reading trap sits in that table. OpenAI’s earlier-model values are retention settings, and the documented maximum is not the working lifetime: 24h keeps entries typically available for around thirty minutes and retains them for up to twenty-four hours, while in_memory is typically five to ten minutes against a one-hour ceiling.49 Plan the follow-up request against the typical figure and treat the ceiling as an upper bound on data retention rather than a promise of a warm hit tomorrow. The ZDR-dependent default also applies only to models offering both settings; GPT-5.5 and 5.5 Pro document 24h alone.
This produces a working-style consequence that nobody considers until the bill arrives: four questions in ten minutes ride one cache; the same four questions spread across an afternoon rebuild it each time. Batching related work now carries a cost consequence on top of the focus one.
Pre-warming is available on Anthropic via a zero-output request, which eliminates the cache-miss latency penalty on the first real interaction:12
# Fire at application startup or on a schedule; costs a cache write, zero output.
client.messages.create(
model="claude-opus-5",
max_tokens=0,
system=[{"type": "text", "text": SYSTEM_PROMPT,
"cache_control": {"type": "ephemeral"}}],
messages=[{"role": "user", "content": "warmup"}],
)
Two constraints govern that call. The breakpoint must be on the block shared with the real request, not on the placeholder message, or the entry is keyed to the placeholder and real traffic never hits it. And the thinking configuration and effort setting must match your real requests because those are rendered into the prompt and a mismatched pre-warm writes an entry nothing will read.
22.7 Routing, and Why Hit Rates Drop Under Load #
One under-documented mechanic carries real production consequences, and which generation you are on decides whether you have to act on it. Cached states live on individual machines. OpenAI documents that traffic above roughly 15 requests per minute can overflow to another machine, and a request only hits if it reaches a machine holding a matching, unexpired entry.49
On models before GPT-5.6 the mitigation is yours to apply, and it takes the form of a routing hint. On GPT-5.6 and later OpenAI routes automatically and states that the key is not needed to optimize caching; there it survives as a way to keep cache accounting separate between tenants, prompt versions or environments.49
const response = await client.responses.create({
model: "gpt-5.6-sol",
prompt_cache_key: "support-v3:tenant-acme", // 5.6+: separate accounting.
prompt_cache_options: { mode: "implicit", ttl: "30m" }, // Earlier models:
input, // this is the routing lever.
});
Key design, from the vendor guidance: combine a prompt version with a stable user, workspace, session, or thread identifier that matches how your application actually reuses context. Keep keys stable rather than generating one per request. And on the earlier generation, if a group gets busy and hit rates decline, shard it deterministically so related requests still land together.49
import hashlib
def cache_key(prompt_version: str, tenant: str, session: str, shards: int = 16) -> str:
digest = hashlib.sha256(f"{tenant}:{session}".encode()).hexdigest()
shard = int(digest[:8], 16) % shards
return f"{prompt_version}:{tenant}:shard-{shard}"
On the generation that needs them, keys influence routing without pinning a request to a machine, and they guarantee nothing about a hit.
22.8 Measuring It #
The usage object is the only ground truth, and the important point about Anthropic’s is that input_tokens does not mean total input. The next two blocks are Anthropic-shaped because the field semantics are not portable and §22.8 is the place that bites people.
total_input = cache_read_input_tokens
+ cache_creation_input_tokens
+ input_tokens ← only tokens AFTER the last breakpoint
def cache_stats(usages: list[dict]) -> dict:
"""Anthropic usage only. The three fields are disjoint, so they add."""
reads = sum(u.get("cache_read_input_tokens") or 0 for u in usages)
writes = sum(u.get("cache_creation_input_tokens") or 0 for u in usages)
fresh = sum(u["input_tokens"] for u in usages)
total = reads + writes + fresh
return {
"hit_rate": reads / total if total else 0.0,
"write_ratio": writes / total if total else 0.0,
"total_input": total,
}
That addition is the part that does not travel, and passing another provider’s usage through it returns zeros rather than an error. OpenAI’s input_tokens is the total and includes both cached and cache-written input, with the split reported underneath it in input_tokens_details as cached_tokens and cache_write_tokens; ordinary input is the total minus those two. So a 15,000-token request with 12,000 cached and 3,000 written is an 80 percent hit rate there, and the helper above reports 0.0 because the two Anthropic-specific cache fields are absent, leaving a zero numerator. The field it does find is input_tokens, which exists under both providers, so fresh picks up the whole 15,000 and total_input comes out right while the split is entirely wrong—a wrong answer that passes a sanity check on the total. Were all three names genuinely missing, the direct subscript on input_tokens would raise KeyError rather than report zero. Normalize to one shape in a provider adapter before computing anything, and keep the field names out of your metrics layer:
def normalize(usage: dict, provider: str) -> tuple[int, int, int]:
"""(reads, writes, fresh) — disjoint, whatever the provider called them."""
if provider == "anthropic":
return (usage.get("cache_read_input_tokens") or 0,
usage.get("cache_creation_input_tokens") or 0,
usage.get("input_tokens") or 0)
d = usage.get("input_tokens_details") or {}
reads, writes = d.get("cached_tokens") or 0, d.get("cache_write_tokens") or 0
return reads, writes, (usage.get("input_tokens") or 0) - reads - writes
How to read the result:
| Signal | Diagnosis |
|---|---|
| Hit rate > 0.85 | Working. Agentic workloads reach the low-to-mid nineties. |
| Hit rate 0.5–0.85 | Something invalidates periodically. Causes are in §23. |
| Hit rate < 0.2 after a week | A prefix-design problem, not a model problem. |
| Writes climbing, reads at zero | The §22.3 trap. Your breakpoint is on changing content. |
| Hit rate falls under load | Routing overflow. On the pre-5.6 OpenAI generation, add a cache key (§22.7). |
Anthropic offers a cache diagnostics beta that compares consecutive requests and reports exactly where the prefix diverged, which short-circuits most of this investigation.12
References cited in this section
2 of 81 · numbering matches the PDF
- 12Anthropic, "Prompt Caching," Claude Platform documentation verified September 8, 2026. Vendor documentation; cited as product fact for mechanism, pricing multipliers, minimum cacheable lengths, invalidation behavior, the 20-block lookback window, pre-warming, and data retention. The pricing table in this reference is the primary source for all Anthropic rates quoted in this document.platform.claude.com/docs/en/build-with-claude/prompt-caching ↗
- 49OpenAI, "Prompt caching," OpenAI API documentation verified September 8, 2026. Vendor documentation; cited as product fact for implicit and explicit breakpoint modes, the 1.25× write and 0.1× read multipliers on GPT-5.6 and later, minimum cacheable lengths, TTL and retention semantics, machine-local cache routing and the ~15 requests-per-minute overflow threshold, prompt_cache_key design guidance, the minimum-cacheable-length break-even formula, and the compaction interaction.developers.openai.com/api/docs/guides/prompt-caching ↗