The Four Token Lanes
Four lanes, a fifty-to-one spread, the context threshold cliff, and why output is where the money goes.
Establish the billing model precisely because every optimization decision in Parts III and IV reduces to moving tokens between these four lanes.
21.1 The Lanes #
Every billed token falls into exactly one of four categories, and on a single model the cheapest and the dearest differ by a factor of fifty or more.
| Lane | What it is | Rate vs. base input |
|---|---|---|
| Cached read | A prefix that matched and was reused | 0.1× (0.025× or 0.05× on some models) |
| Base input | Everything after the last cache hit, and every outright miss | 1.0× |
| Cache write | Storing a prefix for the first time | 1.25× (5-minute) or 2× (1-hour) |
| Output | Everything generated, including reasoning tokens | 5× (6× on some models) |
Worked through on Claude Sonnet 5, verified September 8, 2026:26
| Lane | Price / MTok | Relative |
|---|---|---|
| Cached read | $0.20 | 1× |
| Base input | $2.00 | 10× |
| 5-minute cache write | $2.50 | 12.5× |
| 1-hour cache write | $4.00 | 20× |
| Output | $10.00 | 50× |
The spread between the cheapest and most expensive lane runs fifty to one on Sonnet 5, and two hundred to one on Fable 5.1, whose cached reads bill at a fortieth of base input rather than a tenth (§38.3). That spread is the entire reason token economics warrants a chapter.
21.2 Cross-Provider Comparison #
The lanes are universal, though multipliers and mechanisms vary by provider.
| Anthropic | OpenAI (GPT-5.6+) | OpenAI (earlier) | Google Gemini | |
|---|---|---|---|---|
| Cache read | 0.1× | 0.1× | model-dependent | 0.1× |
| Cache write | 1.25× (5m) / 2× (1h) | 1.25× | none | none (implicit) |
| Explicit control | cache_control breakpoints | prompt_cache_breakpoint | not supported | CachedContent object |
| Automatic mode | Yes, top-level cache_control | Yes, implicit mode | Yes, only mode | Yes, implicit |
| Minimum prefix | 512–4,096 tokens, model-dependent | 1,024 | Varies by request settings | 2,048–4,096, model-dependent |
| Default TTL | 5 minutes, refreshed on hit | 30 minutes | 24h (≈30 min typical), or in_memory under ZDR | 1 hour (explicit) |
| Storage charge | None | None | None | Yes, per token per hour |
Sources: Anthropic prompt caching documentation,12 OpenAI prompt caching documentation,49 Google context caching documentation.50
Three differences deserve particular emphasis.
Cached content bills per token per hour that the cache exists, whether or not it is read—$0.50 per million tokens per hour on Gemini 3.6, 3.7 and 3.8 Flash through December 31, 2026, rising to $1.00 on January 1, 2027, against $1.00 already on 3.5 Flash and the Flash-Lite models and $4.50 on Gemini 3.1 Pro.51 This changes the arithmetic completely: an unused cache keeps billing, where on every other provider it would simply be a sunk cost. Google’s implicit caching has no storage charge and is the right default; explicit caching is for large hot contexts where you want the discount guaranteed.
Reads on 5.6+ moved to 0.1×, which more than compensates. OpenAI states the arithmetic directly: one write plus one full read costs 1.35× ordinary input, against 2× for processing twice uncached; across ten requests, one write and nine reads cost 2.15× against 10×.49
The break-even is under two reads at the 5-minute TTL. A session that makes three or more calls against the same prefix is already ahead.
21.3 The Context Threshold Cliff #
Several models change rate outright above a context threshold. A context threshold is the input-token count at which a provider moves a request onto a second, higher rate, assessed per request against the whole rendered input—system prompt, tool schemas, history, and your message together. Crossing it applies the higher rate to every token in that request, not only to those above the line. This is a cliff rather than a slope, and no amount of caching will carry one back across it.
Verified September 24, 2026:52,51
| Model | Short context | Long context | Threshold |
|---|---|---|---|
| gpt-6-astra | $10.00 / $50.00 | $20.00 / $75.00 | above 272K tokens |
| gpt-5.6-sol | $4.00 / $20.00 | $8.00 / $30.00 | above 272K tokens |
| gpt-5.6-terra | $2.00 / $12.00 | $4.00 / $18.00 | above 272K tokens |
| Gemini 3.1 Pro | $2.00 / $12.00 | $4.00 / $18.00 | above 200K tokens |
One of those rows is promotional rather than standing: GPT-5.6 Sol’s $4.00/$20.00 is documented as available at least through November 21, 2026, so re-check it before building a forecast on it.52
Anthropic’s current generation prices 1M-token context at flat rates without a surcharge on Opus 5, Fable 5.1, and Sonnet 5,26 which is a genuine differentiator when comparing.
On providers carrying a threshold, a session that drifts above it pays 2× on input, cached reads and cache writes, and 1.5× on output, applied to the whole request. A single unnecessary 60K-token file read can push a session that was already near the line across it. The repricing is assessed per request rather than latched for the session, so compaction, pruning or a reset brings the next request back under (§24), where caching alone will not. Section 25.3 covers how to alarm on the crossing.
21.4 Output Is Where the Money Goes #
Output is the dearest lane on every major provider, typically five times base input and sometimes six—the rate cards in this section put GPT-5.6 Terra, GPT-5.6 Luna and Gemini 3.1 Pro’s short-context tier at six, and the long-context tiers lower the ratio again. What holds universally is the ordering, not the multiplier. Caching never discounts output either way: generation bills in full every time, and reasoning tokens bill here.
This inverts the intuition most practitioners bring from the input side. A session with a 94 percent cache hit rate and heavy reasoning can have output as its dominant cost line despite output being under a tenth of the token volume.
def session_cost(usage, rates):
"""rates: dict with base_in, cache_read, cache_write_5m, out — $/MTok."""
return (
usage["cache_read_input_tokens"] * rates["cache_read"] / 1e6 +
usage["cache_creation_input_tokens"] * rates["cache_write_5m"] / 1e6 +
usage["input_tokens"] * rates["base_in"] / 1e6 +
usage["output_tokens"] * rates["out"] / 1e6
)
SONNET_5 = {"base_in": 2.00, "cache_read": 0.20,
"cache_write_5m": 2.50, "out": 10.00} # verified 2026-09-08
# A realistic 40-turn agentic session
usage = {
"cache_read_input_tokens": 1_850_000, # 94% hit rate
"cache_creation_input_tokens": 92_000,
"input_tokens": 34_000,
"output_tokens": 118_000, # includes reasoning
}
print(f"${session_cost(usage, SONNET_5):.2f}")
# cached reads $0.37
# cache writes $0.23
# base input $0.07
# output $1.18 ← 64% of the bill on 6% of the tokens
# total $1.85
The only lever available on output is to ask for less. Prompt discipline, tighter output contracts, and a calibrated thinking budget all reduce it, and caching does none of that work.
21.5 What the Meter Does Not Show #
Tokens are the visible lane, and also the smallest one. Total cost of assisted delivery includes human time spent prompting, reviewing, verifying, and correcting; compute for pipelines and runners; licenses; and the engineering effort spent building the instructions, skills, and evaluations that make any of it work.
At the assisted level, human time typically dominates the total. A program sold on token savings is promising against the smallest term on the sheet and will underdeliver against its business case.
The measurement that puts quality inside the economics rather than beside them: divide total cost by outcomes accepted, not outcomes attempted. If acceptance runs at 60 percent instead of 100, cost per accepted outcome rises roughly 1.7× with nothing about the work having changed. That framing belongs to the governance side and is developed properly elsewhere;1 it is named here so that nobody leaves this chapter believing the token bill is the cost.
References cited in this section
7 of 81 · numbering matches the PDF
- 26Anthropic model pricing, published in the pricing table of reference 12 and verified September 8, 2026. Source for all Claude per-model rates: Fable 5.1 $10/$50, Opus 5 $5/$25, Sonnet 5 $2/$10, Sonnet 4.6 $3/$15, Haiku 4.5 $1/$5 per MTok, with cache multipliers of 1.25× (5m write), 2× (1h write), and 0.1× read (0.025× on Fable 5.1 and Mythos 5.1).platform.claude.com/docs/en/about-claude/pricing ↗
- 12Anthropic, "Prompt Caching," Claude Platform documentation verified September 8, 2026. Vendor documentation; cited as product fact for mechanism, pricing multipliers, minimum cacheable lengths, invalidation behavior, the 20-block lookback window, pre-warming, and data retention. The pricing table in this reference is the primary source for all Anthropic rates quoted in this document.platform.claude.com/docs/en/build-with-claude/prompt-caching ↗
- 49OpenAI, "Prompt caching," OpenAI API documentation verified September 8, 2026. Vendor documentation; cited as product fact for implicit and explicit breakpoint modes, the 1.25× write and 0.1× read multipliers on GPT-5.6 and later, minimum cacheable lengths, TTL and retention semantics, machine-local cache routing and the ~15 requests-per-minute overflow threshold, prompt_cache_key design guidance, the minimum-cacheable-length break-even formula, and the compaction interaction.developers.openai.com/api/docs/guides/prompt-caching ↗
- 50Google, "Context caching," Gemini API documentation and "Context caching overview," Gemini Enterprise Agent Platform documentation. Vendor documentation; cited as product fact for the implicit/explicit distinction, per-model minimum token counts, the 90 percent discount on Gemini 2.5 and later (75 percent on 2.0), the default one-hour TTL on explicit caches, and the statement that storage costs apply to explicit caching only. The minimums are model-specific rather than tier-specific and were re-checked on the vendor page on September 25, 2026: 2,048 on Gemini 2.5 Flash and 2.5 Pro, 4,096 on the 3.x models including Gemini 3.1 Pro.ai.google.dev/gemini-api/docs/caching ↗
- 51Google Gemini API pricing Vendor documentation. Per-token rates verified against Google's own page on September 24, 2026: Gemini 3.1 Pro at $2.00 input, $0.20 cached input and $12.00 output for prompts at or below 200K tokens, and $4.00, $0.40 and $18.00 above it. The Pro storage rate of $4.50 per million tokens per hour was verified against the vendor page on September 25, 2026, along with its Batch table, whose cached-input row reads "same as Standard" and whose storage rate is unchanged—so for that model batching and caching do not compound. That does not generalize, and an earlier revision of this document wrongly rejected an audit finding which said so. Verified on the same page on September 25, 2026: Gemini 3.6/3.7/3.8 Flash Batch halves cached input ($0.075 to $0.0375 in the promotional period), 3.5 Flash Batch halves it ($0.15 to $0.075), and 3.1 Flash-Lite Batch halves both cached input ($0.025 to $0.0125) and storage ($1.00 to $0.50 per million tokens per hour). Batch and cache interaction is a per-model, per-tier fact. Cache storage was re-verified per model and per tier against the same page on September 26, 2026, and it varies along both axes, which is why no single "Flash-tier storage" figure exists. Gemini 3.6, 3.7 and 3.8 Flash are $0.50 per million tokens per hour through December 31, 2026, rising to $1.00 on January 1, 2027, and that promotional row is identical on Standard and Batch—for these models batching does not discount storage, only cached input. Gemini 3.5 Flash is a flat $1.00 on both tiers with no promotional row anywhere in its pricing. Gemini 3.1 Flash-Lite is $1.00 on Standard and $0.50 on Batch, with no promotional row, so here the halving is a tier effect rather than a promotion. Two earlier revisions of this entry each got one axis wrong: one called the $0.50 promotional rate "Flash-tier storage," which collapses the per-model axis and is contradicted by 3.5 Flash; the other carried $1.00 for the 3.6/3.7/3.8 generation from third-party summaries, which turned out to be the post-promotion rate. Read a storage rate off the row for the exact model and tier, and re-read it after January 1, 2027.ai.google.dev/gemini-api/docs/pricing ↗
- 52OpenAI, "Pricing," OpenAI API documentation verified September 8, 2026, with the per-model pages under https://developers.openai.com/api/docs/models/ verified September 24, 2026. Vendor documentation; source for all OpenAI per-model rates including the short-context and long-context tiers. The 272K-token threshold and the 2× input, 2× cache and 1.5× output multipliers are stated on the model pages rather than on the pricing page, which lists the two tiers without naming the boundary. One rate in §21.3's table is promotional rather than standing: the gpt-5.6-sol model page documents its $4.00/$20.00 as available at least through November 21, 2026, verified September 25, 2026, and describes it as a reduction against the prior generation. Re-check it before carrying it into a forecast.developers.openai.com/api/docs/pricing ↗
- 1Joshua Davis, The AI SDLC: An Operating Model, Control Framework, and Maturity Progression for Engineering Organizations Building With Agents, v1.0, September 2026 The companion framework covering governance, controls, and organizational absorption. Cited here for scope boundaries rather than for evidence.jdav.is/ai-sdlc ↗