Gemini
Gemini: two caching mechanisms, the storage meter that changes the arithmetic, hit reliability, and pricing notes.
Cover Google’s context caching model, which is architecturally different from the others in a way that changes the arithmetic.
42.1 Two Caching Mechanisms #
Gemini distinguishes implicit from explicit caching, and the two carry different cost structures.50
Implicit caching is on by default for Gemini 2.5 and newer. When a request reuses a prefix, the discount is passed on automatically. No storage charge, no configuration, no guarantee. Minimum token counts are model-specific rather than tier-specific: 2,048 on Gemini 2.5 Flash and 2.5 Pro, and 4,096 on the 3.x models, Gemini 3.1 Pro included.
Explicit caching is a managed object. You create a CachedContent, reference it by resource name, set a TTL, and delete it when done. It guarantees the discount, 90 percent on Gemini 2.5 and later, and it adds a storage charge.
from google import genai
from google.genai import types
client = genai.Client()
cache = client.caches.create(
model="gemini-3.1-pro-preview",
config=types.CreateCachedContentConfig(
display_name="codebase-context-v3",
system_instruction="You are a code reviewer for this repository.",
contents=[large_document],
ttl="3600s", # storage meter runs for this whole hour
),
)
response = client.models.generate_content(
model="gemini-3.1-pro-preview",
contents="Review the diff below.",
config=types.GenerateContentConfig(cached_content=cache.name),
)
client.caches.delete(name=cache.name) # stop the meter — do not skip this
That final line is what separates Gemini from the other providers. Elsewhere a cache is an optimization you can forget about; here it is an allocated resource with a lifecycle you own. The next subsection explains why forgetting it costs money.
42.2 The Storage Meter Changes Everything #
Explicit caching bills per token per hour that the cache exists, whether or not it is read. The rate is per model and per service tier rather than per family: $0.50 per million tokens per hour on Gemini 3.6, 3.7 and 3.8 Flash through December 31, 2026 and $1.00 from January 1, 2027; $1.00 already on 3.5 Flash and the Flash-Lite models at Standard, though 3.1 Flash-Lite bills $0.50 under Batch; and $4.50 on Gemini 3.1 Pro.51
Worked, on Gemini 3.1 Pro at $2.00/MTok input and $4.50/MTok/hour storage:
100K-token context held for one hour:
storage: 0.1M × $4.50 = $0.45/hour
saving per read: 0.1M × ($2.00 − $0.20) = $0.18
Break-even: 0.45 / 0.18 = 2.5 reads per hour.
Both terms scale with the prefix, so the crossover is the storage rate over the per-token saving and holds at any size inside the tier. Size still decides which tier you are in: above the 200K threshold in §42.4, input and cached input both double while the storage meter does not, and the crossover halves to 1.25 reads an hour. Keep the prefix clear of the line because a 200K cached context plus a question is already over it.
Below roughly three reads an hour, explicit caching loses money. On every other provider a cold cache is a sunk cost; on Gemini explicit caching it is a running meter.
The operating rules follow:
- Default to implicit caching. No meter, no management, no downside.
- Use explicit caching for large, hot contexts where the read rate clearly exceeds break-even and you want the discount guaranteed.
- Delete caches when the job finishes. A forgotten cache is a standing charge. Set short TTLs and treat cache lifecycle as owned resource management.
42.3 Hit Reliability #
Implicit caching is best-effort and controlled upstream. Practitioner reporting suggests hit rates trail the explicit-control providers, which means cost estimates for Gemini workloads should be built at the uncached price with implicit caching treated as upside rather than as budget.
To improve implicit hit rates: put large common content at the beginning of the prompt, and send requests with similar prefixes close together in time.50 Both are the same advice as everywhere else in this document; the difference is that you have no lever beyond it.
42.4 Pricing Notes #
Cached input costs 10 percent of fresh input on Gemini 2.5 and later; the discount was 75 percent on 2.0-generation models. Gemini 3.1 Pro carries a context threshold at 200K tokens, above which input and cached input double and output rises by half—the same multiplier structure as OpenAI’s long-context tier, which begins at 272K (§21.3). Batch halves input and output. Whether it also discounts cached input is model-specific, and the pricing page is the only way to settle it: on Gemini 3.1 Pro the Batch cached-input row reads “same as Standard,” so batching and explicit caching do not compound there, while the 3.6, 3.7 and 3.8 Flash Batch tables halve cached input alongside everything else, and 3.1 Flash-Lite halves the storage meter too.51 Read the row for the model you are billing, not the pattern from another one.
42.5 What Is Distinctive #
The explicit-cache-as-a-managed-object model is simultaneously the most controllable and the most dangerous. A cache write that is never read costs more than ordinary input on Anthropic and on GPT-5.6 and later as well (§21.2), and there it is a one-time premium; here it is a standing charge that accrues until the object is deleted. Gemini is the only provider that meters a cache by the hour, and the only one that requires cache lifecycle management as an operational concern. Teams migrating from Anthropic or OpenAI frequently carry over the assumption that caching is free upside and are surprised.
The 1M-token context on the Pro tier is large and useful for document-scale work—subject to every caveat in §29.5 about effective versus advertised context.
References cited in this section
2 of 81 · numbering matches the PDF
- 50Google, "Context caching," Gemini API documentation and "Context caching overview," Gemini Enterprise Agent Platform documentation. Vendor documentation; cited as product fact for the implicit/explicit distinction, per-model minimum token counts, the 90 percent discount on Gemini 2.5 and later (75 percent on 2.0), the default one-hour TTL on explicit caches, and the statement that storage costs apply to explicit caching only. The minimums are model-specific rather than tier-specific and were re-checked on the vendor page on September 25, 2026: 2,048 on Gemini 2.5 Flash and 2.5 Pro, 4,096 on the 3.x models including Gemini 3.1 Pro.ai.google.dev/gemini-api/docs/caching ↗
- 51Google Gemini API pricing Vendor documentation. Per-token rates verified against Google's own page on September 24, 2026: Gemini 3.1 Pro at $2.00 input, $0.20 cached input and $12.00 output for prompts at or below 200K tokens, and $4.00, $0.40 and $18.00 above it. The Pro storage rate of $4.50 per million tokens per hour was verified against the vendor page on September 25, 2026, along with its Batch table, whose cached-input row reads "same as Standard" and whose storage rate is unchanged—so for that model batching and caching do not compound. That does not generalize, and an earlier revision of this document wrongly rejected an audit finding which said so. Verified on the same page on September 25, 2026: Gemini 3.6/3.7/3.8 Flash Batch halves cached input ($0.075 to $0.0375 in the promotional period), 3.5 Flash Batch halves it ($0.15 to $0.075), and 3.1 Flash-Lite Batch halves both cached input ($0.025 to $0.0125) and storage ($1.00 to $0.50 per million tokens per hour). Batch and cache interaction is a per-model, per-tier fact. Cache storage was re-verified per model and per tier against the same page on September 26, 2026, and it varies along both axes, which is why no single "Flash-tier storage" figure exists. Gemini 3.6, 3.7 and 3.8 Flash are $0.50 per million tokens per hour through December 31, 2026, rising to $1.00 on January 1, 2027, and that promotional row is identical on Standard and Batch—for these models batching does not discount storage, only cached input. Gemini 3.5 Flash is a flat $1.00 on both tiers with no promotional row anywhere in its pricing. Gemini 3.1 Flash-Lite is $1.00 on Standard and $0.50 on Batch, with no promotional row, so here the halving is a tier effect rather than a promotion. Two earlier revisions of this entry each got one axis wrong: one called the $0.50 promotional rate "Flash-tier storage," which collapses the per-model axis and is contradicted by 3.5 Flash; the other carried $1.00 for the 3.6/3.7/3.8 generation from third-party summaries, which turned out to be the post-promotion rate. Read a storage rate off the row for the exact model and tier, and re-read it after January 1, 2027.ai.google.dev/gemini-api/docs/pricing ↗