Attention, KV State, and the Geometry of a Context Window
The KV cache from the inside, why position and length both matter, attention sinks, and effective against advertised context.
Explain the architectural reasons behind the positional and length effects used throughout this document, so the practices follow from mechanism rather than from folklore.
29.1 The KV Cache, From the Inside #
At each layer, a transformer computes query, key, and value projections for every token. Attention weights come from query-key dot products; the output is a weighted sum of values.
During generation, each new token attends to every token preceding it. Recomputing all previous keys and values at each step would be quadratic and wasteful, so they are cached—that is the KV cache. It grows linearly with sequence length, and it is the memory term that varies with the workload: weights are fixed, and KV is what a long conversation adds on top of them. Which term dominates depends on the regime rather than being settled in KV’s favor—70B weights at fp16 run to about 140 GB against the 32.8 GB of KV computed below, and the 4-bit 35B example in §37.3 puts 17.5 GB of weights against 6.4 GB of KV.
Rough size, for intuition:
KV bytes ≈ 2 × layers × kv_heads × head_dim × seq_len × dtype_bytes
A 70B-class model, 80 layers, 8 KV heads (grouped-query attention),
head_dim 128, fp16, 100K tokens:
2 × 80 × 8 × 128 × 100,000 × 2 bytes ≈ 32.8 GB
That figure is conventional full-attention or grouped-query arithmetic rather than a universal law: hybrid, linear-attention and sliding-window designs do not all hold full-length state at every layer, and the Qwen3.6 example in §30.5 is itself a mix of gated-linear and attention layers.
That is thirty-three gigabytes of KV state for a single hundred-thousand-token conversation. That number is why prompt caching exists at all, why providers charge for cache writes, and why cached entries have short TTLs—they occupy scarce GPU memory.
It is also why cache entries are machine-local (§22.7). The state lives in the memory of a specific accelerator; a request routed elsewhere cannot see it.
29.2 Why Position Matters: RoPE and Softmax #
Two mechanisms are the leading explanation for the U-shaped accuracy curve upon which Section 12 relies. Keep the distinction between them and it: what the benchmarks measure is task accuracy by position, and what follows is an interpretation of that measurement rather than a derivation of it.
Rotary position embeddings decay with distance. RoPE encodes position by rotating query and key vectors by an angle proportional to position. A useful property of the construction is long-term decay: the expected dot-product similarity between token pairs decreases as their separation grows.65 Distant tokens are systematically harder to attend to, before any content is considered.
Softmax concentrates. Attention weights are a softmax over scores. Softmax is winner-take-most: it amplifies the highest scores and suppresses the rest. Combined with the positional decay, this reinforces primacy and recency and suppresses the middle.
The empirical result is well replicated. Accuracy is highest for information at the beginning or end of the context and lower when the same information sits in the middle, across six model families and confirmed on additional architectures since.20 What replicates is the shape; the magnitude does not. In the paper’s thirty-document setting the steepest edge-to-middle decline exceeds thirty percent in relative terms, while other models in the same table drop by single digits. The unit matters too: a fall from roughly 73 to roughly 51 is about 22 percentage points and about 31 percent relative, and the two are routinely quoted interchangeably. RoPE’s distance decay is a published property of the construction;65 that it is the whole cause of this curve, primacy included, is not established, and neither paper reports a universal attention-weight profile.
Two design consequences already invoked earlier in this document now have their mechanism: put the most important material at the boundaries (§8.5, §12.5), and treat mid-context placement as active degradation rather than neutral storage.
29.3 Why Length Matters Independent of Position #
Position, however, is not the whole story. Even material at a favorable position degrades as total length grows, and the reason is distributional.
As sequence length increases, the softmax must spread probability mass across more candidates. Each individual token receives less attention on average, and near-miss distractors, content that is semantically similar to the target but wrong, compete more effectively. Systematic testing across 18 frontier models found that performance on focused prompts substantially exceeded performance on full prompts containing the same relevant material plus irrelevant context, and that the gap widened with length.35
This is the mechanism behind the signal-to-noise framing in §12.2. You are not fighting a capacity limit but attention dilution, and the way to win is to remove competitors rather than add capacity.
The related benchmark finding: models advertising 32K or more frequently fail to maintain performance across anything close to their stated limit, with failure modes including a failure to ignore distractors—returning the value attached to an irrelevant key rather than the one asked for—and reverting to parametric knowledge instead of the provided context.36 That last failure mode is the dangerous one because a model answering from training data instead of your document produces a fluent, plausible, unsourced answer.
29.4 Attention Sinks #
One observed phenomenon carries a practical implication: transformers allocate disproportionate attention to the first few tokens of a sequence regardless of content. The prevailing explanation is that softmax must sum to one, so when no position is a good match the model needs somewhere to dump probability mass, and the earliest positions serve as a sink.
The consequence is about eviction rather than about authority. Streaming implementations that drop early tokens degrade sharply unless the sink positions are retained, which is why sink-aware sliding-window designs keep the first few tokens permanently—that retention is the fix those designs introduced, not a property every windowed implementation has.79
What does not follow is that a substantive instruction gains processing by occupying those positions. The reported behavior tracks absolute position rather than semantic content: replacing the initial tokens with newlines preserves the effect.79 The case for putting authoritative material at the top rests on the positional-accuracy findings in §29.2, not on the sink.
29.5 Effective Context Versus Advertised Context #
The gap between these two figures governs everything practical in this section.
| Advertised | Reliable working range (typical) |
|---|---|
| 128K | 30–50K |
| 200K | 50–80K |
| 1M | 100–200K |
These are working estimates rather than published constants, and they vary by task and by model. The pattern is consistent across benchmark work: effective context is a fraction of advertised context, and the fraction shrinks as the task requires disambiguation among distractors rather than simple retrieval.35,36
Design to the second column, and not to a fixed percentage of the first. The fraction is not constant: those ranges are roughly a quarter to two-fifths of a 128K or 200K window and only a tenth to a fifth of a 1M one, because the shrinkage compounds as the window grows. Forty to sixty percent of the advertised figure—a rule this document used to give—lands outside every row in the table.
29.6 What This Says About Compaction #
Section 24 argued for compaction and resets on empirical grounds, and the architecture now supplies the mechanism.
Compaction accomplishes two things simultaneously. It reduces length, which reduces attention dilution. And it moves surviving content closer to the boundaries, which improves its positional weight. A well-executed compaction is not merely a cost operation—it structurally improves the model’s access to the material it retains.
The counterweight, and the reason compaction is not free, is this: the summarizer chooses what survives, and its judgment is the model’s judgment, subject to the same failure modes. Anything that must survive should not be entrusted to it (§24.3).
References cited in this section
5 of 81 · numbering matches the PDF
- 65Jianlin Su, Murtadha Ahmed, Yu Lu, Shengfeng Pan, Wen Bo, and Yunfeng Liu, "RoFormer: Enhanced Transformer with Rotary Position Embedding," Neurocomputing 568 (2024), arXiv:2104.09864. Peer-reviewed. Source of the long-term decay property of rotary position embeddings: reduced expected dot-product similarity between distant token pairs. That is a property of the construction and the leading mechanistic account of the positional findings in reference 20; it is not a proof that RoPE decay causes those findings, it says nothing about primacy, and neither paper publishes a universal attention-weight profile by position. §29.2 is scoped accordingly.arxiv.org/abs/2104.09864 ↗
- 20Nelson F. Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni, and Percy Liang, "Lost in the Middle: How Language Models Use Long Contexts," Transactions of the ACL, 2024. Peer-reviewed. Establishes the U-shaped positional accuracy curve for mid-context material, replicated across six model families and confirmed on additional architectures since. The shape is the cross-model finding; the magnitude is not. In the thirty-document multi-document QA table the steepest edge-to-middle decline exceeds thirty percent relative—a fall of roughly 22 percentage points, from about 73 to about 51—while other models in the same table decline by single-digit percentages. An earlier revision of this document reported the greater-than-thirty-percent figure as a cross-model result; it is a maximum observed effect, and the entry now says which unit it is in, since a relative percent and a percentage point are not the same quantity. The paper reports task accuracy by position and does not publish an attention-weight profile (see reference 65).
- 35Kelly Hong, Anton Troynikov, and Jeff Huber, "Context Rot: How Increasing Input Tokens Impacts LLM Performance," Chroma Research, 2025 Vendor-published research; the publisher sells vector databases and has an interest in the conclusion, which should be stated. Tests 18 frontier models and finds that none use context uniformly, that reliability degrades with input length, and that focused prompts substantially outperform full prompts containing the same relevant material plus distractors. The finding is consistent with independent work on position effects and on effective context length (references 20 and 36) and is the empirical basis for compaction, sub-agent isolation, and aggressive curation.www.trychroma.com/research/context-rot ↗
- 36Cheng-Ping Hsieh, Simeng Sun, Samuel Kriman, Shantanu Acharya, Dima Rekesh, Fei Jia, Yang Zhang, and Boris Ginsburg, "RULER: What's the Real Context Size of Your Long-Context Language Models?," arXiv:2404.06654, 2024. Establishes that models claiming 32K or more frequently fail to maintain performance across their advertised range, with failure modes including a failure to ignore distractors and reverting to parametric knowledge. The distractor mode is easy to write backward, and an earlier revision of this document did: ignoring distractors is the desired behavior, and what the paper reports is models incorrectly retrieving values associated with the distractor keys. Also relevant: Ali Modarressi et al., "NoLiMa: Long-Context Evaluation Beyond Literal Matching," 2025, which demonstrates that non-lexical matching degrades sharply with length.arxiv.org/abs/2404.06654 ↗
- 79Guangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han, and Mike Lewis, "Efficient Streaming Language Models with Attention Sinks," ICLR 2024, arXiv:2309.17453. Peer-reviewed. Source of the attention-sink phenomenon in §29.4 and, importantly, of its limit: the paper attributes the effect to absolute position rather than to semantic content, reports that replacing the initial tokens with newlines preserves it, and introduces sink retention as its own proposed fix for window attention rather than describing existing practice. It therefore supports the eviction argument and does not support any inference that a substantive instruction gains processing by occupying those positions.arxiv.org/abs/2309.17453 ↗