Assembling Context
More context is not better, and it is measured. Signal-to-noise as the metric, retrieval's limits, and a context-budget pattern.
Cover what to put in the window, what to leave out, and why “more context” is one of the most expensive wrong intuitions in the field.
12.1 Why More Is Not Better #
The foundational empirical result for practical prompting is that model performance degrades as input length grows, well before the context window is anywhere near full.
Systematic testing of 18 frontier models established that none of them uses its context uniformly; performance grows increasingly unreliable as input length increases, across every model tested. Critically, the study found significantly higher performance on focused prompts than on full prompts containing the same relevant information plus irrelevant material, meaning that adding irrelevant context, which forces an additional retrieval step, directly degrades reliability.35
Separately, work on positional effects established the U-shaped curve: accuracy is highest for information at the beginning or end of the context and lower in the middle, replicated across six model families, with the magnitude of the drop differing sharply from one model to the next—the worst case exceeds thirty percent relative and the mildest is a few points.20 And benchmark work on effective context length found that models claiming 32K or more frequently fail to maintain performance across anything close to their advertised limit.36
Taken together, these establish that a 200K window states capacity, and that capacity overstates the range in which the model performs reliably. Expect measurable degradation from roughly a third of it, and design accordingly.
12.2 Signal-to-Noise Is the Metric #
What one manages is the ratio of relevant to irrelevant tokens rather than the size of the context.
Ten thousand tokens of exactly the right material will beat a hundred thousand tokens containing the same material plus ninety thousand tokens of near-misses. This is why dumping an entire repository into a window performs worse than supplying four well-chosen files, and why agents that explore broadly before acting often do worse than agents given a narrow starting point.
The operating rule: every token in the window must be there for a reason you could state. If you cannot say why a file is in context, it is noise.
12.3 Retrieval, and Its Measured Limits #
Retrieval is the standard answer to context selection. It is also a genuine failure surface in its own right rather than a solved problem.
Independent benchmark work across 25 repositories reports both halves of the problem, and the population matters for each. On the 287-sample subset with logged agent trajectories, interactive agents never touched any gold file on 27 to 35 percent of samples despite exploration. On the same 287 samples, which also carry line-span annotations, the median labeled evidence is 27 lines, or 4.7 percent of its containing file, and 75.7 percent of span-file pairs use at most a tenth of the file.37 Neither figure is drawn from the full 427-sample benchmark, which also contains no-gold and counterfactual cases. Two consequences follow. First, context acquisition is a distinct and measurable failure mode, independent of whether the model can solve the problem once it holds the right material. Second, file-level retrieval is coarse: a file-level hit shows the document was reachable and says little about locating the useful region inside it, so pulling a whole file to get one relevant function leaves most of what you paid for unaddressed by the annotation. Do not read that as a measured twenty-to-one noise ratio—4.7 percent is a median line fraction of judged evidence, not a token accounting, and the authors warn that useful but unjudged material would otherwise be counted as waste.
The architectural debate between just-in-time agentic exploration and a maintained semantic index has major vendors holding publicly opposite positions, each publishing evidence favoring its own approach. The synthesis the published data supports is scale-dependence: the vendor advocating indexing reports code retention improving 0.3 percent overall but 2.6 percent on codebases over a thousand files.38 Exploration is fine below that scale, and above it an index starts paying for itself.
12.4 Practical Selection #
At the keyboard, in order of value:
Name the files. If you know which files matter, say so. An agent given three specific paths outperforms an agent told to find them because you have eliminated the failure mode in §12.3 entirely.
Read src/auth/session.py and tests/auth/test_session.py.
Do not read anything else unless those two files reference it directly.
Give the interface, not the implementation. Type signatures, schemas, and function stubs carry most of the information at a fraction of the tokens. A 3,000-token module reduces to a 300-token interface summary that answers most questions about how to call it.
Summarize rather than paste, for background. A file the model needs to know exists but not read in detail belongs in one line, not two thousand tokens.
Order by relevance, with the most important adjacent to the question. Directly follows from §12.1’s positional findings.
Prune between tasks. Files loaded for task A are noise for task B. Section 24 covers the mechanics.
12.5 A Context-Budget Pattern #
Concretely, for anyone assembling context programmatically, the pattern runs as follows:
from dataclasses import dataclass
@dataclass
class ContextItem:
content: str
tokens: int
relevance: float # 0..1, from your retriever or heuristics
pinned: bool = False # always include (e.g. the task spec)
def assemble(items: list[ContextItem], budget: int) -> list[ContextItem]:
"""Fill a budget by relevance, keeping pinned items, and place the
highest-relevance item LAST so it sits adjacent to the question (§8.5)."""
pinned = [i for i in items if i.pinned]
rest = sorted((i for i in items if not i.pinned),
key=lambda i: i.relevance, reverse=True)
used, chosen = sum(i.tokens for i in pinned), []
for item in rest:
if used + item.tokens > budget:
continue # skip, don't stop — smaller items may fit
chosen.append(item)
used += item.tokens
# Pinned first (stable, cacheable), then ascending relevance so the
# strongest signal lands closest to the user message.
return pinned + sorted(chosen, key=lambda i: i.relevance)
Two choices in that loop were made on purpose. The loop uses continue rather than break, so a single large high-priority item does not block several smaller ones behind it. And the final ordering is ascending relevance, which puts the best material last—the opposite of the naive ordering, and correct for the reasons in §8.5.
budget bounds the fill, not the total: pinned items go in unconditionally, so pinned content that is itself oversized will exceed it. Set budget inside §29.5’s reliable range for the model you are on—roughly a quarter to two-fifths of a 128K or 200K window, and a tenth to a fifth of a 1M one—rather than at ninety-five percent of it. The remainder absorbs tool results and conversation growth, and you avoid operating in the degraded region.
12.6 Failure Modes #
- Filling the window because it is available. Capacity is not a target.
- Whole-file retrieval for a single function. A file-level hit shows the document was reachable, not that the relevant evidence was localized.37
- Assuming retrieval succeeded. It fails silently on a quarter to a third of hard samples.37
- Burying the question in the middle. The worst position in the window.20
- Never pruning. Context accumulates by default; removal is always an explicit act.
References cited in this section
5 of 81 · numbering matches the PDF
- 35Kelly Hong, Anton Troynikov, and Jeff Huber, "Context Rot: How Increasing Input Tokens Impacts LLM Performance," Chroma Research, 2025 Vendor-published research; the publisher sells vector databases and has an interest in the conclusion, which should be stated. Tests 18 frontier models and finds that none use context uniformly, that reliability degrades with input length, and that focused prompts substantially outperform full prompts containing the same relevant material plus distractors. The finding is consistent with independent work on position effects and on effective context length (references 20 and 36) and is the empirical basis for compaction, sub-agent isolation, and aggressive curation.www.trychroma.com/research/context-rot ↗
- 20Nelson F. Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni, and Percy Liang, "Lost in the Middle: How Language Models Use Long Contexts," Transactions of the ACL, 2024. Peer-reviewed. Establishes the U-shaped positional accuracy curve for mid-context material, replicated across six model families and confirmed on additional architectures since. The shape is the cross-model finding; the magnitude is not. In the thirty-document multi-document QA table the steepest edge-to-middle decline exceeds thirty percent relative—a fall of roughly 22 percentage points, from about 73 to about 51—while other models in the same table decline by single-digit percentages. An earlier revision of this document reported the greater-than-thirty-percent figure as a cross-model result; it is a maximum observed effect, and the entry now says which unit it is in, since a relative percent and a percentage point are not the same quantity. The paper reports task accuracy by position and does not publish an attention-weight profile (see reference 65).
- 36Cheng-Ping Hsieh, Simeng Sun, Samuel Kriman, Shantanu Acharya, Dima Rekesh, Fei Jia, Yang Zhang, and Boris Ginsburg, "RULER: What's the Real Context Size of Your Long-Context Language Models?," arXiv:2404.06654, 2024. Establishes that models claiming 32K or more frequently fail to maintain performance across their advertised range, with failure modes including a failure to ignore distractors and reverting to parametric knowledge. The distractor mode is easy to write backward, and an earlier revision of this document did: ignoring distractors is the desired behavior, and what the paper reports is models incorrectly retrieving values associated with the distractor keys. Also relevant: Ali Modarressi et al., "NoLiMa: Long-Context Evaluation Beyond Literal Matching," 2025, which demonstrates that non-lexical matching degrades sharply with length.arxiv.org/abs/2404.06654 ↗
- 37Bowen Qin and Yi Xie, "Agent Retrieval Bench: Evaluating Repository Context Retrieval for Coding Agents," arXiv:2607.24882, July 2026. Preprint. 427 samples across 25 repositories: 345 positive retrieval examples, 50 natural no-gold cases, and 32 counterfactual wrong-repository controls. Two populations inside it must be kept apart, and an earlier revision of this document attached both figures below to the 427. Both are drawn from the 287-sample code2test / comment2context / trace2code subset: the logged-trajectory result that interactive agents never touched any gold file on 27 to 35 percent of samples (35.2 percent for the OpenAI strict-context agent, 27.2 to 29.3 percent for Codex), and the span annotations whose median labeled evidence is 27 lines, or 4.7 percent of its containing file, with 75.7 percent of span-file pairs using at most a tenth of the file. The 4.7 percent is a median line fraction over 391 span-file pairs, not a token-based signal-to-noise ratio, and the paper explicitly declines to treat unjudged content as waste, so it does not support an inference that the rest of the file is irrelevant. The paper is also explicit that it measures document reachability rather than within-file localization, and that it does not establish that a retrieval miss causes a repair failure.arxiv.org/abs/2607.24882 ↗
- 38Vendor-published comparison of just-in-time agentic exploration against maintained semantic indexing, reporting code retention improving 0.3 percent overall and 2.6 percent on codebases over a thousand files. Two major vendors hold publicly opposite positions on this architecture and each publishes evidence favoring its own; scale-dependence is the only synthesis the evidence supports. Cited via reference 1.