Section 45 of 45 49 min read

References

Sixty-six sources, each stating what kind of source it is, with vendor-published, correlational and unreviewed material marked as such.

Every entry states what kind of source it is. Where a source is vendor-published, self-reported, correlational, or an unreviewed preprint, the entry says so. Rates and product behavior were verified on September 8, 2026, except where an entry names a later date, and move quickly.

  1. 1Joshua Davis, The AI SDLC: An Operating Model, Control Framework, and Maturity Progression for Engineering Organizations Building With Agents, v1.0, September 2026. The companion framework covering governance, controls, and organizational absorption. Cited here for scope boundaries rather than for evidence.
  2. 2OpenAI, “Text generation” and “Counting tokens,” OpenAI API documentation, verified September 25, 2026. Vendor documentation; cited as product fact for the Responses API’s instructions and input parameters and its developer role, which the documentation prioritizes ahead of user messages, and for the POST /v1/responses/input_tokens endpoint. OpenAI describes that count as the exact number the model will receive, including the formatting tokens a local tokenizer cannot see, which is a stronger claim than Anthropic makes for its own counting endpoint.
  3. 3John Yang, Carlos E. Jimenez, Alexander Wettig, Kilian Lieret, Shunyu Yao, Karthik Narasimhan, and Ofir Press, “SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering,” NeurIPS 2024, arXiv:2405.15793. Peer-reviewed. The founding result establishing that the same model scores materially differently under different agent-computer interfaces, and that constrained purpose-built tools with structured feedback outperform raw shell access. The counter-demonstration is equally important, and it comes from the same group’s later mini-swe-agent () rather than from the 2024 paper: a roughly hundred-line harness with no tools but bash, a linear history, no context management, no sub-agents, and no memory scores within a few points of elaborate scaffolds. Anyone selling harness complexity should be asked to beat that baseline.
  4. 4TOON (Token-Oriented Object Notation), specification v3.3 and reference implementation, MIT licensed, verified September 8, 2026. Project documentation and self-reported benchmarks. Cited for the format definition—YAML-style indentation for nested objects, inline name[N]: for primitive arrays, tabular name[N]{fields}: for uniform object arrays—for the concept of tabular eligibility, and for the two-track benchmark results as published in the May 8, 2026 revision of docs/guide/benchmarks.md, pinned at commit 48191288590767858b132732d88f39b0d85a6ae4, across 209 retrieval questions and four models: on the flat-only track CSV totals 63,997 tokens against TOON’s 67,778 (+5.9%); on the mixed-structure track TOON totals 227,830 against formatted JSON’s 291,711 (−21.9%) but compact JSON’s 198,546 (+14.7%); accuracy 76.4 percent against JSON’s 75.0. The benchmark page is versioned with the project, and these figures were replaced on July 24, 2026: from that commit onward the same page reports 244 questions, 72.2 percent TOON accuracy against 71.4 for JSON, and a mixed-structure gap against compact JSON of 1.6 percent rather than 14.7. An earlier revision of this document dated the figures above to a snapshot taken on September 8, 2026, which the repository history rules out—by then the page carried the 244-question run. Cite the commit rather than the page: every figure above is reproducible from the pinned revision and none of them may be mixed with totals from a later one. Also cited for the nested-data results in §7.3—configuration 620 tokens against compact JSON’s 558, event logs 154,084 against 128,529, TOON overheads of 11.1 and 19.9 percent—and for the truncation-detection result in which TOON scored 0 of 4 against CSV’s 4 of 4. The benchmark evaluates six formats, TOON, JSON, YAML, compact JSON, XML and CSV, and does not evaluate TRON; an earlier revision of this document attributed those nested-data figures to “both optimizers” and to “every one of them,” which this source does not support. Token counts use the o200k_base tokenizer. The benchmarks are the project’s own and have not been independently replicated, but they are unusually complete: the documentation publishes the cases where the format loses, separates CSV-eligible from CSV-ineligible data rather than averaging across both, and states plainly that the evaluation tests comprehension rather than generation.
  5. 5Ivan Matveev, “Token-Oriented Object Notation vs JSON: A Benchmark of Plain and Constrained Decoding Generation,” arXiv:2603.03306, February 2026. Preprint, not peer-reviewed. Cited as the counterweight to reference 4: it observes that TOON’s published results test model comprehension rather than generation, and—on four structural cases rather than a sweep of dataset sizes—finds the format’s up-front prompt overhead large enough that TOON consumed more tokens than plain JSON on several models. From that the author advances what he labels a scaling hypothesis: that TOON’s efficiency advantage likely follows a non-linear curve, materializing only past some point where accumulated syntax savings amortize that overhead. It is a hypothesis and is presented as one; the experiment does not measure the curve or locate the threshold, and the paper’s own recommendations call for benchmarking at substantially larger dataset sizes to validate it. Relevant to anyone considering TOON as an output format rather than an input one.
  6. 6Original work on agentic iteration economics establishing that agentic tasks consume roughly a thousand times the tokens of code chat, that input rather than output drives that cost, that runs on the same task differ by up to thirtyfold in total tokens, that accuracy frequently peaks at intermediate cost, and two separate results about anticipating cost that an earlier revision of this document merged into a claim about human forecasting. The paper tests model self-prediction directly: correlations up to 0.39, with systematic underestimation. Its human data are SWE-bench-Verified’s expert estimates of how long a professional developer would need to resolve each issue, compared against agent token consumption; it finds that difficulty category is a weak predictor of spend. No human was asked to forecast agent tokens or cost, so the study does not establish that expert humans forecast task cost badly—only that human-effort difficulty transfers weakly as a proxy for it. Cited via reference 1, which contains the full source annotation. The thirtyfold variance figure is the single most consequential number for capacity planning in this document.
  7. 7Vendor modeling of token accumulation across a fifty-turn agentic session: roughly 5,000 input tokens per turn for turns one through ten, 20,000 for turns eleven through thirty, and 35,000 for turns thirty-one through fifty, with input outnumbering output twenty to twenty-five times. Vendor-published; the shape is the point rather than the specific figures. Cited via reference 1.
  8. 8Peer-reviewed work identifying self-conditioning: models become more likely to err when the context contains their own prior errors. Not a long-context artifact—injecting artificial error histories reproduces the effect, and larger models are more susceptible despite better long-context handling. The same work finds that reasoning-trained models eliminate self-conditioning entirely, which is a meaningful qualification on the practices in §24. Cited via reference 1.
  9. 9Multi-turn conversational degradation of approximately 39 percent, decomposed into a minor aptitude loss and a large increase in unreliability, driven by models making early assumptions and over-relying on them. The study tested conversational generation rather than agentic coding trajectories, so transfer to coding is plausible and unproven. Cited via reference 1.
  10. 10Failure taxonomy across 1,794 complete agent trajectories and more than 63,000 execution steps, seven models and three scaffolds. Source of the false-premise rate (30.7 percent), the epistemic/competence/environment breakdown (57.9 / 32.8 / 9.4 percent), the finding that 82 percent of failed trajectories continue executing after the failure is empirically unrecoverable, that the first observable signal surfaces roughly ten steps after the decisive error, and that 71 percent of successful trajectories recover from at least one error. The paper separates three events, and an earlier revision of this document collapsed two of them: the decisive error, t_lock (the point after which no correct recovery is observed), and the first observable signal. The 82 percent continued-execution figure is measured from t_lock; the ten-step lag is measured from the decisive error. Also the source of the prefix-monitor results cited in §25.4: roughly 2 to 3 percent false positives and about 82 percent precision at recognizing a locked-in failure, against recall under thirty percent and a median lead time of zero relative to t_lock, with only 3.7 to 8.7 percent of failures flagged before lock-in. That is failure confirmation rather than loop detection, and §25.4 is scoped accordingly. The strongest published failure taxonomy for coding agents. Cited via reference 1.
  11. 11Anthropic, “Data residency,” Claude Platform documentation, verified September 20, 2026. Vendor documentation; cited as product fact for the inference_geo request parameter and its two accepted values—"global", the default, documented as running inference in any available geography, and "us", which confines it to US-based infrastructure—for the usage.inference_geo response field reporting where inference actually ran, for the workspace-level default_inference_geo and allowed_inference_geos settings, and for the 1.1x multiplier applied to US-only inference on Claude 4.6 and later across input tokens, output tokens, cache writes, and cache reads. Scope matters for §5.3’s advice: the parameter exists on the Claude API and Claude Platform on AWS only, and on Claude 4.6 and later models only, returning a 400 on earlier ones. Bedrock and Google Cloud determine the inference region from the endpoint URL or inference profile instead, and Microsoft Foundry from the deployment type, which means the geography is pinnable on those platforms but is not reported back in the response. Cited here because it is the one point in the stack where a serving-infrastructure variable is exposed to the caller rather than inferred from behavior. The 1.1x multiplier is a billing fact and should be re-checked alongside the price tables.
  12. 12Anthropic, “Prompt Caching,” Claude Platform documentation, verified September 8, 2026. Vendor documentation; cited as product fact for mechanism, pricing multipliers, minimum cacheable lengths, invalidation behavior, the 20-block lookback window, pre-warming, and data retention. The pricing table in this reference is the primary source for all Anthropic rates quoted in this document.
  13. 13Work combining serial iteration with parallel candidate generation, where each trajectory generates a test script alongside its draft edit, reaching 57.4 percent at roughly $4.60 per instance; selection across edits drawn from top existing submissions reached 66.2 percent, outperforming the best individual ensemble member. Establishes that selection, not generation, is the under-invested problem. Cited via reference 1.
  14. 14Julia Kasper, Megan Rogge, and Aaron Munger, “The Coding Harness Behind GitHub Copilot in VS Code,” Visual Studio Code blog, May 15, 2026. Vendor-published engineering account; cited as product fact for the three harness responsibilities (context assembly, tool exposure, tool execution), the turn/round/run vocabulary, and the per-model divergence in system prompts, tool sets, and conversation management—including that Claude models are given replace_string_in_file while GPT models are given apply_patch, that Gemini requires reminders to use tool-calling and breaks on orphaned tool calls, and that the harness selects a different system prompt per model version. The same post describes VSC-Bench and reports an effort setting that consumed more tokens while resolving slightly fewer tasks, which independently corroborates the saturation finding in reference 6. As a vendor account of its own product it is authoritative on mechanism and self-interested on quality; the mechanism is what is cited here.
  15. 15Visual Studio Code Team, “How Prompt Tuning Improved GPT-5.5 in VS Code,” Visual Studio Code blog, July 6, 2026. Vendor-published controlled experiment: a two-week online A/B test on live agent traffic, one control and two system-prompt treatments at a 25/25/25 split, with per-metric effect sizes and p-values. Source of the figures in §6.3. Unusually credible for vendor research because it reports an unfavorable movement (a small, marginally significant decline in ten-minute code survival under the shipped treatment) alongside the favorable ones, and because the prompt text of both treatments is published in the open-source repository. Limitations to carry: a single model on a single product, quality proxied by code-survival rather than correctness, and no independent replication.
  16. 16TRON (Token Reduced Object Notation) specification and reference implementations, with SDKs, verified September 8, 2026. Project documentation and self-reported figures. Cited for the format definition—an optional header of class declarations followed by a data section using ClassName(v1,v2,…) instantiation syntax—for its status as a superset of JSON in which any valid JSON document is already valid TRON, and for the claimed 20 to 40 percent token reduction achieved by hoisting repeated object schemas into the header. That magnitude comes from the PyPI package description, pinned at tron-python 0.1.0 since a registry version is immutable; the figure appears there and not in the specification, which commits to no figure and states that the more aggressive encoding strategy “does not always guarantee fewer tokens than pure JSON encoding.” The reduction claim is the project’s own and has not been independently replicated at that magnitude; the agentic evaluation in reference 17 measures 0 to 27 percent in end-to-end loops. One documentation claim warrants skepticism: the body section is described as readable by existing JSON parsers without modification, yet A("x",1) is not valid JSON, so a TRON-aware parser is required in practice for any document that actually uses the class syntax.
  17. 17“Notation Matters: A Benchmark Study of Token-Optimized Formats in Agentic AI Systems,” arXiv:2605.29676, June 2026. Preprint, not peer-reviewed, and the only independent end-to-end evaluation of these formats located. Evaluates TOON and TRON across four agentic benchmarks (BFCL, MCPToolBenchPP, MCP-Universe, StableToolBench) and five configurations of four open-weight models—Mistral-Small-24B, Qwen3-32B with thinking on and again with it off, DeepSeek-R1-Distill-Qwen-32B, and Llama-4-Scout-17B-16E—which the paper itself describes as “five configurations span four model families.” An earlier revision of this document called them five models; the thinking toggle on Qwen3-32B is one of the study’s own comparisons, so the distinction is load-bearing rather than pedantic. Deliberately decouples input compression from output compression so that comprehension and generation are measured separately. Source of the per-model token reductions (TOON 2 to 18 percent, TRON 0 to 27 percent), the finding that the dominant pattern is per-benchmark rather than per-format, the TRON accuracy drops of 27 to 44 percentage points on BFCL under reasoning configurations, which fall in the study’s input-only condition with tool calls still emitted as JSON, the parsing-cascade effect observed for TOON on the two multi-turn benchmarks, the observation that tool-call training transfers to comprehension but not to robust generation in unfamiliar formats, and the conclusion identifying TRON as a defensible drop-in for JSON in token-sensitive agentic systems. The most important single reference in §7, and the only one measuring these formats where practitioners actually deploy them.
  18. 18Jia He, Mukund Rungta, David Koleczek, Arshdeep Sekhon, Franklin S. Wang, and Sadid Hasan, “Does Prompt Formatting Have Any Impact on LLM Performance?,” arXiv:2411.10541, November 2024. Preprint. Formats identical content into plain text, Markdown, JSON, and YAML and evaluates across natural language reasoning, code generation, translation, and named-entity tasks on four OpenAI models. Source of the abstract’s claim that GPT-3.5-turbo varies by as much as 40 percent on a code translation task purely by template, and of the finding that format sensitivity is statistically significant (p < 0.05) on every dataset and model pairing tested except one. Carry the 40 percent as an unreconciled abstract claim, because the paper’s own translation results do not display it: the widest GPT-3.5 spread in Tables 7–8 is Java-to-C# BLEU running from 66.46 under plaintext to 78.40 under JSON, which is 11.94 points or roughly 18 percent relative, and the paper’s other headline number—accuracy up 42 percent for JSON over Markdown—is from MMLU international law rather than code translation. This is an internal inconsistency in the source rather than a misquotation, and no calculation reconciling the abstract with the tables appears in the paper. Where a defensible magnitude is needed, cite the table result; do not silently swap in the 42 percent figure, which is a different task and a different metric. Also the source for the observation that larger models are more robust without being insensitive, and that the best-performing template does not transfer reliably between models. The study varies whole-prompt template rather than inline emphasis markup, which is the limitation relevant to §8.4.
  19. 19Melanie Sclar, Yejin Choi, Yulia Tsvetkov, and Alane Suhr, “Quantifying Language Models’ Sensitivity to Spurious Features in Prompt Design: Or, How I Learned to Start Worrying About Prompt Formatting,” ICLR 2024, arXiv:2310.11324. Peer-reviewed. Introduces FormatSpread, which searches over semantically equivalent prompt formats rather than testing a handful by hand, and reports accuracy differences of up to 76 percentage points attributable to formatting alone on tasks drawn from SuperNatural-Instructions. Cited for the magnitude of format sensitivity and for the methodological point that evaluating a model on a single format reports one sample from a wide distribution—the same argument §34.4 makes about running an eval once.
  20. 20Nelson F. Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni, and Percy Liang, “Lost in the Middle: How Language Models Use Long Contexts,” Transactions of the ACL, 2024. Peer-reviewed. Establishes the U-shaped positional accuracy curve for mid-context material, replicated across six model families and confirmed on additional architectures since. The shape is the cross-model finding; the magnitude is not. In the thirty-document multi-document QA table the steepest edge-to-middle decline exceeds thirty percent relative—a fall of roughly 22 percentage points, from about 73 to about 51—while other models in the same table decline by single-digit percentages. An earlier revision of this document reported the greater-than-thirty-percent figure as a cross-model result; it is a maximum observed effect, and the entry now says which unit it is in, since a relative percent and a percentage point are not the same quantity. The paper reports task accuracy by position and does not publish an attention-weight profile (see reference 65).
  21. 21Field study of 15,549 agentic pull requests across 148 projects, finding near-symmetric outcomes after instruction-file introduction: 27.7 percent of projects raised merge rate by at least twenty percent, 26.4 percent lowered it by at least twenty percent. Source of the median word counts (976 for improved projects, 569 for declined). Cited via reference 1.
  22. 22Controlled study across benchmark tasks and developer-committed issues, using both generated and human-written context files across multiple models and agents, finding that context files did not generally improve task success rates while increasing inference cost by over twenty percent on average. The null result held across models, agents, and file provenance. Cited via reference 1.
  23. 23Paired within-task study of 124 pull requests finding median wall-clock time down 28.6 percent and median output tokens down 16.6 percent with an instruction file present. The study explicitly did not evaluate task success, which is the material limitation. Cited via reference 1.
  24. 24Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed Chi, Quoc Le, and Denny Zhou, “Chain-of-Thought Prompting Elicits Reasoning in Large Language Models,” NeurIPS 2022, arXiv:2201.11903. Peer-reviewed. The foundational chain-of-thought result. Discovered on models that did not reason unless prompted; §10.1 covers what changed.
  25. 25OpenAI, “Reasoning best practices,” OpenAI API documentation, and Anthropic extended thinking guidance. Vendor documentation. Cited for the reasoning-model prompting differences in §8.7, which follow directly from what post-training rewarded and are therefore unusually reliable vendor advice.
  26. 26Anthropic model pricing, published in the pricing table of reference 12 and, verified September 8, 2026. Source for all Claude per-model rates: Fable 5.1 $10/$50, Opus 5 $5/$25, Sonnet 5 $2/$10, Sonnet 4.6 $3/$15, Haiku 4.5 $1/$5 per MTok, with cache multipliers of 1.25× (5m write), 2× (1h write), and 0.1× read (0.025× on Fable 5.1 and Mythos 5.1).
  27. 27Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa, “Large Language Models are Zero-Shot Reasoners,” NeurIPS 2022, arXiv:2205.11916. Peer-reviewed. The zero-shot chain-of-thought trigger result.
  28. 28Peer-reviewed test-driven agentic development framework supplying tests followed by remediation loops, moving one model from 69.7 to 82.5 to 87.7 percent on one Python benchmark and 78.7 to 87.8 to 93.3 percent on another. Gains shrink sharply at file scope (23.0 to 30.3 percent). The study names three limits: more tests can hurt through lost-in-the-middle effects, solutions sometimes satisfy only supplied tests, and the returns on additional tests are dataset-dependent. That last one is easy to misreport, and an earlier revision of this document did, as a general plateau after roughly three tests. The paper says the opposite for one of its two datasets: on the 143-problem HumanEval subset it states that it does not observe a plateau, while the 398-problem MBPP subset is already declining at three tests—which the authors attribute to MBPP’s simpler problems, noting that the first test alone reaches full line coverage on 92.4 percent of MBPP cases against 75 percent on HumanEval. Report the dataset, not a threshold. Cited via reference 1.
  29. 29Jie Huang, Xinyun Chen, Swaroop Mishra, Huaixiu Steven Zheng, Adams Wei Yu, Xinying Song, and Denny Zhou, “Large Language Models Cannot Self-Correct Reasoning Yet,” ICLR 2024, arXiv:2310.01798. Peer-reviewed. Establishes that models struggle to self-correct without external feedback and that performance sometimes degrades after self-correction, with prior positive results depending on oracle labels. This is the citation to use against any vendor claim that an agent reviews its own work.
  30. 30CR-Bench, arXiv:2603.11078. Preprint. Measurement of critic subagents on CR-Bench-verified, the 174-case verified subset of a 584-case corpus of real pull-request defects; the verified subset, not the full corpus, is the population every figure here is drawn from. Be precise about the context the reviewers got, which an earlier revision of this document described as full repository context: the prompts in Appendix B supply the repository name, PR number, title, description and diff, and nothing else from the tree. The benchmark’s own comparison table claims “Full PR Context,” meaning the whole pull request rather than isolated diff hunks, which is a different thing. The paper is explicit about the consequence—it attributes weak recall on usability and functional-suitability defects to context “not fully contained within the PR diff,” calling these “closed-context code review agents, lacking access to the broader system state.” That makes the measurement a floor for diff-scoped review rather than a verdict on what a repository-aware reviewer could do. Single-shot reviewer at 27.0 percent recall and 3.6 percent precision; iterative self-critique raising recall to 32.8 percent while collapsing signal-to-noise from 5.11 to 1.95 on the large model and 2.89 to 0.91 on the small one. Signal-to-noise is bug hits plus valid suggestions over noise, so below 1.0 the reviewer emits more noise than useful output of any kind, and precision counts confirmed bugs alone against a usefulness rate of 83.6 percent on the same row. Also cited via reference 1.
  31. 31Ziqi Yin, Hao Wang, Kaito Horio, Daisuke Kawahara, and Satoshi Sekine, “Should We Respect LLMs? A Cross-Lingual Study on the Influence of Prompt Politeness on LLM Performance,” Proceedings of SICon 2024, arXiv:2402.14531; Om Dobariya and Akhil Kumar, “Mind Your Tone: Investigating How Prompt Politeness Affects LLM Accuracy,” arXiv:2510.04950, 2025—whose own arXiv metadata describes it as a short paper under submission to Findings of ACL 2025, with no proceedings record located, so it carries no venue claim here; Cai et al., “Does Tone Change the Answer? Evaluating Prompt Politeness Effects on Modern LLMs: GPT, Gemini, and LLaMA,” arXiv:2512.12812; and Om Dobariya and Akhil Kumar, “Mind Your Tone: Does Tone Alter LLM Performance?,” Proceedings of AMCIS 2026, arXiv:2605.29027—the same authors’ peer-reviewed full-paper extension of their own short paper, and the largest measurement in this group. Two peer-reviewed papers and two preprints, cited together because they disagree and the disagreement is the finding. Yin et al. report impolite prompts often degrading performance with an optimum that varies by language; Dobariya and Kumar report the reverse on GPT-4o, 80.8 percent under very polite phrasing against 84.8 percent under very rude, across fifty questions in ten runs; Cai et al. test three model families on MMLU and find small, mostly non-significant effects favoring neutral and polite phrasing where significant at all, concentrated in Philosophy and Professional Law, with Gemini showing no significant sensitivity in any comparison. The 2025 Dobariya and Kumar result circulated widely and rests on fifty questions, a single model, and no peer review; it is cited here for completeness rather than as guidance, and that assessment of it stands. Their 2026 extension is a different artifact and carries the weight instead: 570 MMLU questions across all 57 subjects, seven tones including Sycophantic and Threatening, four models, ten runs, within-subjects paired tests with Holm correction. It reports ChatGPT-5-nano ranging from 82.37 percent (Neutral, its best) to 71.25 (Threatening), Gemini 2.5 Flash Lite spanning 12.46 points, and a significant global tone effect even on the least sensitive model. An earlier revision of this document concluded from the first three sources that tone effects do not survive aggregation across domains; this one aggregates across every MMLU subject and they survive. It also undercuts the rude-is-better reading its own predecessor produced, since neutral phrasing wins where the effect is largest.
  32. 32Fixed three-phase pipeline (localize, repair, validate) using regression tests plus generated reproduction tests, with no autonomous tool selection, resolving 32 percent of SWE-bench Lite at roughly $0.70 per issue. The authors scope both claims and an earlier revision of this document did not: the accuracy lead is over the evaluated open-source approaches, and the cost is below most prior agent-based approaches rather than all. The paper’s own Table 1 is explicit about it—CodeStory Aide solves 43.00 percent, and Moatless runs at $0.17 and $0.14—and the paper states in terms that 32.00 percent “is not the highest percentage of problems solved on SWE-bench Lite.” The argument does not need universal superiority: a fixed pipeline with execution verification beating every open-source agent at a cost most of them exceed is the finding. The strongest available evidence that verification rather than autonomy carries the value. Cited via reference 1.
  33. 33Literature on correct-to-incorrect sycophancy, comprising Ethan Perez et al., “Discovering Language Model Behaviors with Model-Written Evaluations,” 2023; Mrinank Sharma et al. on sycophancy in RLHF-trained assistants; “Why LLMs Give In: Conversational Factors and Reasoning Behind Medical Sycophancy,” arXiv:2608.01017, which uses a fully crossed design over user role, evidence, challenge turn, and grounding, and takes a pre-challenge competence baseline so capitulation can be distinguished from ignorance; “Sycophancy is an Educational Safety Risk,” arXiv:2605.14604, source of the finding that models resisting context-switch frame attacks still capitulate under authority claims and social-affective pressure; and “What Counts as AI Sycophancy? A Taxonomy and Expert Survey,” arXiv:2605.21778, which surveys 44 papers operationalizing factual capitulation and reports that training models toward warmth and empathy substantially increases sycophancy, amplified when users express vulnerability. Mixed peer-reviewed and preprint. The medical and educational studies are domain-specific and their effect sizes should not be read as transferring to software engineering; the qualitative findings §10.5 relies on are that capitulation is governed more by the question than the model, that pressure mode is a dominant driver, and that warmth increases susceptibility. Which pressure mode dominates is not consistent across models and should not be reported as though it were: in the educational study’s Table 5, authority and social-affective pressure are the effective levers on GPT-5.2 (16.8 and 18.1 percent against 7.7 for context switching) while Claude Sonnet 4.5 inverts it (17.9 percent on context switching, 8.9 on social-affective). An earlier revision of this document generalized the GPT-5.2 ordering to both. The authors explicitly decline to read the difference as a model ranking, treating it instead as evidence that similar aggregate rates conceal different fragility profiles.
  34. 34Zhi Rui Tam, Cheng-Kuang Wu, Yi-Lin Tsai, Chieh-Yen Lin, Hung-yi Lee, and Yun-Nung Chen, “Let Me Speak Freely? A Study on the Impact of Format Restrictions on Performance of Large Language Models,” Proceedings of EMNLP 2024, Industry Track, pages 1218–1236, DOI 10.18653/v1/2024.emnlp-industry.91, arXiv:2408.02442. Peer-reviewed. Establishes that format restrictions degrade reasoning while improving classification accuracy. Cited for that direction and not for a monotonic strictness ordering, which an earlier revision of this document reported as constrained-decoding > format-restricting-instructions > natural-language-then-convert. The published Table 2 contradicts it: on gpt-4o-mini the JSON-Schema condition, which is the stricter constraint, scores 91.71 / 81.77 / 86.07 on GSM8K, Shuffled Objects and Last Letter against the format-restricting instruction’s 87.17 / 81.46 / 84.73, and on Last Letter exceeds the natural-language mean of 83.11. The JSON-mode condition is the worst in all three rows, so the study’s poor JSON-mode results are what the ordering was built on, and they do not generalize to schema-constrained decoding. The reported standard deviations are wide and these comparisons are not tests of pairwise significance; they are sufficient to refute a universal ordering, not to establish the reverse one. Two further limitations are load-bearing: the study predates current-generation reasoning models, and the authors have published updates in response to methodological critique, which do not establish an intrinsic monotonic penalty either. The direction is well established; the effect size on current models is unconfirmed.
  35. 35Kelly Hong, Anton Troynikov, and Jeff Huber, “Context Rot: How Increasing Input Tokens Impacts LLM Performance,” Chroma Research, 2025. Vendor-published research; the publisher sells vector databases and has an interest in the conclusion, which should be stated. Tests 18 frontier models and finds that none use context uniformly, that reliability degrades with input length, and that focused prompts substantially outperform full prompts containing the same relevant material plus distractors. The finding is consistent with independent work on position effects and on effective context length (references 20 and 36) and is the empirical basis for compaction, sub-agent isolation, and aggressive curation.
  36. 36Cheng-Ping Hsieh, Simeng Sun, Samuel Kriman, Shantanu Acharya, Dima Rekesh, Fei Jia, Yang Zhang, and Boris Ginsburg, “RULER: What’s the Real Context Size of Your Long-Context Language Models?,” arXiv:2404.06654, 2024. Establishes that models claiming 32K or more frequently fail to maintain performance across their advertised range, with failure modes including a failure to ignore distractors and reverting to parametric knowledge. The distractor mode is easy to write backward, and an earlier revision of this document did: ignoring distractors is the desired behavior, and what the paper reports is models incorrectly retrieving values associated with the distractor keys. Also relevant: Ali Modarressi et al., “NoLiMa: Long-Context Evaluation Beyond Literal Matching,” 2025, which demonstrates that non-lexical matching degrades sharply with length.
  37. 37Bowen Qin and Yi Xie, “Agent Retrieval Bench: Evaluating Repository Context Retrieval for Coding Agents,” arXiv:2607.24882, July 2026. Preprint. 427 samples across 25 repositories: 345 positive retrieval examples, 50 natural no-gold cases, and 32 counterfactual wrong-repository controls. Two populations inside it must be kept apart, and an earlier revision of this document attached both figures below to the 427. Both are drawn from the 287-sample code2test / comment2context / trace2code subset: the logged-trajectory result that interactive agents never touched any gold file on 27 to 35 percent of samples (35.2 percent for the OpenAI strict-context agent, 27.2 to 29.3 percent for Codex), and the span annotations whose median labeled evidence is 27 lines, or 4.7 percent of its containing file, with 75.7 percent of span-file pairs using at most a tenth of the file. The 4.7 percent is a median line fraction over 391 span-file pairs, not a token-based signal-to-noise ratio, and the paper explicitly declines to treat unjudged content as waste, so it does not support an inference that the rest of the file is irrelevant. The paper is also explicit that it measures document reachability rather than within-file localization, and that it does not establish that a retrieval miss causes a repair failure.
  38. 38Vendor-published comparison of just-in-time agentic exploration against maintained semantic indexing, reporting code retention improving 0.3 percent overall and 2.6 percent on codebases over a thousand files. Two major vendors hold publicly opposite positions on this architecture and each publishes evidence favoring its own; scale-dependence is the only synthesis the evidence supports. Cited via reference 1.
  39. 39Linux Foundation, “Linux Foundation Announces the Formation of the Agentic AI Foundation (AAIF), Anchored by New Project Contributions Including Model Context Protocol (MCP), goose and AGENTS.md,” December 2025; and the AGENTS.md specification. Primary announcement plus specification. Establishes neutral governance for both MCP and AGENTS.md, which is the durable signal for an organization standardizing on either.
  40. 40Study collecting 2,303 context files across 1,925 repositories and three agentic coding tools. Separate the corpus from the analysis population, which an earlier revision of this document did not: the content percentages come from a manually labeled subset of 332 Claude Code files, labeled by two inspectors with a third resolving 438 disagreements, and not from a content analysis of all 2,303. On that subset, testing content appears in 75 percent, implementation detail in 69.9 percent, and architecture in 67.7 percent, with security and performance each appearing in only 14.5 percent. The source carries some internal count inconsistencies around label totals; do not infer a replacement denominator from the rounded percentages, and do not substitute a later revision for the pinned v1, since a restatement there would not be the figures quoted here. The artifact steering every agent almost never mentions the two quality attributes that degrade fastest. Cited via reference 1; the percentages quoted here match version 1 of the preprint, which is the version to pin, since a later revision may restate them.
  41. 41Agent Skills specification, and vendor documentation for skill directories on Claude Code, GitHub Copilot, and OpenAI Codex. Open specification plus vendor documentation; cited as product fact for the progressive-disclosure mechanism.
  42. 42Model Context Protocol specification, revision 2026-07-28, verified September 11, 2026. Specification. Cited for transport (Streamable HTTP current, HTTP-with-SSE formally deprecated under the feature lifecycle policy), the OAuth resource-server authorization model, and the primitive set. This revision made the protocol core stateless, removing the initialize handshake and the Mcp-Session-Id header; required Mcp-Method and Mcp-Name headers on Streamable HTTP requests; added ttlMs and cacheScope to list and resource-read results; and deprecated OAuth 2.0 Dynamic Client Registration in favor of Client ID Metadata Documents. Servers written against 2025-11-25 or earlier will not be conformant without transport work. The community registry remains in preview with the API frozen at v0.1.
  43. 43Anthropic, “Subagents,” Claude Code documentation. Vendor documentation; cited as product fact for the isolation mechanism—fresh context window, only the final message returning to the parent, and tool grants bounded by the definition. No published measurement exists of whether sub-agent isolation improves task success on software engineering work; the mechanism is documented and the benefit is asserted.
  44. 44Anthropic, “How we built our multi-agent research system,” Anthropic Engineering, 2025. Vendor-published. Reports a 90.2 percent improvement over a single agent on an internal research evaluation, and in the same publication notes roughly fifteen times the token usage of chat for multi-agent systems and about four times chat for single agents, that token usage alone explained 80 percent of performance variance on the relevant benchmark, and that most coding tasks involve fewer truly parallelizable subtasks than research. Both denominators are chat; neither is a matched comparison against one coding agent, and an earlier revision of this document quoted the fifteen as though it were. The vendor with the strongest published multi-agent result explicitly carves coding out of it.
  45. 45Peer-reviewed study of over 1,600 annotated execution traces across seven multi-agent frameworks, producing a fourteen-mode failure taxonomy in three categories (system design, inter-agent misalignment, task verification) with inter-annotator agreement at κ = 0.88, and finding that better prompts and added structure produce partial gains only. Establishes that multi-agent failure is architectural rather than promptable. Cited via reference 1.
  46. 46Study comparing single-agent and multi-agent systems under matched thinking-token budgets across five multi-agent variants and three models, finding single-agent best or statistically indistinguishable at every budget except the lowest. The word “only” was wrong in an earlier revision of this document: in the paper’s Table 13, sequential agents are already level with single agents at 50 percent corruption—0.223 versus 0.223 under masking, 0.219 versus 0.216 under deletion with overlapping bootstrap intervals—and the 70 percent results are mixed rather than a clean handover, with sequential ahead under masking and substitution but behind under deletion. The aggregate single-agent result stands and individual table cells do not overturn it; what they overturn is a clean corruption threshold. Covers multi-hop question answering rather than software engineering, which should be stated. Cited via reference 1.
  47. 47Vendor guidance establishing the read/write discriminator for parallelism: read actions parallelize and write actions do not. Vendor-published; the reasoning is mechanical rather than measured, and it is the most useful architectural heuristic available on this question. Cited via reference 1.
  48. 48Claude Code documentation (settings, hooks, sub-agents, and skills references), verified September 11, 2026, cross-checked against an independently compiled feature and settings snapshot. Vendor documentation plus a third-party catalog that links each row back to the official docs. Cited for the settings precedence tree (user, project, project-local, CLI flags, enterprise managed, in ascending precedence, with the managed layer a floor that CLI flags cannot relax for scalar values and deny rules—list-valued keys such as permissions.allow and the sandbox allow and exclusion arrays merge across scopes instead, so lower scopes can add entries and widen access, which allowManagedPermissionRulesOnly exists to prevent for permission rules), the hook event catalog including PostCompact and its auto/manual matcher, the hook exit-code semantics, subagent frontmatter fields and isolation: "worktree", and the documented routing of subagent permission prompts—foreground subagents pass prompts through to the user, background subagents surface them in the main session naming the asking subagent, and auto-denial is a permission-mode behavior rather than a property of delegation. An earlier revision of this entry asserted that subagents cannot raise interactive prompts at all, so approval-required calls always resolve as denials; that was wrong, and §18.5 was corrected before this entry was. The same revision compressed the exit-code semantics to “0 allow, 1 allow with warning, 2 deny,” which conflates the handler’s process status with the event’s decision, and the hook printed in §19.2 is the counterexample: it emits a permissionDecision of deny and exits 0. Exit 0 means the handler succeeded and Claude Code reads the decision from stdout JSON—silence is not approval, it is merely no decision, and the call continues through the normal permission flow. Exit 1 is a non-blocking error that Claude Code proceeds past, not a warning-flavored allow. Exit 2 blocks, but which events can block is event-specific: PreToolUse and UserPromptSubmit block, while PermissionRequest, PostToolUse, Notification, SessionStart and others do not honor it. Read the per-event table rather than a three-value mapping. Also cited, against the memory page and the v2.1.277 release notes of September 18, 2026, for native AGENTS.md loading and its conditions: by default Claude reads AGENTS.md only where no CLAUDE.md, .claude/CLAUDE.md or CLAUDE.local.md sits in the working directory or above it, while a user-level CLAUDE.md, a managed one and .claude/rules/ files do not count against it; a Project instructions setting in /config selects other modes, including loading both; and nested and subdirectory files load on access. The provider limitation this entry previously recorded as current is now version-scoped: the memory page places it before v2.1.281, published September 23, 2026, and directs affected Bedrock users to update rather than describing an ongoing platform gap. The verification date in this entry was accurate when made; this is a product change after it, not a correction to it. The conditions that remain current are an installation before v2.1.277, a disabled agents-md plugin, and in some cases the first session after an upgrade. An earlier revision of this document said Claude Code simply does not read AGENTS.md and presented the import line as a universal requirement; §15.2, §38.6, §43.1 and Appendix D were corrected together. The third-party snapshot is dated May 2026 and its model-name rows are consequently stale against the lineup in reference 12; the mechanism rows cited here were re-checked against the current official pages.
  49. 49OpenAI, “Prompt caching,” OpenAI API documentation, verified September 8, 2026. Vendor documentation; cited as product fact for implicit and explicit breakpoint modes, the 1.25× write and 0.1× read multipliers on GPT-5.6 and later, minimum cacheable lengths, TTL and retention semantics, machine-local cache routing and the ~15 requests-per-minute overflow threshold, prompt_cache_key design guidance, the minimum-cacheable-length break-even formula, and the compaction interaction.
  50. 50Google, “Context caching,” Gemini API documentation, and “Context caching overview,” Gemini Enterprise Agent Platform documentation. Vendor documentation; cited as product fact for the implicit/explicit distinction, per-model minimum token counts, the 90 percent discount on Gemini 2.5 and later (75 percent on 2.0), the default one-hour TTL on explicit caches, and the statement that storage costs apply to explicit caching only. The minimums are model-specific rather than tier-specific and were re-checked on the vendor page on September 25, 2026: 2,048 on Gemini 2.5 Flash and 2.5 Pro, 4,096 on the 3.x models including Gemini 3.1 Pro.
  51. 51Google Gemini API pricing. Vendor documentation. Per-token rates verified against Google’s own page on September 24, 2026: Gemini 3.1 Pro at $2.00 input, $0.20 cached input and $12.00 output for prompts at or below 200K tokens, and $4.00, $0.40 and $18.00 above it. The Pro storage rate of $4.50 per million tokens per hour was verified against the vendor page on September 25, 2026, along with its Batch table, whose cached-input row reads “same as Standard” and whose storage rate is unchanged—so for that model batching and caching do not compound. That does not generalize, and an earlier revision of this document wrongly rejected an audit finding which said so. Verified on the same page on September 25, 2026: Gemini 3.6/3.7/3.8 Flash Batch halves cached input ($0.075 to $0.0375 in the promotional period), 3.5 Flash Batch halves it ($0.15 to $0.075), and 3.1 Flash-Lite Batch halves both cached input ($0.025 to $0.0125) and storage ($1.00 to $0.50 per million tokens per hour). Batch and cache interaction is a per-model, per-tier fact. Cache storage was re-verified per model and per tier against the same page on September 26, 2026, and it varies along both axes, which is why no single “Flash-tier storage” figure exists. Gemini 3.6, 3.7 and 3.8 Flash are $0.50 per million tokens per hour through December 31, 2026, rising to $1.00 on January 1, 2027, and that promotional row is identical on Standard and Batch—for these models batching does not discount storage, only cached input. Gemini 3.5 Flash is a flat $1.00 on both tiers with no promotional row anywhere in its pricing. Gemini 3.1 Flash-Lite is $1.00 on Standard and $0.50 on Batch, with no promotional row, so here the halving is a tier effect rather than a promotion. Two earlier revisions of this entry each got one axis wrong: one called the $0.50 promotional rate “Flash-tier storage,” which collapses the per-model axis and is contradicted by 3.5 Flash; the other carried $1.00 for the 3.6/3.7/3.8 generation from third-party summaries, which turned out to be the post-promotion rate. Read a storage rate off the row for the exact model and tier, and re-read it after January 1, 2027.
  52. 52OpenAI, “Pricing,” OpenAI API documentation, verified September 8, 2026, with the per-model pages under verified September 24, 2026. Vendor documentation; source for all OpenAI per-model rates including the short-context and long-context tiers. The 272K-token threshold and the 2× input, 2× cache and 1.5× output multipliers are stated on the model pages rather than on the pricing page, which lists the two tiers without naming the boundary. One rate in §21.3’s table is promotional rather than standing: the gpt-5.6-sol model page documents its $4.00/$20.00 as available at least through November 21, 2026, verified September 25, 2026, and describes it as a reduction against the prior generation. Re-check it before carrying it into a forecast.
  53. 53Vendor documentation warning that as the context window fills, older messages are replaced with a summary and that specific instructions from early in a conversation may not be preserved. Unresolved attribution: the product and page behind this exact wording could not be re-identified, and it is retained only as the basis for the constraint re-injection pattern in §24.3, not as support for any broader claim. Where a host documents its own behavior, that documentation governs: Claude Code states that the project-root instruction file is re-read from disk and re-injected after compaction, that nested files and path-scoped rules reload on access, and that an instruction which disappears was given only in conversation.48 §24.3 is scoped to conversation-only instructions accordingly.
  54. 54Nicholas Carlini, “Building a C Compiler with a Team of Parallel Claudes,” Anthropic Engineering, February 5, 2026. Vendor-published, n = 1. Describes externalizing specification, plan, implementation notes, and a running decision log into repository markdown, with verification commands and repair after each milestone. An anecdote rather than a measurement—sixteen agents across nearly two thousand Claude Code sessions over two weeks, consuming 2 billion input tokens and generating 140 million output tokens at a total just under $20,000, but the clearest published account of the long-horizon file-state pattern, and the architectural shape is independently justified by references 6, 7, and 8.
  55. 55Public agentic-coding leaderboard cost comparison: one model-scaffold pair at $366.81 for a 27.2 percent resolve rate against another at $67.09 for 38.0 percent. A counterexample to the assumption that spending more buys accuracy, and nothing more: an earlier revision of this entry said the pair establishes that cost and accuracy are not correlated, which two observations cannot do—they have a sample correlation of exactly −1, and one broken monotonicity says nothing about a population relationship. The associated methodological critique should be read in full: accuracy-only evaluation, inadequate holdout sets producing shortcut-taking, and a lack of standardization, with the finding that state-of-the-art agents are needlessly complex and costly. Cited via reference 1. The two price/score pairs were recorded without model or scaffold names, dataset variant, cost denominator, or a dated snapshot, and the leaderboard route they were read from now returns 404 (the surviving view, /swebench_verified_mini, is a different 50-task board). They are reproduced in §25.1 as an illustration rather than as a citable record, and that is their settled status rather than a gap awaiting an archived table: the pair reached this document through reference 1’s first edition, recorded without the board, model-scaffold pair, dataset variant or snapshot date that would let a particular table be retrieved again. The methodological critique is what the reference actually supports.
  56. 56Huiqiang Jiang, Qianhui Wu, Chin-Yew Lin, Yuqing Yang, and Lili Qiu, “LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models,” EMNLP 2023, arXiv:2310.05736; Zhuoshi Pan et al., “LLMLingua-2: Data Distillation for Efficient and Faithful Task-Agnostic Prompt Compression,” ACL 2024 Findings, arXiv:2403.12968; and Huiqiang Jiang et al., “LongLLMLingua,” ACL 2024. Peer-reviewed. Source of the up-to-20× compression claim, the GSM8K exact-match deltas of 1.44 and 1.52 points at 14× and 20×, the token-classification reformulation in LLMLingua-2, and LongLLMLingua’s reported performance improvement of up to 21.4 percent at roughly 4× fewer tokens on NaturalQuestions. The improvement result is the important one: it shows compression can raise quality by removing distractors, consistent with the signal-to-noise finding in reference 35. Limitation to carry, stated precisely: these are LongBench-family and question-answering benchmarks, not agentic coding trajectories, and the compressors were evaluated against models older than the current generation. They are not, however, silent on code—LLMLingua-2’s Table 2 and LongLLMLingua’s results both carry a LongBench Code column, and LongLLMLingua includes an LCC code-completion example—so an earlier revision of this document overstated the gap by saying nothing cited measured code. What remains unmeasured is long-horizon agentic repository editing and debugging, which is the workload §26.7 warns against transferring these trade-offs to.
  57. 57Yun-Hao Cao, Yangsong Wang, Shuzheng Hao, Zhenxing Li, Chengjun Zhan, Sichao Liu, and Yi-Qi Hu, “EFPC: Towards Efficient and Flexible Prompt Compression,” arXiv:2503.07956v1, 2025. Preprint. Cited for the re-tabulated LLMLingua-2 single-document QA baseline—35.5 at 3× against 29.8 at 5×, which is 5.7 score points and a 16.1 percent relative decline. Its Table 3 marks those rows as taken from Pan et al. (2024), which is reference 56, rather than re-measured here; an earlier revision of this document called them an independent measurement, and they are not one. What the two numbers support is a 5.7-point drop between those two ratios on that task. They do not support a steepening: two points fit a line as well as a curve, and an earlier revision of this entry claimed the curve steepens between 3× and 5× on their strength, which does not follow. The paper also proposes a competing method, so its choice of baseline is not disinterested, and neither the magnitude nor any knee location generalizes to other tasks, models and compressors.
  58. 58Practitioner surveys of the prompt-compression tool landscape, including PointFive, “Top 10 Prompt Compression Solutions (2026),”, and NeuralTrust, “Prompt Compression: Cut Token Costs Without Losing Quality,”, both verified September 8, 2026. Vendor-adjacent commercial content; both publishers sell related products, which should be stated. Cited only for the qualitative shape of practitioner opinion—useful reduction at moderate ratios, a quality cost at aggressive ones—and for the observation that extreme published ratios carry large accuracy drops. An earlier revision of this entry attributed a 30-to-70-percent input-reduction band with minimal accuracy loss, and a tradeoff threshold at roughly 80 percent, to these two pages; neither establishes that band, and the PointFive page now carries a different title and is an evaluation guide. The numeric band in §26.2 is the author’s own heuristic and is labeled as one. Treat these pages as a summary of practice rather than as measurement.
  59. 59Julius Brussee, Caveman, with docs/HONEST-NUMBERS.md and docs/WRAP-BENCHMARK.md, verified September 8, 2026, and the pipeline diagram docs/assets/pixel-pipeline.svg, pinned at commit bb1cfc8b5ce5 of the same date. Project documentation, self-reported. Source of the withdrawn 65 percent output ratio, the 33.2 percent fewer provider-reported input tokens in a pinned 54-run Claude Code benchmark passing 18 of 18 exact-answer checks, the per-content-type compression targets, the pixel-mode figures (55,413 estimated text tokens to 11,402 estimated image tokens on a dense payload of minified JSON and long-line logs, which the diagram labels −79 percent and inferred; a skill body from 1,069 to 415, which the README carries), the split MIT / BSL-1.1 licensing with automatic Apache-2.0 conversion, and the default-on anonymous telemetry with a documented opt-out. Cited here as a worked example rather than as an endorsement. The project is unusually rigorous about its own limits. The first of those two files, as corrected on September 8, 2026, states that the skill shrinks output tokens only, that no reviewed aggregate output-reduction result is published and the earlier fixed 65 percent ratio had no committed reviewed result behind it, that the input cost the skill adds is not measured there, that whole-session savings can go net negative on terse workloads, and that local results are labeled inferred rather than verified. That self-reporting discipline is the reason it is usable as an example; the numbers themselves remain vendor-reported and unreplicated. Two notes on retrieval, because the dense-payload pair is easy to conclude is missing. It sits in the text of a diagram rather than in prose, so reading the Markdown files above will not find it. And the project is actively developed: the pair is absent from the current revision of every file here, which is why the diagram is pinned to a commit rather than cited at main.
  60. 60Isaac Ong, Amjad Almahairi, Vincent Wu, Wei-Lin Chiang, Tianhao Wu, Joseph E. Gonzalez, M. Waleed Kadous, and Ion Stoica, “RouteLLM: Learning to Route LLMs with Preference Data,” ICLR 2025, arXiv:2406.18665. Peer-reviewed. Source of the over-85-percent cost reduction on MT-Bench at 95 percent of the strong model’s performance with 14 percent of queries routed to the frontier tier, the roughly 45 percent saving on MMLU with a classifier router, the cross-model-pair transfer finding, and the augmentation result: roughly 1,500 golden-labeled MMLU examples, described by the authors as less than 2 percent of the overall training data, added on top of roughly 65,000 public preference comparisons. Limitation, and it is load-bearing: the benchmarks are conversational, use a specific model pair, and make no claim about structured agent task completion.
  61. 61Lingjiao Chen, Matei Zaharia, and James Zou, “FrugalGPT: How to Use Large Language Models While Reducing Cost and Improving Performance,” arXiv:2305.05176, 2023. Preprint. Source of the cascade-routing approach and reported reductions up to 98 percent on the evaluated benchmarks. The headline figure is benchmark-specific and depends heavily on how skewed the query distribution is toward easy requests.
  62. 62Practitioner accounts of production routing deployments. Blog-level sources of varying rigor, verified September 8, 2026; cited for the shape of the cost-quality dial rather than as measurement. Unresolved attribution: an earlier revision of this entry reported a conservative band of 25 to 35 percent savings at roughly 99 percent quality retention, and classifier latency of 300 to 500 milliseconds on the request path. No author, deployment, URL, measurement window or denominator could be recovered for any of those three figures, so they have been removed from this entry and from the §27.7 latency row rather than retained as a typical setting. Measure routing savings, quality retention and router latency on your own deployment; nothing in this document’s routing recommendations depends on the withdrawn numbers.
  63. 63Analysis of routing failure modes specific to agent pipelines, verified September 8, 2026. Vendor-adjacent commentary rather than measurement, and cited as reasoning rather than evidence. The argument is mechanical and holds independently: a routing failure in a conversational product yields a worse answer the user re-asks, while a routing failure inside an agent loop yields a malformed tool call whose error propagates downstream and whose retry costs the frontier call anyway. The related point, that role-based structural routing captures most of the saving without classifier overhead or routing failure modes, is the basis for the recommendation in §27.4.
  64. 64GitHub, “Project HydraFusion: Frontier quality via multi-model orchestration,” The GitHub Blog, September 4, 2026, and the accompanying research-preview announcement, both verified September 11, 2026. Vendor announcement and vendor-published evaluation. Cited for the three composed execution patterns (single, cascade, and a critique pattern using an independent read-only critic from a different model family), for the statement that first-turn single-prompt tasks are the recommended starting point with multi-turn performance identified as future work, and for the cost-accounting methodology counting every invoked leg including drafting, critique, revision, escalation, retry, and fallback. The reported results—improved verified task quality at substantially lower estimated cost against frontier baselines—are GitHub’s own, measured against GitHub-selected baselines under GitHub-selected conditions, and have not been independently replicated. GitHub discloses that one of the three benchmarks is relatively saturated and that a second is internal and derived from its own Copilot sessions, which no external party can reproduce. Neither page carries adoption or volume statistics: an earlier revision of this document attributed a June 2026 request count and a share of paying users to them, and that attribution was wrong. Do not restore figures of that shape without a source that names the window and the denominator. The feature is an experimental research preview limited to the Copilot CLI at the time of writing; availability, pricing, execution patterns, and composed models are all subject to change.
  65. 65Jianlin Su, Murtadha Ahmed, Yu Lu, Shengfeng Pan, Wen Bo, and Yunfeng Liu, “RoFormer: Enhanced Transformer with Rotary Position Embedding,” Neurocomputing 568 (2024), arXiv:2104.09864. Peer-reviewed. Source of the long-term decay property of rotary position embeddings: reduced expected dot-product similarity between distant token pairs. That is a property of the construction and the leading mechanistic account of the positional findings in reference 20; it is not a proof that RoPE decay causes those findings, it says nothing about primacy, and neither paper publishes a universal attention-weight profile by position. §29.2 is scoped accordingly.
  66. 66Qwen team model card for Qwen3.6-35B-A3B, together with Unsloth’s local-deployment documentation for Qwen3-Coder, verified September 8, 2026. Vendor documentation and self-reported benchmarks. The parameter counts, Apache 2.0 licensing and reported benchmark figure come from the model card; the Unsloth page is deployment guidance and is not the source of any number here. Cited as product fact for the total-versus-active parameter naming convention, the Apache 2.0 licensing, and the availability of 35B-A3B and 480B-A35B configurations. Note that the card describes a mixed architecture, gated-linear layers interleaved with attention layers, which is why §29.1’s KV formula is labeled as full-attention or grouped-query arithmetic rather than as a universal law. The 73.4 percent verified-software-engineering-benchmark figure is the model authors’ own reported number, is not independently replicated, and should not be read as a prediction about any particular repository—see reference 55 on why public coding benchmarks travel badly.
  67. 67Provider prompt-optimization tooling, including OpenAI’s prompt optimizer () and open-source frameworks for bootstrapped demonstration selection and multi-stage pipeline optimization. Vendor documentation and open-source projects. No independent evaluation establishing generalization from optimized prompts to production distributions was located, which is why §35.3 and §35.4 treat held-out testing as mandatory rather than advisable.
  68. 68Published incident analysis of the compromised Nx build package (GHSA-cxm3-wv7p-598c, August 27, 2025) and subsequent vendor telemetry analysis, documenting malware that invoked locally installed coding agents, using the agents’ own filesystem access as the attack primitive. Attribute the two halves correctly, which an earlier revision of this document did not: the advisory’s prompt casts the agent as a file-search agent and asks for “a newline-separated inventory of full file paths,” explicitly adding “only list file paths — do not include file contents,” while the accompanying malware code performs the reading of those files and the upload of the results. Delivery was an npm postinstall hook, not a fetched README. Cited as an existence proof of AI-assisted supply-chain malware—that an installed agent’s privileges are a reachable attack surface—and not as a documented instance of the retrieved-README injection chain in §36.2, which it does not establish. First-party advisory plus vendor-published analysis; the refusal-rate findings in the telemetry analysis are uncorroborated and the publishing firm revised its own repository count upward mid-investigation. The scope of what it proves is stated above.
  69. 69Practitioner guidance on local inference hardware sizing, including the weights estimate of parameters × bits ÷ 8, 4-bit Q4_K_M as the standard quality-versus-memory compromise, and KV cache quantization requiring flash attention support. Community and blog-level sources verified September 8, 2026; the arithmetic is verifiable independently and the quantization-quality claim is a rule of thumb rather than a measurement. Two limits belong with it. The ideal 0.5 bytes per parameter at four bits is a floor: Q4_K-family files carry per-block scale and minimum metadata, so an actual GGUF is larger than the formula. And the flash-attention requirement for KV-cache quantization is a property of particular runtimes and build options rather than a universal one—name the runtime and check its flags. Reference 78 is the vLLM side of that question. Any quantization change should be treated as a model change and re-evaluated (§34.8).
  70. 70Anthropic, “Tool use with prompt caching,” Claude Platform documentation. Vendor documentation; cited as product fact for defer_loading, breakpoint placement on mcp_toolset entries, and the automatic 5-minute breakpoint applied to server tool results.
  71. 71OpenAI, “AGENTS.md,” Codex documentation. Vendor documentation; cited as product fact for Codex’s native AGENTS.md support and nesting behavior.
  72. 72GitHub, “Adding repository custom instructions for GitHub Copilot,” GitHub Docs, together with the surface-specific pages for the IDE and for Copilot CLI, verified September 8, 2026. Vendor documentation; cited as product fact for the three instruction types, the .github/instructions/NAME.instructions.md location, the applyTo glob frontmatter, the excludeAgent keyword, and instruction-file discovery order. Cited also for the limitation: path-specific instruction support is surface-dependent rather than universal, and the documentation is split across surface-specific pages rather than published as one specification, so behavior must be confirmed per surface.
  73. 73GitHub, “GitHub Copilot is moving to usage-based billing,” The GitHub Blog, April 27, 2026, and “What changed with Copilot billing (legacy),” GitHub Docs, both verified September 11, 2026. First-party announcement plus vendor documentation. Source of the June 1, 2026 transition from premium request units to GitHub AI Credits, the one-credit-equals-one-cent conversion, the calculation of consumption from input, output, and cached tokens at published per-model API rates, the unchanged seat pricing ($19 Business including $19 of credits, $39 Enterprise including $39), the pooling of credits at the organization level, the exclusion of code completions and Next Edit Suggestions from credit consumption, the replacement of the lower-cost-model fallback with balance and administrator budget controls, and the survival of model multipliers only as a legacy concept for annual Pro and Pro+ subscribers within an existing term. This is the most perishable product fact in the document: it replaced its predecessor roughly three months before publication, and that predecessor had itself been in place barely a year.
  74. 74GitHub, “Using custom instructions to unlock the power of Copilot code review,” GitHub Docs, verified September 8, 2026. Vendor documentation. Cited for the troubleshooting guidance naming instruction files over a thousand lines, vague or ambiguous instructions, conflicting instructions, and repository-wide rules that belong in path-specific files as causes of instructions not being followed—a vendor-side corroboration of the length and allocation discipline argued in §15.
  75. 75Cursor, “Rules,” Cursor documentation, verified September 8, 2026. Vendor documentation. Cited as product fact for the .mdc frontmatter schema, the four activation modes, and the interaction between .cursor/rules/ and AGENTS.md.
  76. 76Anthropic, “Structured outputs,” Claude Platform documentation, together with the Claude Sonnet 5 migration guide, and “Extended thinking,”, all verified September 25, 2026. Vendor documentation. Cited as product fact for four things the surrounding chapters depend on: that the structured-output API is constrained decoding and is documented in those words (§11.1); the documented capitalization caveat on enum, under which a normally completed response can return a casing the schema does not allow, with no special stop reason to signal it (§32.2); that thinking is on by default on Opus 5 and Sonnet 5, so a thinking block can precede the first text block and response parsing must select by block type rather than by index (§10.4); and that assistant prefill returns HTTP 400 on Sonnet 5, Opus 5 and the Fable and 4.6-through-4.8 families, with structured outputs or a system instruction as the documented replacement (§31.5). Numbered after reference 75 because these pages were added during a later revision; the reference numbering is append-only and no longer strictly follows first citation.
  77. 77OpenAI, “Function calling,” OpenAI API documentation, verified September 25, 2026. Vendor documentation. Cited for the custom-tools section, which accepts a context-free grammar in lark syntax or a regex to constrain a tool’s input, and which documents that the API rejects a grammar it considers too complex. This is the hosted counterexample §32.5 needed: grammar constraint over tool input is available without self-hosting, while an arbitrary grammar over the final assistant text is not.
  78. 78vLLM documentation: “Batch invariance,”, “Reproducibility,”, and “Offline inference,”, all verified September 25, 2026. Project documentation. Cited for two corrections. Batch-invariant serving is an opt-in mode with an enumerated list of tested models, DeepSeek and Qwen MoE families among them, which is why §30.4 describes hosted nondeterminism as a property of default serving rather than an impossibility in principle; the project also tracks configurations where the guarantee does not yet hold, so the mode is scoped rather than universal, and reproducibility additionally requires pinning kernels, sampling, hardware and versions. And vLLM exposes offline in-process inference through LLM.generate() and LLM.chat(), so it is not API-only (§37.4). Also cited, against the penalty implementation in vllm/model_executor/layers/utils.py, for the §31.2 distinction between the two count-based penalties: the frequency term multiplies an occurrence-count tensor and the presence term a binary occurrence mask, so only the former scales with repetition. An earlier revision of this document described both as scaled by how often the token occurred.
  79. 79Guangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han, and Mike Lewis, “Efficient Streaming Language Models with Attention Sinks,” ICLR 2024, arXiv:2309.17453. Peer-reviewed. Source of the attention-sink phenomenon in §29.4 and, importantly, of its limit: the paper attributes the effect to absolute position rather than to semantic content, reports that replacing the initial tokens with newlines preserves it, and introduces sink retention as its own proposed fix for window attention rather than describing existing practice. It therefore supports the eviction argument and does not support any inference that a substantive instruction gains processing by occupying those positions.
  80. 80Cursor, “Plugins,” Cursor documentation, and Gemini CLI, “Subagents,”, both verified September 25, 2026; and Gemini CLI, “Provide context with GEMINI.md files,” the project’s own documentation, verified September 27, 2026. Vendor documentation. Cited for the two §43.1 crosswalk cells that older comparisons show empty: Cursor packages rules, skills, agents, commands, hooks and MCP servers as plugins distributed through a reviewed marketplace and team marketplaces, and Gemini CLI defines subagents in .gemini/agents/*.md with automatic delegation and explicit @-invocation. The third source is what supports §43.1’s path-scoped Gemini cell, which the first two do not: it sets out the three-tier context hierarchy—global, workspace, and just-in-time—and documents that when a tool accesses a file or directory the CLI “automatically scans for GEMINI.md files in that directory and its ancestors up to a trusted root,” loading them only when needed. An earlier revision of this document made that claim against this entry before the supporting page was in it.
  81. 81Fireworks AI, “Grammar mode,” Fireworks documentation, verified September 25, 2026. Vendor documentation. Cited as product fact for §32.5’s hosted-grammar claim: grammar mode accepts a GBNF grammar—an extension of BNF with regex-like features, following llama.cpp’s implementation—passed in the request’s response-format field on any Fireworks model, and it constrains the generated assistant text rather than only a tool’s input. This is the entry that claim needs; reference 77 covers OpenAI’s tool-input grammars only, and an earlier revision of this document attached the Fireworks sentence to it.

Version 2.1, September 2026. This guide is offered as practitioner reference and does not constitute legal, security, or financial advice. Model prices were verified against primary sources on September 8, 2026; product behavior, billing models, and specification versions were re-verified on September 11, 2026; the data-residency controls in §5.2 and §5.3, and the multiplier that prices them, were verified on September 20, 2026; the OpenAI and Gemini per-model rates were re-verified on September 24, 2026; and the vendor and primary-source corrections applied in this revision—among them MCP authorization roles, Claude Code subagent permission behavior, Copilot instruction-surface support, the Cursor and Gemini CLI capability rows, OpenAI cache retention, Anthropic structured-output and prefill behavior, vLLM batch invariance and offline inference, the LLMLingua-2 baseline provenance, and the pinned TOON benchmark commit—were verified on September 25, 2026. All of it moves quickly, so re-verify before relying on any figure. Where a source is vendor-published, self-reported, correlational, or an unreviewed preprint, the reference entry says so. Practices marked [unmeasured] in the body have mechanical justification and no published measurement of effectiveness. Appendix D lists the events that would occasion a revision.

References cited in this section

1 of 81 · numbering matches the PDF

  1. 48Claude Code documentation (settings, hooks, sub-agents, and skills references), verified September 11, 2026, cross-checked against an independently compiled feature and settings snapshot at https://hidekazu-konishi.com/entry/claude_code_features_settings_reference_2026.html. Vendor documentation plus a third-party catalog that links each row back to the official docs. Cited for the settings precedence tree (user, project, project-local, CLI flags, enterprise managed, in ascending precedence, with the managed layer a floor that CLI flags cannot relax for scalar values and deny rules—list-valued keys such as permissions.allow and the sandbox allow and exclusion arrays merge across scopes instead, so lower scopes can add entries and widen access, which allowManagedPermissionRulesOnly exists to prevent for permission rules), the hook event catalog including PostCompact and its auto/manual matcher, the hook exit-code semantics, subagent frontmatter fields and isolation: "worktree", and the documented routing of subagent permission prompts—foreground subagents pass prompts through to the user, background subagents surface them in the main session naming the asking subagent, and auto-denial is a permission-mode behavior rather than a property of delegation. An earlier revision of this entry asserted that subagents cannot raise interactive prompts at all, so approval-required calls always resolve as denials; that was wrong, and §18.5 was corrected before this entry was. The same revision compressed the exit-code semantics to "0 allow, 1 allow with warning, 2 deny," which conflates the handler's process status with the event's decision, and the hook printed in §19.2 is the counterexample: it emits a permissionDecision of deny and exits 0. Exit 0 means the handler succeeded and Claude Code reads the decision from stdout JSON—silence is not approval, it is merely no decision, and the call continues through the normal permission flow. Exit 1 is a non-blocking error that Claude Code proceeds past, not a warning-flavored allow. Exit 2 blocks, but which events can block is event-specific: PreToolUse and UserPromptSubmit block, while PermissionRequest, PostToolUse, Notification, SessionStart and others do not honor it. Read the per-event table rather than a three-value mapping. Also cited, against the memory page and the v2.1.277 release notes of September 18, 2026, for native AGENTS.md loading and its conditions: by default Claude reads AGENTS.md only where no CLAUDE.md, .claude/CLAUDE.md or CLAUDE.local.md sits in the working directory or above it, while a user-level CLAUDE.md, a managed one and .claude/rules/ files do not count against it; a Project instructions setting in /config selects other modes, including loading both; and nested and subdirectory files load on access. The provider limitation this entry previously recorded as current is now version-scoped: the memory page places it before v2.1.281, published September 23, 2026, and directs affected Bedrock users to update rather than describing an ongoing platform gap. The verification date in this entry was accurate when made; this is a product change after it, not a correction to it. The conditions that remain current are an installation before v2.1.277, a disabled agents-md plugin, and in some cases the first session after an upgrade. An earlier revision of this document said Claude Code simply does not read AGENTS.md and presented the import line as a universal requirement; §15.2, §38.6, §43.1 and Appendix D were corrected together. The third-party snapshot is dated May 2026 and its model-name rows are consequently stale against the lineup in reference 12; the mechanism rows cited here were re-checked against the current official pages.docs.claude.com/en/docs/claude-code ↗
PDF↓