Section 44 of 45 6 min read

Appendices

The one-page version, a glossary, a diagnostic index from symptom to cause, and the events that would occasion a revision.

Appendix A—The One-Page Version #

For a reader who will not get past this page.

Foundations. The prompt you write is a fraction of what the model reads. Order is tools, then system, then history, then your message. Everything before your message is prefix, and prefix is where cost and cache live.

The harness. On any coding agent, a layer between you and the model rewrites your request, chooses your tools, and selects a different system prompt depending on the model you picked. Attribute failures to your configuration first, the harness second, the model last.

Craft. Specify the deliverable, not the topic. Supply what the model cannot infer and delete what it can read. Give it an escape hatch for the empty case. Put the important material at the boundaries, never in the middle.

Verification. Verification against an external oracle works. Verification against the model’s own judgment does not. The highest-leverage change to any prompt is usually giving the agent something that can tell it when it is wrong.

Context. More is not better. Performance degrades with input length regardless of relevance, well before the window is full. Budget to §29.5’s reliable range—around a third of a 200K window, far less of a 1M one—not to a fixed share of the advertised figure.

The loop. Long sessions accumulate commitment, not understanding. Two failures means clear the context, not explain a third time. Write the handoff to a file.

Economics. Four lanes: cached read at 0.1× (0.025× or 0.05× on some models), base input at 1×, cache write at 1.25×, output at five or six. Caching is a prefix match that holds up to the first change. Put the breakpoint on the last block that does not change.

Configuration. Always-on mechanisms are subscriptions; on-demand mechanisms are purchases. Procedures go in skills, not instruction files. File-type rules go in globs. Rules that must hold go in hooks.

Security. Every string in the window is a potential instruction. Defense is capability restriction and deterministic enforcement, not instruction.

Evaluation. Thirty cases, run five times each, gated per case—and the gate is an alarm, not a test, so confirm what it flags before calling it a regression. Every production failure becomes a permanent test.

Appendix B—Glossary #

Attention sink Early tokens that absorb attention mass regardless of content; evicting them degrades performance sharply.

BPE Byte-pair encoding. The dominant tokenization scheme, though not the only one: WordPiece and Unigram are the other subword families in common use.

Cache breakpoint A marker designating the end of a cacheable prefix. Writes happen only at breakpoints.

Cache write / cache read Storing a prefix (1.25–2× base input) versus reusing one (0.1×).

Compaction Replacing conversation history with a summary. Reduces length, breaks cache, and loses instructions that existed only in the conversation; re-read files and re-sent system fields survive it (§24.3).

Constrained decoding Masking logits so invalid tokens are unreachable. Guarantees form on a completed generation, with documented exceptions (§32.2); degraded reasoning in the models and methods where it was measured (§32.3).

Context rot Degradation in output quality as input length grows, independent of window capacity.

Effective context The length at which a model actually performs reliably. A fraction of advertised, and not a constant one: it shrinks as the window grows, so read §29.5’s per-window ranges rather than applying a single ratio.

Fixed block The tokens present on every turn before any work: provider content, tool schemas, instruction files.

Grammar masking See constrained decoding.

KV cache Cached key and value tensors for a prefix. What prompt caching stores; not readable text.

Logprob Log probability of a token. One candidate confidence signal; like verbalized confidence, it needs calibrating against ground truth before a threshold means anything (§31.4).

Lost in the middle The U-shaped positional accuracy curve: performance is worse when the relevant information sits in the middle of the context than at either edge. It is a measured task result, not an attention-weight profile.

MCP Model Context Protocol. The tool-connection standard, now under the Agentic AI Foundation.

Oracle Something outside the model that can determine correctness: tests, compiler, type checker, schema.

Prefix Everything in the rendered context before the varying portion. The cacheable part.

Progressive disclosure Loading metadata always and content on demand. The mechanism behind skills.

Prompt cache key OpenAI’s prompt_cache_key. On models before GPT-5.6, a routing hint influencing which machine serves a request, and so hit rates under load; from GPT-5.6 routing is automatic and the key instead separates cache accounting (§22.7, §39.3).

Rendered context The flat token sequence the model actually receives.

Self-conditioning Increased error rate when the context contains the model’s own prior errors.

Skill A directory with frontmatter loaded always and a body loaded on demand.

Subagent A separate model execution with its own context window; only its final message returns. A fresh one starts empty and pays a cold prefix; a fork inherits the parent conversation and reuses its cache.

Termination reason The typed cause of loop exit. The most useful and most commonly discarded telemetry.

Token healing Re-tokenizing at a prompt boundary so natural merges remain available.

Appendix C—Diagnostic Index #

SymptomLikely causeSection
Output degrades late in a sessionSelf-conditioning, context length§4.3, §24
Model ignores a rule stated only in chatCompaction summarized it away§24.3
Cache hit rate near zeroBreakpoint on changing content§22.3
Cache hit rate fell under loadRouting overflow§22.7
Bill doubled with no usage changeCrossed a context threshold§21.3
Output is well-formed and wrongNo semantic validation layer§11.4
Reasoning got worse after adding a schemaFormat restriction§11.2, §32.3
Agent invents a root causeNo nullable path, no escape hatch§8.1, §11.3
Same prompt, different answersSampling plus infrastructure variance§5.2
Model cannot count charactersTokenization, not reasoning§28.3
Agent used a tool it should not haveTools list omitted or too broad§18.5
Instruction file made things worseContent the model could infer§8.6, §15.4
Retrieval “worked” but the answer is wrongGold file never accessed§12.3
Costs vary wildly across identical tasksCost is stochastic§25.1
Fix for one case broke three othersNo regression suite§34.6
Same prompt behaves differently in two productsDifferent harness, not a worse model§6.2
Quality changed and nothing on your side didHarness update, not only a model update§6.7

Appendix D—Revision Triggers #

This document should be revised when any of the following occurs:

  • A vendor changes its billing model, not merely its rates. GitHub’s replacement of premium request units with token-metered AI Credits on June 1, 2026 invalidated every word of §40.3 as previously written, and nothing about the rate card signaled it in advance.
  • The Model Context Protocol publishes a revision. The 2026-07-28 revision made the protocol core stateless and changed the transport contract; servers built against earlier revisions are not conformant without work.
  • A major provider changes cache multipliers, TTL semantics, or minimum cacheable length.
  • A provider adds or removes a context-length threshold.
  • AGENTS.md gains required fields, or Claude Code’s conditional native loading (§15.2) becomes unconditional.
  • The MCP specification revises the transport or authorization model.
  • A published study measures sub-agent context isolation on software engineering tasks—currently unmeasured (§18.3).
  • A published study measures selective context pruning against full reset—currently unmeasured (§24.6).
  • A detection heuristic for runaway agent loops is published with a measured false-positive rate—currently unmeasured (§25.4).
  • Constrained-decoding reasoning degradation is re-measured on current-generation reasoning models. The governing study predates them.34

All prices in this document were verified on September 8, 2026, apart from the data-residency multiplier in §5.3, verified on September 20, 2026, and the OpenAI and Gemini per-model rates, re-verified on September 24, 2026. Re-verify any of them before use in a business case.

References cited in this section

1 of 81 · numbering matches the PDF

  1. 34Zhi Rui Tam, Cheng-Kuang Wu, Yi-Lin Tsai, Chieh-Yen Lin, Hung-yi Lee, and Yun-Nung Chen, "Let Me Speak Freely? A Study on the Impact of Format Restrictions on Performance of Large Language Models," Proceedings of EMNLP 2024, Industry Track, pages 1218–1236, DOI 10.18653/v1/2024.emnlp-industry.91, arXiv:2408.02442. Peer-reviewed. Establishes that format restrictions degrade reasoning while improving classification accuracy. Cited for that direction and not for a monotonic strictness ordering, which an earlier revision of this document reported as constrained-decoding > format-restricting-instructions > natural-language-then-convert. The published Table 2 contradicts it: on gpt-4o-mini the JSON-Schema condition, which is the stricter constraint, scores 91.71 / 81.77 / 86.07 on GSM8K, Shuffled Objects and Last Letter against the format-restricting instruction's 87.17 / 81.46 / 84.73, and on Last Letter exceeds the natural-language mean of 83.11. The JSON-mode condition is the worst in all three rows, so the study's poor JSON-mode results are what the ordering was built on, and they do not generalize to schema-constrained decoding. The reported standard deviations are wide and these comparisons are not tests of pairwise significance; they are sufficient to refute a universal ordering, not to establish the reverse one. Two further limitations are load-bearing: the study predates current-generation reasoning models, and the authors have published updates in response to methodological critique, which do not establish an intrinsic monotonic penalty either. The direction is well established; the effect size on current models is unconfirmed.arxiv.org/abs/2408.02442 ↗
PDF↓