Resources Version 2.1 · September 2026

Prompting Mechanics.

A technical reference on producing reliable output from language models and coding agents: prompting, context engineering, token economics, and the mechanics underneath.

Parts
6
Sections
45
References
81
Status
Released

The thesis

Prompting is what one writes, and context engineering is what arrives. The prompt itself is a small fraction of what the model receives. To produce the best results, one must understand the underlying mechanics of the overall system and each of the levers at play.

Everything one can control falls into one of two categories: the words in the request, and what else is in the window. Beginners spend nearly all their effort on the first, though the measurable wins are overwhelmingly in the second.

A third category sits outside both, and practitioners routinely forget it. On any coding agent, a layer between the user and the model rewrites the request before it is sent, and that layer measurably changes the answer.

Read Section 1 →

The two sets

Two sets run through every chapter: the first divides what a practitioner controls from what the harness decides before the model sees the request, and the second accounts for where every billed token lands.

Three surfaces
PROMPT
CONTEXT
HARNESS

A model answers from everything that reaches it, so context deserves at least the deliberate design the wording gets. The harness assembles and rewrites that context before the request is sent, and it is the one surface a practitioner does not own.

Four token lanes
CACHED
INPUT
WRITE
OUTPUT

Every billed token lands in exactly one lane, and the lanes differ in price by more than an order of magnitude. Optimization is the work of moving tokens toward the cheaper ones: keeping a prefix stable enough to be cached, and generating only the output the task actually requires.

The six parts

45 sections

Contents · Prompting Mechanics

v2.1
Part I — Foundations §1–7
01 Purpose and How to Read This 5 min 02 What a Prompt Actually Is 6 min 03 Tokens, and Why They Are the Unit of Everything 28 min 04 The Turn, the Loop, and the Session 6 min 05 Why the Same Prompt Gives You a Different Answer 6 min 06 The Harness 12 min 07 Choosing a Data Format 8 min
Part II — The Craft §8–13
08 Instruction Design 10 min 09 Examples, and When They Stop Helping 5 min 10 Eliciting Reasoning 10 min 11 Output Contracts 8 min 12 Assembling Context 5 min 13 Turning a Prompt Into an Artifact 3 min
Part III — The Environment §14–20
14 The Configuration Surface 6 min 15 AGENTS.md and the Instruction File Problem 10 min 16 Skills and Progressive Disclosure 6 min 17 Tools, Schemas, and MCP 9 min 18 Subagents and Delegation 8 min 19 Hooks: One Deterministic Control 4 min 20 Packaging and Distribution 2 min
Part IV — Economics and the Context Lifecycle §21–27
21 The Four Token Lanes 6 min 22 Prefix Caching, Mechanically 10 min 23 Invalidation: Everything That Breaks a Cache 7 min 24 Recycling Context 9 min 25 Budgeting and Instrumentation 9 min 26 Compression and Token-Reduction Tools 11 min 27 Model Routing 13 min
Part V — Advanced Mechanics §28–37
28 Inside the Tokenizer 7 min 29 Attention, KV State, and the Geometry of a Context Window 6 min 30 Dense and Mixture-of-Experts Architectures 5 min 31 Sampling, Logprobs, and Decoding Control 6 min 32 Constrained Decoding and Grammar Masking 9 min 33 Post-Training, and Why Certain Phrasings Work 5 min 34 Evaluation 9 min 35 Automated Prompt Optimization 4 min 36 Adversarial Mechanics 8 min 37 Running Models Locally 9 min
Part VI — Platforms §38–43
38 Claude and Claude Code 5 min 39 OpenAI and Codex 4 min 40 GitHub Copilot 7 min 41 Cursor 2 min 42 Gemini 4 min 43 Writing Once and Running Everywhere 9 min
Appendices and References §44–45
44 Appendices 6 min 45 References 49 min

The complete reference, typeset for print.

All 45 sections, the six platform chapters, the appendices, and 81 references in a single PDF.