This part goes below the API surface. It assumes Parts I through IV and it does not repeat them.
Inside the Tokenizer
How BPE builds a vocabulary, the digit problem, character-level blindness, the non-English cost asymmetry, and encoding costs by format.
Attention, KV State, and the Geometry of a Context Window
The KV cache from the inside, why position and length both matter, attention sinks, and effective against advertised context.
Dense and Mixture-of-Experts Architectures
What separates dense from mixture-of-experts, what follows from it, the nondeterminism consequence, and when it should change a decision.
Sampling, Logprobs, and Decoding Control
From logits to a token, what each sampling parameter does, reading logprobs, calibration, stop sequences, and seeds.
Constrained Decoding and Grammar Masking
How grammar masking works, what it guarantees and what it does not, the measured reasoning cost, and the provider surfaces.
Post-Training, and Why Certain Phrasings Work
Three stages of post-training, what preference training explains, why format matters more than wording, and the status of magic words.
Evaluation
The two errors, building the set, grading, running it with variance, how many runs, regression gating, and what to evaluate beyond correctness.
Automated Prompt Optimization
What automated optimization is good at, what it is bad at, how to do it safely, and the honest position on what it returns.
Adversarial Mechanics
The core problem, the taxonomy, defenses ordered by strength, the layered implementation, prompt extraction, and how to test it.
Running Models Locally
The five reasons ranked by how often each is the real one, what is available, the hardware arithmetic, the serving stack, and the cost sheet.