Part V · §28–37

Advanced Mechanics

10 sections 68 min read

This part goes below the API surface. It assumes Parts I through IV and it does not repeat them.

§28 7 min

Inside the Tokenizer

How BPE builds a vocabulary, the digit problem, character-level blindness, the non-English cost asymmetry, and encoding costs by format.

§29 6 min

Attention, KV State, and the Geometry of a Context Window

The KV cache from the inside, why position and length both matter, attention sinks, and effective against advertised context.

§30 5 min

Dense and Mixture-of-Experts Architectures

What separates dense from mixture-of-experts, what follows from it, the nondeterminism consequence, and when it should change a decision.

§31 6 min

Sampling, Logprobs, and Decoding Control

From logits to a token, what each sampling parameter does, reading logprobs, calibration, stop sequences, and seeds.

§32 9 min

Constrained Decoding and Grammar Masking

How grammar masking works, what it guarantees and what it does not, the measured reasoning cost, and the provider surfaces.

§33 5 min

Post-Training, and Why Certain Phrasings Work

Three stages of post-training, what preference training explains, why format matters more than wording, and the status of magic words.

§34 9 min

Evaluation

The two errors, building the set, grading, running it with variance, how many runs, regression gating, and what to evaluate beyond correctness.

§35 4 min

Automated Prompt Optimization

What automated optimization is good at, what it is bad at, how to do it safely, and the honest position on what it returns.

§36 8 min

Adversarial Mechanics

The core problem, the taxonomy, defenses ordered by strength, the layered implementation, prompt extraction, and how to test it.

§37 9 min

Running Models Locally

The five reasons ranked by how often each is the real one, what is available, the hardware arithmetic, the serving stack, and the cost sheet.

PDF↓