Inside the Tokenizer
How BPE builds a vocabulary, the digit problem, character-level blindness, the non-English cost asymmetry, and encoding costs by format.
Give a working model of byte-pair encoding precise enough to predict and diagnose tokenization-induced failures rather than misattributing them to reasoning.
28.1 How BPE Builds a Vocabulary #
Section 3.1 walked through the training procedure using a runnable toy corpus. This section assumes it and goes to the properties that only matter once you are optimizing against the tokenizer rather than merely budgeting for it.
Production BPE starts from raw bytes rather than characters—256 base symbols, so any input at all is representable—and there is therefore no out-of-vocabulary case. Encoding applies the merges greedily in learned order and decoding is a table lookup. The property that governs everything downstream is that merge order is fixed at training time and determines every segmentation, which is why tokenization is deterministic while remaining thoroughly unintuitive.
Vocabulary size is a per-model parameter, not a constant, and the range has been moving. OpenAI’s current encoding is o200k_base at roughly 200,000; Qwen3.6-35B-A3B’s tokenizer defines 248,070 IDs—248,044 vocabulary entries plus 26 added tokens, contiguous rather than padded, counted from its tokenizer.json at revision 995ad96e, since a model repository is mutable and the count is only meaningful against a stated one—and Google lists Gemma 4 at a 262K vocabulary. Treat 100,000 to 200,000 as where the previous generation clustered rather than as a bound, and read the number off the model you are actually using. Larger vocabularies compress better, yielding fewer tokens per character, at the cost of a larger embedding table and softmax.
import tiktoken
enc = tiktoken.get_encoding("o200k_base")
for s in ["hello", "Hello", " hello", "hello world", "hllo"]:
ids = enc.encode(s)
print(f"{s!r:16} {len(ids)} token(s) {[enc.decode([i]) for i in ids]}")
# 'hello' 1 token(s) ['hello']
# 'Hello' 1 token(s) ['Hello'] ← a different token from 'hello'
# ' hello' 1 token(s) [' hello'] ← the leading space is inside the token
# 'hello world' 2 token(s) ['hello', ' world']
# 'hllo' 2 token(s) ['h', 'llo'] ← a typo fragments
The comments show decoded pieces rather than numeric IDs because the IDs are specific to one vocabulary and carry no meaning across models. What transfers is the segmentation.
Three things visible in that output that explain a lot of downstream behavior.
hello and Hello share no representation at the token level. The model learns their relationship through embeddings, but they are not the same input.
hello is one token, not two. This is why a prompt ending in a trailing space is awkward: the natural next token starts with a space, and you have already emitted it.
hllo becomes two low-frequency tokens. The model must reconstruct meaning from fragments it has seen rarely in that combination, which is part of why models handle typos less gracefully than humans expect.
28.2 The Digit Problem #
Number handling is where tokenization most visibly alters capability. Encoding a short series of numbers makes the problem legible in about six lines.
for n in ["7", "42", "1234", "12345", "1,234", "1 2 3 4"]:
ids = enc.encode(n)
print(f"{n!r:10} {len(ids)} tokens: {[enc.decode([i]) for i in ids]}")
# Run this yourself; the grouping differs between vocabularies. A typical result:
# '7' 1 tokens: ['7']
# '42' 1 tokens: ['42']
# '1234' 2 tokens: ['123', '4']
# '12345' 2 tokens: ['123', '45'] ← the split ignores place value
# '1,234' 3 tokens: ['1', ',', '234']
# '1 2 3 4' 7 tokens: ['1', ' ', '2', ' ', '3', ' ', '4']
The fourth line is the one to notice. 12345 splits as 123 + 45, a boundary falling three digits in for no reason connected to the number’s magnitude. Now consider what column alignment does to a table of figures, or what happens when the same quantity appears as 12345 in one place and 12,345 in another: those are different token sequences describing the same number, and the model has to recover the identity rather than read it.
The grouping bears no alignment to place value. 12345 might split as 123+45 or 1234+5, and neither corresponds to thousands and units. Arithmetic on such a representation requires the model to first recover place value from an arbitrary segmentation, and only then compute.
In practice this means:
- Digit-separated input helps. Presenting
1 2 3 4gives one token per digit with unambiguous position. This is why some numeric prompts work dramatically better with spaced digits. - Writing out the calculation helps. Long-form arithmetic converts a single hard step into several easy ones, each operating on freshly emitted, individually addressable digits.
- A calculator tool beats both. For anything that must be right, do not ask the model to compute. Give it a tool. This is the §10.4 oracle principle applied to arithmetic.
28.3 Character-Level Blindness #
A model cannot see the characters inside a token any more than a reader can see the individual phonemes in a word taken in at a glance. Counting letters, reversing strings, and checking for character membership are all operations on a representation where characters are not addressable.
Fails: How many 'r's are in "strawberry"?
Works: Spell "strawberry" one letter per line, then count the r's.
The second version forces the model to emit each character as its own token, which makes them addressable, and then counts on a representation it can actually see. This is not a trick; it is a change of representation.
The same reasoning explains why models struggle with rhyme in some languages, with acrostics, and with precise character-count constraints. Where you need character-level precision, either force the spelling step or use a tool.
28.4 Non-English and the Cost Asymmetry #
Tokenizers are trained on corpora dominated by English, and the compression ratio reflects that fact.
| Language | Approximate tokens for the same meaning |
|---|---|
| English | 1.0× (baseline) |
| Spanish, French, German | 1.2–1.5× |
| Russian, Greek | 1.5–2.5× |
| Chinese, Japanese, Korean | 1.5–2.5× |
| Thai, Hindi, Burmese | 2.5–4× |
Those ratios are order-of-magnitude guidance rather than measurement. They move with the tokenizer, and this document supplies no aligned corpus you could reproduce them from, so count your own with the script in §28.6 before budgeting on them.
The asymmetry itself holds all the same. A document in Thai can cost two to four times what its English equivalent costs to process, and consume that much more of the context window, which pushes it into the degraded region (§12.1) sooner.
For multilingual applications, two mitigations apply. Instructions can be in English even when content is not—the model handles the mix, and English instructions are cheaper. And context budgets must be set per language, not globally, or your Thai users hit window pressure at a third of the content your English users do.
28.5 Prompt Boundaries and Token Healing #
This level exposes one subtle mechanic. If a prompt ends mid-token—say it ends with "The answer is 4" and the natural continuation is 42—the model must continue from 4 as a separate token rather than emitting 42 as one. That path has different probabilities than the natural one.
Some inference stacks implement token healing: backing up one token from the end of the prompt and re-tokenizing the boundary so that the natural merge is available. Most hosted APIs do not expose control over this, but the phenomenon explains a class of odd behavior at prompt boundaries, particularly in completion-style or fill-in-the-middle usage.
The practical rule remains simple. End prompts at natural boundaries. No trailing whitespace, no partial words, no dangling punctuation that the model must complete.
28.6 Structured Data Encoding Costs #
For anyone assembling context programmatically, format choice has a measurable price. The script below encodes fifty records of the same shape three ways and reports what each costs.
import json, tiktoken
enc = tiktoken.get_encoding("o200k_base")
rows = [{"user_identifier": i, "display_name": f"User {i}",
"account_status": "active"} for i in range(50)]
pretty = json.dumps(rows, indent=2)
compact = json.dumps(rows, separators=(",", ":"))
csv = "id,name,status\n" + "\n".join(
f"{r['user_identifier']},{r['display_name']},{r['account_status']}"
for r in rows)
for label, s in [("json pretty", pretty), ("json compact", compact), ("csv", csv)]:
ids = enc.encode(s)
print(f"{label:14} {len(ids):>6} tokens {len(s)/len(ids):.2f} chars/token")
# Indicative magnitudes for 50 records; run it for exact numbers:
# json pretty ~1,500 tokens compact is ~43% smaller
# json compact ~850 tokens csv is ~73% smaller than pretty
# csv ~400 tokens
Dropping to CSV for tabular data saves roughly three-quarters. The cost is that CSV carries no type information and no nesting, so it is right for uniform tables and wrong for anything else.
The general guidance is this: repeated keys are the expensive part. Fifty records with three long key names each pay for 150 key repetitions. If the shape is uniform, state the shape once and send the values.
Format: id,name,status (one record per line)
1,User 1,active
2,User 2,active
...
The saving compounds with row count, which makes this a decision to make once in whatever code assembles your context rather than case by case. Three-quarters off a tabular payload is the largest single-line change available, and unlike most token reductions it costs nothing in quality—the information is identical, only the framing is gone.