Tokens, and Why They Are the Unit of Everything
What a token is, why estimating instead of measuring goes wrong, where token boundaries change behavior, and the reasoning-token line item.
Give you an accurate working model of tokens, enough arithmetic to estimate cost and window pressure without a calculator, and the specific places where token boundaries change model behavior.
3.1 What a Token Is #
A token is a unit of text that a model’s tokenizer maps to an integer ID in a fixed vocabulary. Its boundaries land in neither word nor character positions, and they differ between vendors. The dominant family for the models in this document is byte-pair encoding, which builds that vocabulary by repeatedly merging frequently co-occurring byte sequences into single units, and it is the one this section walks through—though it is a choice of algorithm rather than the definition of a token, and WordPiece and Unigram, which do not use that pair-merging procedure, are the other two in common use. Consequently, common English words usually cost one token, while rare words, identifiers, and most non-Latin scripts break into several.
The rules of thumb that hold well enough for estimation:
| Content | Tokens per unit |
|---|---|
| English prose | ~1.3 tokens/word, ~0.75 words/token, ~4 characters/token |
| Source code | ~0.4–0.5 tokens/character—denser than prose |
| JSON with verbose keys | ~0.3 tokens/character—worst case |
| Minified JSON | ~0.4 tokens/character |
| CJK text | ~1–1.5 tokens/character |
| Base64 / hashes / UUIDs | ~0.3 tokens/character—near worst case |
For anyone assembling context programmatically, a 200 KB JSON payload is not “about 50,000 tokens.” It runs closer to 65,000, and where the keys are long and repetitive it can exceed 70,000. Therefore, estimate on the pessimistic side, or measure.
3.1.1 How the Vocabulary Is Built
Before a model is trained, someone runs a program over a large corpus to decide what the token vocabulary will be. That program starts with nothing but individual characters and repeatedly merges whichever adjacent pair occurs most often. The output is an ordered list of merges, and that list is frozen forever.
The process is small enough to watch end to end, using a four-sentence corpus:
The cache stores the prefix. The prefix is the cached part.
Caching the prefix makes the next request cheaper.
A cached request costs less than an uncached request.
First the trainer counts words. Twenty-eight words, nineteen of them distinct:
the 4 · prefix 3 · request 3 · The 2 · cached 2 · cache 1 · stores 1
is 1 · part 1 · Caching 1 · makes 1 · next 1 · ... (7 more, once each)
Then it splits every word into characters, counts every adjacent pair across the whole corpus, and merges the most frequent one. Then it does that again, as many times as you tell it to.
That merge count is the single parameter of the entire procedure, and it is most of what practitioners mean by vocabulary size—the published figure also counts the initial single-character symbols and any special tokens, so it runs a few hundred above the merge count rather than equaling it. Twenty merges are used below, which is enough for the interesting behavior to appear and few enough to read in one screen. The choice is revisited with measurements once the mechanism is clear. Real output follows:
merge 1: 'h' + 'e' -> 'he' (seen 11 times)
merge 2: 'he' + '</w>' -> 'he</w>' (seen 7 times)
merge 3: 'r' + 'e' -> 're' (seen 7 times)
merge 4: 'a' + 'c' -> 'ac' (seen 5 times)
merge 5: 's' + 't' -> 'st' (seen 5 times)
merge 6: 's' + '</w>' -> 's</w>' (seen 5 times)
merge 7: 'c' + 'ac' -> 'cac' (seen 4 times)
merge 8: 't' + 'he</w>' -> 'the</w>' (seen 4 times)
merge 9: 'p' + 're' -> 'pre' (seen 3 times)
merge 10: 'pre' + 'f' -> 'pref' (seen 3 times)
merge 11: 'pref' + 'i' -> 'prefi' (seen 3 times)
merge 12: 'prefi' + 'x' -> 'prefix' (seen 3 times)
merge 13: 'prefix' + '</w>'-> 'prefix</w>' (seen 3 times)
merge 14: 'cac' + 'he' -> 'cache' (seen 3 times)
merge 15: 'cache' + 'd' -> 'cached' (seen 3 times)
merge 16: 'cached' + '</w>'-> 'cached</w>' (seen 3 times)
merge 17: 're' + 'q' -> 'req' (seen 3 times)
merge 18: 'req' + 'u' -> 'requ' (seen 3 times)
merge 19: 'requ' + 'e' -> 'reque' (seen 3 times)
merge 20: 'reque' + 'st' -> 'request' (seen 3 times)
The first few merges show the procedure clearly, and the arithmetic checks against the corpus at every step.
Merge 1 finds h followed by e eleven times, more than any other pair. The eleven are the four times, The twice, cached twice, and cache, cheaper and uncached once each. So he becomes a vocabulary entry. Caching is not among them: it contains hi, not he, which is a useful reminder that the trainer is matching literal character pairs and has no idea cache and Caching are the same word.
Merge 2 asks where those eleven sit. Seven of them fall at the end of a word—the four times, The twice, cache once—so word-final he earns an entry of its own. The other four are word-internal and stay as plain he.
By merge 8 the trainer has t and he</w> available and can assemble the</w> in one step. The most common word in the corpus is now a single token, and it took eight merges to get there.
</w> marks the end of a word. It exists so that the tokenizer can distinguish cache standing alone from cache embedded within cached, which is the same office that leading spaces perform in production vocabularies.
Merges 9 through 13 assemble prefix, and 17 through 20 assemble request. Both start from a piece the trainer already had: merge 9 is p + re and merge 17 is re + q, and re was earned at merge 3 by appearing seven times across the corpus. request gets a second reuse at merge 20, where st from merge 5 finishes it in one step instead of two. Reuse is the mechanism working as intended.
What it costs is still the point. prefix consumes five merges, 9 through 13, the last of which attaches its end-of-word marker. request consumes four, 17 through 20, and runs out of budget before its marker is merged at all—which is why the encoder in the next subsection still splits it. Every merge spent on one word is a merge not available to another, and at twenty merges the second long word is already going short.
3.1.2 What the Merge List Does to Real Words
Encoding applies the previous twenty merges, in order, to whatever you hand it:
the -> ['the</w>'] 1 token
The -> ['T', 'he</w>'] 2 tokens
prefix -> ['prefix</w>'] 1 token
cached -> ['cached</w>'] 1 token
request -> ['request', '</w>'] 2 tokens
cache -> ['cac', 'he</w>'] 2 tokens
uncached -> ['u', 'n', 'cached</w>'] 3 tokens
Caching -> ['C', 'ac', 'h', 'i', 'n', 'g', '</w>'] 7 tokens
Four results in that list recur at production scale, and each warrants a closer look.
the costs one token and The costs two. They are different character sequences and therefore different entries. The corpus contained the four times and The only twice, so the earned a merge while The did not. Capitalizing a word can cost a token, and the model has no notion that the two are the same word until training teaches it.
request appeared three times and still costs two tokens. Merge 20 assembled the letters, and the trainer exhausted its budget before it could attach </w>. One further merge would have made it a single token. Where the vocabulary stops is arbitrary from the practitioner’s vantage and permanent from the model’s.
cache costs two tokens even though cache is a vocabulary entry. Merge 14 created cache, but merge 2 had already created he</w>, and merges apply strictly in order. By the time the encoder reaches merge 14, no bare he remains to attach to cac. The entry exists and the word cannot reach it, and the only route to one token is a later merge of cac+he</w> built from scratch—which this budget never reaches and §3.1.5 shows arriving at merge 26. The mechanism is working as designed, and it is the reason tokenization remains deterministic while staying impossible to predict by eye.
3.1.3 Being in the Corpus Is Not Enough; Frequency Is What Counts
Caching is in the corpus, and it is the most expensive word in the list at seven tokens. It appeared once, and one occurrence never wins a pair contest.
Seven is the surprising number here because the obvious objection is that merge 4 created ac, so Caching should surely reach ach and land at six. It does not, and tracing why is the clearest look at merge order in this whole example.
After merges 1 and 2 turned he and word-final he</w> into entries, the cac family looked like this:
cache c ac he</w>
cached c ac he d </w>
uncached u n c ac he d </w>
Caching C ac h i n g </w> <- the only bare 'h' left
The h in cache, cached and uncached was swallowed into he at merge 1 because in those words an e follows it. In Caching an i follows, so its h survives alone. That leaves the pair ac+h with a count of exactly one, against c+ac at four. It never wins, ach is never created, and Caching keeps a loose h forever.
Furthermore, the capital costs it a second time. Merge 7 built cac from lowercase c+ac, and Caching starts with C, so it cannot use that entry either. Two accidents of spelling, and a word that shares an obvious root with cache and cached to any human shares almost nothing with them in the vocabulary.
3.1.4 Words the Trainer Never Saw Still Encode
Neither of these appears anywhere in the corpus:
prefixes -> ['prefix', 'e', 's</w>'] 3 tokens, reusing the whole-word piece
recache -> ['re', 'cac', 'he</w>'] 3 tokens, reusing two learned fragments
Nothing is unrepresentable, since the fallback descends all the way to individual characters. Common input is cheap, rare input is expensive, and impossible input does not exist. This property is what makes the entire approach work: a vocabulary fixed years ago still handles a variable name invented this morning.
A production vocabulary is this exact procedure run to a few hundred thousand entries over a corpus of internet text. Every property in the following subsections follows from it.
Below is the trainer that produced every number above, in full. Change the corpus and watch the vocabulary rearrange.
import re
from collections import Counter
CORPUS = """The cache stores the prefix. The prefix is the cached part.
Caching the prefix makes the next request cheaper.
A cached request costs less than an uncached request."""
def train(word_counts: Counter, num_merges: int):
"""Learn a merge list. Every word starts as characters plus an end marker."""
vocab = {" ".join(w) + " </w>": c for w, c in word_counts.items()}
merges = []
for step in range(num_merges):
pairs = Counter()
for word, freq in vocab.items():
syms = word.split()
for i in range(len(syms) - 1):
pairs[(syms[i], syms[i + 1])] += freq # count adjacent pairs
if not pairs:
break
best, count = pairs.most_common(1)[0] # most frequent wins
merges.append(best)
# The lookarounds matter: a plain str.replace would rewrite 'cac he</w>'
# into 'cache</w>', merging into the middle of a symbol the encoder
# treats as one piece.
pat = re.compile(r"(?<!\S)" + re.escape(" ".join(best)) + r"(?!\S)")
vocab = {pat.sub("".join(best), w): c for w, c in vocab.items()}
print(f" merge {step+1:2}: {best[0]!r} + {best[1]!r} -> "
f"{''.join(best)!r} (seen {count} times)")
return merges
def encode(word: str, merges) -> list[str]:
"""Apply the frozen merge list, in learned order, to any word."""
syms = list(word) + ["</w>"]
for a, b in merges:
i = 0
while i < len(syms) - 1:
if syms[i] == a and syms[i + 1] == b:
syms[i:i + 2] = [a + b] # merge in place
else:
i += 1
return syms
counts = Counter(re.findall(r"[A-Za-z]+", CORPUS))
merges = train(counts, 20)
for w in ["the", "The", "prefix", "cached", "request", "cache", "uncached", "Caching",
"prefixes", "recache"]:
print(f" {w:10} -> {encode(w, merges)}")
3.1.5 Why Twenty, and What the Number Actually Controls
Twenty was a legibility choice. The merge count is the only knob on this procedure, so the useful question is what different settings buy.
Below, the same corpus is trained at seven different budgets. Each cell is the number of tokens that word costs to encode at that budget. The last column is the cost of encoding all 28 words of the corpus, counting repeats, so it is the number you would actually be billed for.
The table rests on two facts about the encoding. Every word ends with the </w> marker, and that marker is itself a token until something merges it away, so at zero merges a three-letter word like the costs four: t, h, e, </w>. And the merges run out at 63 because after that, every distinct word in the corpus has been built into a single entry, and there are no adjacent pairs left to merge.
| Merges | the lc | The cap | prefix | cached | request | cache | Whole corpus |
|---|---|---|---|---|---|---|---|
| 0 | 4 | 4 | 7 | 7 | 8 | 6 | 161 |
| 5 | 2 | 2 | 6 | 5 | 6 | 3 | 126 |
| 10 | 1 | 2 | 4 | 4 | 6 | 2 | 107 |
| 20 | 1 | 2 | 1 | 1 | 2 | 2 | 77 |
| 30 | 1 | 1 | 1 | 1 | 1 | 1 | 61 |
| 40 | 1 | 1 | 1 | 1 | 1 | 1 | 51 |
| 63 | 1 | 1 | 1 | 1 | 1 | 1 | 28 |
The first row represents the state before any merges exist. With no vocabulary beyond single characters, every word is spelled out letter by letter, plus its end marker: the is 4 tokens, request is 8, and the whole corpus costs 161 tokens to say 28 words.
The last row, by contrast, represents the opposite extreme. At 63 merges every distinct word has earned its own entry, so all 28 words cost exactly 1 token each and the same text costs 28. Reaching that floor takes three merges more than the shape of the corpus suggests because cache spends one of them undoing the shadowing described in §3.1.2.
Between those rows is the whole range available on this corpus: 161 tokens down to 28, a factor of 5.75. The row in bold is the twenty-merge budget used in the walkthrough above, which lands at 77—roughly halfway, which is why it shows some words resolved and others still in pieces.
Three things in that table carry forward into production reasoning.
The returns diminish sharply. The first twenty merges save 84 tokens, while the next forty-three save only 49. Frequent sequences are absorbed first by construction; therefore, early merges buy considerably more than late ones. This is why real vocabularies stop where they do rather than continuing: past some point each additional entry earns almost nothing while still costing an embedding row and a slot in the output softmax. Where that point falls is a per-model decision, and current vocabularies are larger than the 100,000-to-200,000 band the previous generation settled into—Qwen3.6-35B-A3B defines 248,070 IDs and Gemma 4 is listed at 262K.
Words become cheap in frequency order. the reaches one token by merge 10, prefix and cached by merge 20, and request by merge 30. The arrives only at merge 30 because it appeared half as often as the. Consequently, the cost of any content is set by how far its vocabulary overlaps with whatever the trainer saw most.
A shadowed word pays twice. cache stays at two tokens from merge 14 until merge 26, when the pair cac+he</w> finally wins a contest of its own and rebuilds the word the encoder could not assemble from the entries it already had (§3.1.2). Frequency alone does not predict when a word becomes cheap; merge order does, and a word the trainer has effectively spelled twice has consumed two merges that a differently ordered vocabulary would have spent elsewhere.
Scale the whole table up by four orders of magnitude and you have the real decision a lab makes. A larger vocabulary means fewer tokens per character, which means more text fits a context window and every request costs less—paid for with a larger embedding table, a larger softmax, and more parameters spent on representing rare strings. §28.1 picks that trade-off up on the production side.
3.1.6 What the Model Actually Receives
Each vocabulary entry carries a number, and encoding produces a list of those numbers. The model never sees characters at all; it sees integers. The same string always produces the same integers, which is why tokenization remains deterministic even though generation does not (§5).
A common confusion lives here, and one sentence surfaces it:
The quick brown fox jumps over the lazy dog.
| Piece | Same vocabulary entry as… | Why |
|---|---|---|
The | — | capital T, no leading space |
quick | — | leading space is inside the token |
the | not The | different characters entirely |
dog | — | |
. | — | punctuation is its own entry |
The word “the” appears twice in that sentence and costs two tokens, but they are not the same token. The at the start and the in the middle differ in capitalization and in whether a space is attached, so they are two separate entries in the table with two different numbers. The model learns they are related through training; the tokenizer does not know they are related at all.
That sentence therefore costs ten tokens for nine words (eight distinct). English prose lands near one token per word because English is what the vocabulary was optimized for.
3.1.7 Why Other Content Costs More
Everything else constitutes a departure from that optimum, and the reason is invariably the same: sequences rare in the corpus never earned a merge.
Source code breaks into far more pieces because identifiers are rarer in the training corpus than English words are:
def get_user_by_id(user_id: int) -> User:
pieces: def | get | _user | _by | _id | ( | user | _id | : | int | ) | -> | User | :
Roughly 14 tokens for 41 characters, or about 2.9 characters per token against prose’s 4. Two factors drive this. First, identifiers split at word boundaries the tokenizer learned from English, so getUserById becomes get+User+By+Id—four tokens where a human sees one name. Second, punctuation is individually expensive: (, ), :, and -> constitute four tokens carrying no information the model could not have inferred from position.
The remaining examples all carry the same two records, five fields each, covering an integer, two strings, a boolean and a float. Only the syntax around the data changes between them.
JSON is the worst of the common formats, and the reason is repetition:
[
{
"id": 4821,
"service": "auth-api",
"severity": "sev2",
"open": true,
"latency_p95": 812.4
},
{
"id": 4822,
"service": "billing-api",
"severity": "sev3",
"open": false,
"latency_p95": 204.1
}
]
244 characters for two records. Every field name is spelled out once per record, every string is quoted, and the braces, brackets, colons and commas carry a real share of the count—though not one token apiece, since the tokenizer merges punctuation with the whitespace beside it and emits pieces like {\n, ", ": and ,\n. Across fifty records, therefore, the five key names are paid for fifty times over while saying nothing new after the first.
TOML drops the braces and commas, leaves keys unquoted, keeps quoting on strings, and puts spaces around the delimiter, which the tokenizer renders as pieces like = and ":
[[incidents]]
id = 4821
service = "auth-api"
severity = "sev2"
open = true
latency_p95 = 812.4
[[incidents]]
id = 4822
service = "billing-api"
severity = "sev3"
open = false
latency_p95 = 204.1
193 characters. Bare keys and the absence of braces and commas are where it beats JSON, which quotes every key; the brackets it does spend go on one table header per record rather than around every value. String values stay quoted, which is where it loses to YAML.
YAML drops the quoting as well, keeping only the colon and indentation:
- id: 4821
service: auth-api
severity: sev2
open: true
latency_p95: 812.4
- id: 4822
service: billing-api
severity: sev3
open: false
latency_p95: 204.1
167 characters. The field names are still repeated per record, which is the remaining redundancy and the one the last two formats attack.
CSV states the field names once and then never repeats them:
id,service,severity,open,latency_p95
4821,auth-api,sev2,true,812.4
4822,billing-api,sev3,false,204.1
100 characters, and the gap widens with every added record since the header is paid only once. What CSV surrenders, however, is everything that is not a flat table. It has no nesting, no types, and no unambiguous null. The boolean true and the string "true" both arrive as the same four-character field, and the model is left to guess which was meant.
TOON is that same table with the missing guarantees added back. It was designed for this problem, and it encodes the identical data as:4
incidents[2]{id,service,severity,open,latency_p95}:
4821,auth-api,sev2,true,812.4
4822,billing-api,sev3,false,204.1
119 characters. The field contents of each row are identical to CSV’s; what TOON adds is the richer header plus two spaces of indentation on every row. [2] declares how many rows follow and {...} declares the field list, so a model can confirm it received the whole table rather than a truncated one.
Those two costs scale differently, and the difference matters more than the headline number. The header is paid once—fifteen characters here—and stops growing. The indentation is paid per row, forever. Of the nineteen extra characters in the two-record example, fifteen are header and four are indentation; at fifty records the header contributes sixteen and the indentation one hundred, so TOON runs 1,752 characters against CSV’s 1,636—about seven percent more, which sits inside the five to ten percent overhead the format’s own documentation claims.4 The overhead converges on a flat two characters per row rather than amortizing away.
3.1.8 What Happens When the Data Is Not Flat
Everything above is a flat table, which is the easy case and the one where CSV wins. Real tool results are rarely flat. Take the same incidents wrapped in the query that produced them, with a tag list alongside:
{
"query": {
"service": "auth-api",
"window": "24h"
},
"tags": [
"auth",
"regression"
],
"incidents": [
{
"id": 4821,
"severity": "sev2",
"open": true,
"latency_p95": 812.4
},
{
"id": 4822,
"severity": "sev3",
"open": false,
"latency_p95": 204.1
}
]
}
343 characters, and it now contains three different shapes: a nested object, an array of primitives, and an array of uniform objects.
YAML handles all three with indentation, at 222 characters:
query:
service: auth-api
window: 24h
tags:
- auth
- regression
incidents:
- id: 4821
severity: sev2
open: true
latency_p95: 812.4
- id: 4822
severity: sev3
open: false
latency_p95: 204.1
TOON treats each shape differently, which is the whole design, at 156 characters:
query:
service: auth-api
window: 24h
tags[2]: auth,regression
incidents[2]{id,severity,open,latency_p95}:
4821,sev2,true,812.4
4822,sev3,false,204.1
The nested object receives YAML-style indentation. The primitive array collapses onto a single line with a declared count. Only incidents, the one array of uniform objects, receives the tabular treatment. That selective behavior is what the format’s documentation calls tabular eligibility, and it is why TOON’s advantage shrinks as data gets less uniform.
CSV cannot represent this at all. The closest it gets is a lossy flattening:
query.service,query.window,tags,id,severity,open,latency_p95
auth-api,24h,"auth;regression",4821,sev2,true,812.4
auth-api,24h,"auth;regression",4822,sev3,false,204.1
165 characters, and three properties were destroyed in the process. The nesting became a dotted naming convention that the model has to infer. query is now duplicated on every row, so it grows with the incident count even though there is only one query. And the tag array was crushed into a semicolon-delimited string inside a quoted field, which is no longer an array at all. And it is not even cheaper. The flattened CSV is 165 characters against TOON’s 156, and the gap widens with every row: at fifty incidents it is 2,685 against 1,285, more than double, since the duplicated query is paid fifty times. On flat data, CSV wins on size; on nested data, it loses on size and fidelity.
That is the distinction in one comparison. On flat tables, CSV is smaller, and TOON’s header is pure overhead. The moment anything nests, CSV stops being a competitor and starts being a lossy export format.
The project publishes its benchmarks in two tracks, split precisely along that line, and the split is more informative than the headline.4
On flat tabular data, where CSV applies, CSV totals 63,997 tokens against TOON’s 67,778—TOON about six percent more expensive, buying the count and field declarations. Both are far below JSON at 164,452.
On mixed and nested structures, where CSV is excluded because it cannot represent them, TOON totals 227,830 tokens against formatted JSON’s 291,711, a 21.9 percent saving. But against compact JSON at 198,546, TOON is 14.7 percent worse. On the deeply nested configuration dataset specifically, compact JSON wins by eleven percent, and on semi-uniform event logs by twenty.
That, then, is the shape of it. TOON beats formatted JSON everywhere, beats compact JSON only where arrays are uniform enough to go tabular, and never beats CSV on data CSV can hold. The accuracy figures move less than the token figures: 76.4 percent against JSON’s 75.0 across four models, with the same 209 questions.4 Every figure in this subsection comes from one snapshot of a benchmark the project re-runs, and the project has since published a larger suite with different numbers; §7.7 is about exactly this, and the reference pins which run these came from.
Three caveats belong with all of it. These are the project’s own benchmarks. They measure the model reading data, which the documentation states explicitly, and independent work testing TOON for generation found the picture less favorable.5 And the structural guardrails do not always deliver: on a dataset with three rows removed from the end, TOON’s declared [N] count should have made truncation obvious, and it scored zero out of four while CSV scored four out of four.4
The same split shows up in the two payloads used above, measured at two records and at fifty. Flat first, the incident table on its own:
| Format | 2 records | 50 records | vs CSV |
|---|---|---|---|
| JSON | 244 | 6,052 | 3.7× |
| TOML | 193 | 4,849 | 3.0× |
| JSON compact | 171 | 4,251 | 2.6× |
| YAML | 167 | 4,199 | 2.6× |
| TOON | 119 | 1,752 | 1.1× |
| CSV | 100 | 1,636 | 1.0× |
Then the nested payload, the same incidents wrapped in a query object with a tag array, where CSV can only manage the lossy flattening shown above:
| Format | 2 records | 50 records | vs CSV |
|---|---|---|---|
| JSON | 343 | 5,359 | 2.0× |
| TOML | 224 | 3,800 | 1.4× |
| YAML | 222 | 3,606 | 1.3× |
| JSON compact | 215 | 3,215 | 1.2× |
| CSV (lossy) | 165 | 2,685 | 1.0× |
| TOON | 156 | 1,285 | 0.5× |
CSV moves from cheapest to dearer than TOON the moment nesting appears, since the wrapper is duplicated on every row, and it gives up fidelity to get there. TOON moves in the opposite direction, from a 1.1× overhead on CSV to half its size. Furthermore, it is the only format in either table that is simultaneously lossless and smaller than the flattening. That inversion is the whole argument for the format, and it is invisible on flat data.
These are characters rather than tokens, since exact token counts depend on the vocabulary, and the ratios are close but the ordering does not survive intact. Counting the same payloads in o200k_base, compact JSON overtakes YAML in both tables; and in the nested table at two records the flattened CSV overtakes TOON, 61 tokens against 66, even though TOON is nine characters shorter. At fifty records TOON wins that pair by a wide margin. Which reorderings appear is therefore a property of the payload size as much as of the format, and the script below gives the real token figures for whichever model you use.
The rule this produces: you are paying for repeated structure, not for information. That is why §28.6 treats format choice on a large payload as a larger lever than anything you can do to the prose around it.
Measure rather than trust the ratios. Every number above depends on which vocabulary a given model uses. This script gives you the real figures for yours:
# pip install tiktoken (OpenAI vocabularies)
import json, tiktoken
enc = tiktoken.get_encoding("o200k_base")
def show(label: str, text: str) -> None:
ids = enc.encode(text)
pieces = [enc.decode([i]) for i in ids]
print(f"{label:12} {len(ids):>5} tokens {len(text)/len(ids):.2f} chars/token")
print(" " + " | ".join(pieces[:18]) + (" ..." if len(ids) > 18 else ""))
FIELDS = ["id", "service", "severity", "open", "latency_p95"]
def rec(i):
return {"id": 4821 + i, "service": "auth-api" if i % 2 else "billing-api",
"severity": "sev2" if i % 2 else "sev3", "open": bool(i % 2),
"latency_p95": 812.4 - i}
def lit(v): # unquoted scalar form
return str(v).lower() if isinstance(v, bool) else str(v)
N = 50
flat = [rec(i) for i in range(N)] # a plain table
nested = { # the same rows, wrapped
"query": {"service": "auth-api", "window": "24h"},
"tags": ["auth", "regression"],
"incidents": [{k: v for k, v in r.items() if k != "service"} for r in flat],
}
nfields = [f for f in FIELDS if f != "service"]
show("prose", "The quick brown fox jumps over the lazy dog.")
show("code", "def get_user_by_id(user_id: int) -> User:")
# --- flat payload -----------------------------------------------------------
show("json", json.dumps(flat, indent=2))
show("compact", json.dumps(flat, separators=(",", ":")))
show("toml", "\n".join("[[incidents]]\n" + "\n".join(
f'{k} = ' + (f'"{v}"' if isinstance(v, str) else lit(v))
for k, v in r.items()) for r in flat))
show("yaml", "\n".join("- " + "\n ".join(
f"{k}: {lit(v)}" for k, v in r.items()) for r in flat))
show("toon", f"incidents[{N}]{{{','.join(FIELDS)}}}:\n" + "\n".join(
" " + ",".join(lit(v) for v in r.values()) for r in flat))
show("csv", ",".join(FIELDS) + "\n" + "\n".join(
",".join(lit(v) for v in r.values()) for r in flat))
# --- nested payload; CSV can only approximate it ----------------------------
show("json/n", json.dumps(nested, indent=2))
show("toon/n", "query:\n service: auth-api\n window: 24h\n"
"tags[2]: auth,regression\n"
f"incidents[{N}]{{{','.join(nfields)}}}:\n" + "\n".join(
" " + ",".join(lit(v) for v in r.values())
for r in nested["incidents"]))
show("csv/n", "query.service,query.window,tags," + ",".join(nfields) + "\n" + "\n".join(
'auth-api,24h,"auth;regression",' + ",".join(lit(v) for v in r.values())
for r in nested["incidents"]))
The chars-per-token column is the figure to remember. It reveals whether a context budget is being spent on content or on syntax, and it is the only way to know before sending a payload rather than after the bill arrives.
3.2 Measure, Do Not Estimate #
Every provider exposes a counting path, and it should be used in any code that assembles context, since a bad estimate silently becomes a context overflow in production.
Anthropic provides a dedicated endpoint that runs the real tokenizer over a fully-formed request, including tools and system content:
count = client.messages.count_tokens(
model="claude-opus-5",
system=[{"type": "text", "text": SYSTEM_PROMPT}],
tools=TOOLS,
messages=[{"role": "user", "content": user_message}],
)
print(count.input_tokens) # includes tool schema overhead
This is the only reliable way to learn what your tool definitions actually cost, and the number is usually larger than people expect. A twelve-tool surface with detailed descriptions and nested schemas runs 8,000 to 15,000 tokens before a single tool is called. Anthropic labels the result an estimate that can differ from billed usage by a small amount, which is a rounding concern against a rule of thumb that can be wrong by a third.
OpenAI exposes counting through the API as well, at POST /v1/responses/input_tokens, and describes that count as exact rather than as an estimate.2 For local estimation against OpenAI models, tiktoken gives you the tokenizer directly:
import tiktoken
enc = tiktoken.get_encoding("o200k_base")
print(len(enc.encode(text)))
Local encoding reveals the cost of the text itself. It does not, however, reveal the cost of the rendered request, which includes message framing, tool schemas, and provider content. Therefore, count the request when budgeting, and count the text when optimizing a document.
A helper for any production codebase, in C#:
public sealed class ContextBudget
{
private readonly int _windowSize;
private readonly int _reserveForOutput;
public ContextBudget(int windowSize, int reserveForOutput = 8_192)
{
_windowSize = windowSize;
_reserveForOutput = reserveForOutput;
}
public int Available => _windowSize - _reserveForOutput;
public bool Fits(int promptTokens) => promptTokens <= Available;
/// Fraction of usable window consumed. Degradation does not begin at a
/// fixed fraction: §29.5 gives the per-window reliable ranges, which are
/// well under 0.6 on a 1M model. Alarm high, budget low.
public double Pressure(int promptTokens) => (double)promptTokens / Available;
}
That reserve matters, since a request that fills the entire window leaves no room for the response, and providers differ on whether the result is an error or a truncated generation.
3.3 Where Token Boundaries Change Behavior #
Most of the time, tokenization amounts to an accounting detail. In a few specific places, however, it changes the entire capability of the model. Those places matter, since they present as model stupidity when they are in fact artifacts of representation.
Tokenizers group digits inconsistently. 1234 may be one token or two or three depending on the vocabulary and the surrounding characters, and the grouping is not aligned with place value. Arithmetic performed on numbers whose digits are grouped arbitrarily is harder than arithmetic on digits presented individually. This is why models are better at math when they write the calculation out and why comma-separated or space-separated digits sometimes help.
Counting letters in a word, reversing a string, or checking whether a word contains a particular character are all tasks the model performs on a representation where the characters are not individually addressable. Failures here are famous and misinterpreted as reasoning failures. They are representation failures. The fix is to have the model spell the word out first, which converts the task into one it can see.
A prompt ending in a space or a partial token puts the model in an awkward position: the natural continuation may be a token that begins mid-word, and the sampler has to choose among low-probability continuations. Do not end a prompt with trailing whitespace. This is a small effect and an easy one to eliminate.
Delimiters are not free. Wrapping every field in XML tags costs meaningfully more tokens than a compact format. On the other hand, structure helps the model separate instruction from data, which matters for both quality and injection resistance. Section 8.4 works through the trade-off with numbers.
3.4 The Reasoning-Token Line Item #
Reasoning models emit tokens that never appear in the response and that you pay for at the output rate. On any provider that supports extended thinking or a reasoning-effort setting, the invoice line for output includes them.
This produces the defining budgeting surprise of 2026: a task whose visible answer is four hundred tokens can bill for eleven thousand output tokens because the model reasoned for ten thousand six hundred of them. Output is the most expensive lane on every provider, typically five times base input.
Every major provider now exposes a control:
# Anthropic: adaptive thinking, effort setting
response = client.messages.create(
model="claude-opus-5",
max_tokens=4096,
thinking={"type": "adaptive"},
output_config={"effort": "low"},
messages=[{"role": "user", "content": prompt}],
)
// OpenAI: effort setting
const response = await client.responses.create({
model: "gpt-5.6-sol",
reasoning: { effort: "low" },
input: [{ role: "user", content: prompt }],
});
Two warnings that Section 23 develops. First, changing the thinking configuration mid-conversation invalidates the cache—on some models only the message cache, on others the system and tool caches too—and a request-level effort change behaves the same way. Effort has documented exceptions that thinking does not: setting it to the model’s default does not invalidate, models supporting per-message effort can carry the change in a role: "system" message and keep the prefix, and Claude Code takes that path on Opus 5.5 and Fable 5.1 under the conditions §23.1 lists. Check which path your harness uses before budgeting for the miss. Second, the setting is a ceiling on effort, not a target; the model may spend less than you allow, and allowing more does not reliably buy correctness. Accuracy frequently peaks at intermediate spend and saturates above it.6
3.5 Estimating Without a Calculator #
Three numbers let you sanity-check a budget conversation in real time:
- A page of prose is about 500 tokens. A hundred-page document is roughly 50,000.
- A thousand lines of code is about 10,000 tokens. A medium source file is 2,000 to 5,000.
- One million tokens at $3/MTok base input costs $3. Everything else is a multiplier on that: cached reads at a tenth, cache writes at 1.25, output at five.
From those three you can estimate almost any real conversation about spend. A team of twenty developers each running eight agentic sessions a day, each session consuming roughly 400,000 tokens of mixed input at a high cache-hit rate and 30,000 tokens of output, lands in the low thousands of dollars a month. If someone quotes you a number an order of magnitude away from that, one of you is wrong about the workload.
References cited in this section
4 of 81 · numbering matches the PDF
- 4TOON (Token-Oriented Object Notation), specification v3.3 and reference implementation, MIT licensed and https://toonformat.dev, verified September 8, 2026. Project documentation and self-reported benchmarks. Cited for the format definition—YAML-style indentation for nested objects, inline name[N]: for primitive arrays, tabular name[N]{fields}: for uniform object arrays—for the concept of tabular eligibility, and for the two-track benchmark results as published in the May 8, 2026 revision of docs/guide/benchmarks.md, pinned at commit 48191288590767858b132732d88f39b0d85a6ae4, across 209 retrieval questions and four models: on the flat-only track CSV totals 63,997 tokens against TOON's 67,778 (+5.9%); on the mixed-structure track TOON totals 227,830 against formatted JSON's 291,711 (−21.9%) but compact JSON's 198,546 (+14.7%); accuracy 76.4 percent against JSON's 75.0. The benchmark page is versioned with the project, and these figures were replaced on July 24, 2026: from that commit onward the same page reports 244 questions, 72.2 percent TOON accuracy against 71.4 for JSON, and a mixed-structure gap against compact JSON of 1.6 percent rather than 14.7. An earlier revision of this document dated the figures above to a snapshot taken on September 8, 2026, which the repository history rules out—by then the page carried the 244-question run. Cite the commit rather than the page: every figure above is reproducible from the pinned revision and none of them may be mixed with totals from a later one. Also cited for the nested-data results in §7.3—configuration 620 tokens against compact JSON's 558, event logs 154,084 against 128,529, TOON overheads of 11.1 and 19.9 percent—and for the truncation-detection result in which TOON scored 0 of 4 against CSV's 4 of 4. The benchmark evaluates six formats, TOON, JSON, YAML, compact JSON, XML and CSV, and does not evaluate TRON; an earlier revision of this document attributed those nested-data figures to "both optimizers" and to "every one of them," which this source does not support. Token counts use the o200k_base tokenizer. The benchmarks are the project's own and have not been independently replicated, but they are unusually complete: the documentation publishes the cases where the format loses, separates CSV-eligible from CSV-ineligible data rather than averaging across both, and states plainly that the evaluation tests comprehension rather than generation.github.com/toon-format/toon ↗
- 5Ivan Matveev, "Token-Oriented Object Notation vs JSON: A Benchmark of Plain and Constrained Decoding Generation," arXiv:2603.03306, February 2026. Preprint, not peer-reviewed. Cited as the counterweight to reference 4: it observes that TOON's published results test model comprehension rather than generation, and—on four structural cases rather than a sweep of dataset sizes—finds the format's up-front prompt overhead large enough that TOON consumed more tokens than plain JSON on several models. From that the author advances what he labels a scaling hypothesis: that TOON's efficiency advantage likely follows a non-linear curve, materializing only past some point where accumulated syntax savings amortize that overhead. It is a hypothesis and is presented as one; the experiment does not measure the curve or locate the threshold, and the paper's own recommendations call for benchmarking at substantially larger dataset sizes to validate it. Relevant to anyone considering TOON as an output format rather than an input one.arxiv.org/abs/2603.03306 ↗
- 2OpenAI, "Text generation" and "Counting tokens," OpenAI API documentation and https://developers.openai.com/api/docs/guides/token-counting, verified September 25, 2026. Vendor documentation; cited as product fact for the Responses API's instructions and input parameters and its developer role, which the documentation prioritizes ahead of user messages, and for the POST /v1/responses/input_tokens endpoint. OpenAI describes that count as the exact number the model will receive, including the formatting tokens a local tokenizer cannot see, which is a stronger claim than Anthropic makes for its own counting endpoint.developers.openai.com/api/docs/guides/text ↗
- 6Original work on agentic iteration economics establishing that agentic tasks consume roughly a thousand times the tokens of code chat, that input rather than output drives that cost, that runs on the same task differ by up to thirtyfold in total tokens, that accuracy frequently peaks at intermediate cost, and two separate results about anticipating cost that an earlier revision of this document merged into a claim about human forecasting. The paper tests model self-prediction directly: correlations up to 0.39, with systematic underestimation. Its human data are SWE-bench-Verified's expert estimates of how long a professional developer would need to resolve each issue, compared against agent token consumption; it finds that difficulty category is a weak predictor of spend. No human was asked to forecast agent tokens or cost, so the study does not establish that expert humans forecast task cost badly—only that human-effort difficulty transfers weakly as a proxy for it. Cited via reference 1, which contains the full source annotation. The thirtyfold variance figure is the single most consequential number for capacity planning in this document.