Choosing a Data Format
Five data shapes, schema hoisting, and what the agentic evidence shows once a format claim is tested inside a loop rather than in isolation.
Turn the format comparison in §3 into a decision procedure, and establish what the independent evidence actually supports about the token-optimized formats now competing for that decision.
Section 3 established that format choice is a token decision and demonstrated five encodings of the same data. This section takes the next step and answers the practical question: given a payload, which format should carry it? The short answer is that the shape of the data decides, and preference has almost nothing to do with it.
7.1 Five Shapes, and Only Five #
Nearly every payload an agent reads falls into one of five shapes. Identifying the shape is the whole of the decision, since each shape has a format that fits it and several that do not.
| Shape | Example | What makes it hard |
|---|---|---|
| Uniform flat table | Log export, query result, metrics series | Nothing. This is the easy case. |
| Uniform nested | Records with a consistent sub-object | The nesting must survive |
| Heterogeneous, repeated schemas | Tool definitions, event streams, trace spans | Several schemas, interleaved |
| Deeply nested, schema-less | Configuration trees, arbitrary API responses | No repetition to exploit |
| Prose or code | Documentation, source files, transcripts | Not structured data at all |
The third row is the one that trips people up. A trace containing HTTP spans, database spans and cache spans is not one table and it is not schema-less either. It is several schemas interleaved, and that shape has its own answer.
7.2 Schema Hoisting: TRON #
TOON (§3.1) attacks repetition by declaring a field list once and streaming rows beneath it. That works beautifully for one uniform array, and it works less well the moment a payload carries several different schemas.
TRON attacks the same repetition from the other direction. Instead of grouping rows under a table header, it hoists each repeated schema into a class definition at the top of the document and then instantiates it inline, which leaves the original record order intact.16
class H: kind,method,path,status,ms
class D: kind,table,op,rows,ms
class C: kind,key,hit,ms
[H("http","GET","/v1/items/0",200,12),
D("db","items","select",40,3),
C("cache","items:0",true,1),
H("http","GET","/v1/items/1",200,13), ...]
The format is a superset of JSON, so any valid JSON document is already valid TRON, and adoption can therefore be incremental.16 The published claim is a 20 to 40 percent token reduction against JSON through schema hoisting alone.16
Measuring that trace payload at eighteen interleaved records—the four shown above continued in the same rotation, elided here for space—gives the following, in characters:
| Format | Characters | vs JSON | Preserves record order |
|---|---|---|---|
| JSON | 1,724 | 1.00× | yes |
| JSON compact | 1,111 | 0.64× | yes |
| YAML | 1,061 | 0.62× | yes |
| TRON | 664 | 0.39× | yes |
| TOON, regrouped | 450 | 0.26× | no |
TOON is smaller, and it achieves that only by regrouping the records into three separate tables by kind. For a trace, where the interleaving is the information, that regrouping destroys the thing you were sending. TRON is larger and keeps the sequence. TOON can keep it too, in the list form it falls back to when an array fails the tabular test, and that form forfeits the saving the row above is measuring.
That is the actual trade-off in the heterogeneous case, and it is not a token comparison at all. It is a question of whether record order carries meaning.
7.3 The Selection Table #
| Data shape | Format that fits | Why |
|---|---|---|
| Massive flat tables | CSV, or TOON where structural guarantees matter | No repeating syntax; the header is paid once |
| Uniform nested records | TOON | Tabular rows for the uniform part, indentation for the rest |
| Heterogeneous, order-significant | TRON | Hoists each schema once; record sequence survives |
| Heterogeneous, order-irrelevant | TOON, regrouped by schema | Regrouping is free when order carries nothing |
| Deeply nested or schema-less | JSON compact, or YAML for readability | No repetition to hoist, so the optimizers lose |
| Prose or code | Raw text | Structure adds tokens and no information |
Two entries deserve their caveats. CSV wins the first row on size alone, and it surrenders types, nesting, and any unambiguous null to get there (§3.1); where those matter, the TOON overhead of roughly six percent buys them back. And the fifth row is the one practitioners resist, since a format adopted for its savings tends to get applied everywhere. On deeply nested data TOON loses to compact JSON outright, by eleven percent on configuration trees—620 tokens against 558—and twenty percent on semi-uniform event logs, 154,084 against 128,529.4 That benchmark evaluates TOON, JSON, YAML, compact JSON, XML and CSV; it does not evaluate TRON, so read the row as a measured result for TOON and a reasonable expectation for TRON, which hoists schemas and has correspondingly less to hoist here. Measure it before relying on it.
7.4 What the Agentic Evidence Actually Shows #
Everything above concerns size. Size is the easy half, and it is the half every vendor benchmark measures. The harder question is whether these formats hold up inside an agent loop, where the model must not only read the payload but also generate structured calls back.
Independent work published in 2026 evaluated both TOON and TRON across four agentic benchmarks and five configurations of four open-weight models—Qwen3-32B appears twice, with thinking enabled and disabled—deliberately separating input compression from output compression so that comprehension and generation could be measured apart.17 Four findings from it should govern any adoption decision.
Averaged per model across the four benchmarks, TOON reduced total tokens by 2 to 18 percent and TRON by 0 to 27 percent.17 Those figures sit well below the 40 percent that circulates from single-payload comparisons.
The authors state plainly that the dominant pattern is per-benchmark rather than per-format.17 A format that helps on one workload does nothing on another, which means published percentages transfer poorly and your own measurement is not optional.
On one function-calling benchmark, reasoning configurations showed TRON accuracy drops of 27 to 44 percentage points—in the study’s input-only condition, with the model still emitting its tool calls as JSON.17 That is a catastrophic failure rather than a tuning problem, and it appeared specifically in the models one would expect to handle an unfamiliar notation best. The related finding is that tool-call training transfers to comprehension but not to robust generation in unfamiliar formats.17
On the two multi-turn benchmarks, TOON exhibited what the authors term a parsing-cascade effect, in which a single malformed structure propagates through subsequent turns.17 That is §27.4’s routing argument arriving from a different direction: inside a loop, a format error does not merely produce a worse answer but corrupts everything downstream of it.
The paper’s own conclusion is measured, and its shape carries more than its words: it identifies TRON as a defensible drop-in for JSON in token-sensitive agentic systems, which is an endorsement bounded by two qualifiers.17
7.5 Input Format and Output Format Are Different Decisions #
This is the central distinction here, and the one most often collapsed.
Asking a model to read a compact format is a comprehension task. The formats were designed for it, the benchmarks measure it, and the models are broadly capable of it.
Asking a model to write a compact format is a generation task in a notation that appeared rarely or never in training. Separate work found the picture materially less favorable for generation than for comprehension.5
The safe default follows from the lane arithmetic: compress what the model reads, and let it write JSON. Input compression captures most of the available token reduction, since input outnumbers output twenty to twenty-five times in agentic work (§4.2). Whether it captures most of the spend is a different question, and the answer moves with your cache-hit rate: at Sonnet 5’s $0.20 cached read against $10.00 output, twenty-five cached input tokens come to half the charge of the one output token beside them. §21.4 is that inversion in full. Output compression captures the smaller share while carrying the generation risk above.
One qualifier belongs with that default, and it is the reason step five of the procedure below is measurement rather than trust. Reading a compact format is the safer half and it is not a free one: the 27 to 44 point collapses in §7.4 were measured with compact input and JSON output, which is exactly the configuration this rule recommends.
7.6 A Decision Procedure #
(§7.1). If it is prose or code, stop and send raw text.
If it does and the schemas vary, TRON is the compact option that preserves it; JSON and YAML preserve it at more tokens, and TOON preserves it only in list form rather than the regrouped tables measured in §7.2.
If it nests deeply and irregularly, use compact JSON and stop; the optimizers lose here.
If they do, CSV is out.
, since the benchmark evidence says workload dominates format.17
Let the model emit JSON.
(§34), since a format change is a prompt change.
Steps five and seven are the ones that often get skipped. Unfortunately, they are the two that separate a saving from a regression.
7.7 A Note on Format Churn #
TOON, TRON, ZON, and several others emerged within roughly a year of each other, all making broadly similar claims against JSON. That churn is itself information. The category is young, the specifications are moving, tooling maturity varies by an order of magnitude between them, and independent evaluation is thin and recent.
None of that argues against adoption. It argues for a particular posture: keep JSON as the representation in your code, treat any of these as an encoding applied at the prompt boundary, and make sure the encoder is a function you can remove easily. §3.1 called TOON a translation layer, and the same framing protects you here. A translation layer can be swapped when a better one arrives, or deleted when the claimed savings fail to appear on your traffic.
7.8 Failure Modes #
- Choosing on published percentages. The dominant pattern is per-benchmark, not per-format.17
- Asking the model to generate the compact format. Comprehension and generation are different tasks, and generation in an unfamiliar notation measures worse.5
- Applying one format everywhere. TOON loses to compact JSON on deeply nested data, measurably; TRON is untested there (§7.3).4
- Regrouping records to fit a tabular format. Where order carries meaning, that is data loss disguised as compression.
- Adopting without re-running evals. A format change is a prompt change (§34.8).
- Treating CSV as a general option. It cannot carry nesting, types, or an unambiguous null, and it fails silently on all three.
References cited in this section
4 of 81 · numbering matches the PDF
- 16TRON (Token Reduced Object Notation) specification and reference implementations with SDKs at https://www.npmjs.com/package/@tron-format/tron and https://pypi.org/project/tron-python/, verified September 8, 2026. Project documentation and self-reported figures. Cited for the format definition—an optional header of class declarations followed by a data section using ClassName(v1,v2,…) instantiation syntax—for its status as a superset of JSON in which any valid JSON document is already valid TRON, and for the claimed 20 to 40 percent token reduction achieved by hoisting repeated object schemas into the header. That magnitude comes from the PyPI package description, pinned at tron-python 0.1.0 since a registry version is immutable; the figure appears there and not in the specification, which commits to no figure and states that the more aggressive encoding strategy "does not always guarantee fewer tokens than pure JSON encoding." The reduction claim is the project's own and has not been independently replicated at that magnitude; the agentic evaluation in reference 17 measures 0 to 27 percent in end-to-end loops. One documentation claim warrants skepticism: the body section is described as readable by existing JSON parsers without modification, yet A("x",1) is not valid JSON, so a TRON-aware parser is required in practice for any document that actually uses the class syntax.tron-format.github.io ↗
- 4TOON (Token-Oriented Object Notation), specification v3.3 and reference implementation, MIT licensed and https://toonformat.dev, verified September 8, 2026. Project documentation and self-reported benchmarks. Cited for the format definition—YAML-style indentation for nested objects, inline name[N]: for primitive arrays, tabular name[N]{fields}: for uniform object arrays—for the concept of tabular eligibility, and for the two-track benchmark results as published in the May 8, 2026 revision of docs/guide/benchmarks.md, pinned at commit 48191288590767858b132732d88f39b0d85a6ae4, across 209 retrieval questions and four models: on the flat-only track CSV totals 63,997 tokens against TOON's 67,778 (+5.9%); on the mixed-structure track TOON totals 227,830 against formatted JSON's 291,711 (−21.9%) but compact JSON's 198,546 (+14.7%); accuracy 76.4 percent against JSON's 75.0. The benchmark page is versioned with the project, and these figures were replaced on July 24, 2026: from that commit onward the same page reports 244 questions, 72.2 percent TOON accuracy against 71.4 for JSON, and a mixed-structure gap against compact JSON of 1.6 percent rather than 14.7. An earlier revision of this document dated the figures above to a snapshot taken on September 8, 2026, which the repository history rules out—by then the page carried the 244-question run. Cite the commit rather than the page: every figure above is reproducible from the pinned revision and none of them may be mixed with totals from a later one. Also cited for the nested-data results in §7.3—configuration 620 tokens against compact JSON's 558, event logs 154,084 against 128,529, TOON overheads of 11.1 and 19.9 percent—and for the truncation-detection result in which TOON scored 0 of 4 against CSV's 4 of 4. The benchmark evaluates six formats, TOON, JSON, YAML, compact JSON, XML and CSV, and does not evaluate TRON; an earlier revision of this document attributed those nested-data figures to "both optimizers" and to "every one of them," which this source does not support. Token counts use the o200k_base tokenizer. The benchmarks are the project's own and have not been independently replicated, but they are unusually complete: the documentation publishes the cases where the format loses, separates CSV-eligible from CSV-ineligible data rather than averaging across both, and states plainly that the evaluation tests comprehension rather than generation.github.com/toon-format/toon ↗
- 17"Notation Matters: A Benchmark Study of Token-Optimized Formats in Agentic AI Systems," arXiv:2605.29676, June 2026. Preprint, not peer-reviewed, and the only independent end-to-end evaluation of these formats located. Evaluates TOON and TRON across four agentic benchmarks (BFCL, MCPToolBenchPP, MCP-Universe, StableToolBench) and five configurations of four open-weight models—Mistral-Small-24B, Qwen3-32B with thinking on and again with it off, DeepSeek-R1-Distill-Qwen-32B, and Llama-4-Scout-17B-16E—which the paper itself describes as "five configurations span four model families." An earlier revision of this document called them five models; the thinking toggle on Qwen3-32B is one of the study's own comparisons, so the distinction is load-bearing rather than pedantic. Deliberately decouples input compression from output compression so that comprehension and generation are measured separately. Source of the per-model token reductions (TOON 2 to 18 percent, TRON 0 to 27 percent), the finding that the dominant pattern is per-benchmark rather than per-format, the TRON accuracy drops of 27 to 44 percentage points on BFCL under reasoning configurations, which fall in the study's input-only condition with tool calls still emitted as JSON, the parsing-cascade effect observed for TOON on the two multi-turn benchmarks, the observation that tool-call training transfers to comprehension but not to robust generation in unfamiliar formats, and the conclusion identifying TRON as a defensible drop-in for JSON in token-sensitive agentic systems. The most important single reference in §7, and the only one measuring these formats where practitioners actually deploy them.arxiv.org/abs/2605.29676 ↗
- 5Ivan Matveev, "Token-Oriented Object Notation vs JSON: A Benchmark of Plain and Constrained Decoding Generation," arXiv:2603.03306, February 2026. Preprint, not peer-reviewed. Cited as the counterweight to reference 4: it observes that TOON's published results test model comprehension rather than generation, and—on four structural cases rather than a sweep of dataset sizes—finds the format's up-front prompt overhead large enough that TOON consumed more tokens than plain JSON on several models. From that the author advances what he labels a scaling hypothesis: that TOON's efficiency advantage likely follows a non-linear curve, materializing only past some point where accumulated syntax savings amortize that overhead. It is a hypothesis and is presented as one; the experiment does not measure the curve or locate the threshold, and the paper's own recommendations call for benchmarking at substantially larger dataset sizes to validate it. Relevant to anyone considering TOON as an output format rather than an input one.arxiv.org/abs/2603.03306 ↗