Section 26 of 45 11 min read

Compression and Token-Reduction Tools

Four different things called compression, the research base, a worked example, what the tools cost you, and how to evaluate a savings claim.

Objective

Cover the tools that shrink what you send, what they actually do, what they cost beyond their price, and how to evaluate a savings claim before you believe it.

26.1 Four Different Things Called Compression #

The category is muddled because four unrelated techniques share a single name. Each compresses something different, saves in a different lane, and fails in a different way.

TechniqueWhat it shrinksLane affectedRecoverable
Output brevityWhat the model writesOutput (5×)N/A
Extractive compressionTokens in the promptInput (1× or 0.1×)No
Structural compressionRedundancy in tool results and filesInputUsually, via a handle
Modality shiftText re-rendered as imagesInput, at image ratesYes, via a handle

Sorting a tool into the right row is the first analytical step because the economics differ by an order of magnitude. Output brevity acts on the most expensive lane, and generation itself is never discounted—caching reduces input only. Input compression acts on the cheapest lane and may be fighting a cached prefix that already costs a tenth of the base rate (§22), essentially reducing its actual benefit if not eliminating it.

That last point deserves emphasis because it is the interaction nobody advertises. Compressing a prefix that is already a cache hit is close to pointless: the cache has taken it from $2.00 per million tokens to $0.20, and halving what remains saves $0.10. The same cut is worth $1.00 on uncached input and $5.00 on output. Worse, if the compressor’s output varies between requests, it rewrites the prefix and destroys the cache, converting cheap reads into full-price misses. A compressor must be deterministic over stable content or it costs more than it saves.

Fig. 9—Compression by LaneWhat a 50% token cut is worth in each lane

26.2 The Research Base #

Prompt compression has a genuine academic literature behind it, and it proves more encouraging than most practitioners assume.

LLMLingua, from Microsoft Research, uses a small language model to identify and drop low-information tokens, reporting up to 20× compression with limited quality loss—on GSM8K, exact-match scores fell by 1.44 and 1.52 points at 14× and 20× respectively.56 LLMLingua-2 reformulates compression as token classification, training a BERT-scale encoder by distillation, and reports better quality than its predecessor at equal ratios with lower latency.56 LongLLMLingua adds question-aware compression that scores segment relevance against the query, reporting performance improvements of up to 21.4 percent at roughly 4× fewer tokens on NaturalQuestions.56

That last result is the interesting one, and the mechanism is simply §12.2 restated: compression that removes distractors raises the signal-to-noise ratio, and the model performs better on the smaller prompt. Compression and curation are the same operation performed by different agents.

The counterweight comes from inside that same literature rather than from outside it. LLMLingua-2’s own reported single-document QA score falls from 35.5 at 3× to 29.8 at 5×—5.7 score points, a 16.1 percent relative decline—over a change in ratio that vendors describe as conservative. Those two figures are easiest to read where a later paper proposing a competing method re-tabulates them as its baseline, but that paper’s own table marks them as taken from the LLMLingua-2 authors rather than re-measured, so they are the method’s published behavior and not independent confirmation of it.56,57 Practitioner writing on the tool landscape agrees qualitatively—useful reduction at moderate ratios, a quality cost at aggressive ones—without converging on numbers anyone can cite.58 My own working band, offered as a heuristic and not as measurement, is to expect a third to two-thirds reduction to be roughly free on prose-heavy context, and to treat anything past four-fifths as a change requiring its own evaluation. [unmeasured]

The working rule is this: compression trades quality for ratio, and on the one extractive method measured above, raising the ratio from 3× to 5× cost 5.7 score points. Below some ratio you are mostly removing noise; above it you are removing signal. Be exact about what two scores can carry because it is less than the shape of that sentence suggests. They establish a drop between two settings and nothing more: two points fit a straight line as well as a curve—35.5 and 29.8 sit exactly on q(x) = 44.05 − 2.85x—so they do not establish curvature, a slope change, a cliff, or the location of a knee. The shape in figure 10 is drawn to illustrate the trade-off and is not traced from data. So treat “quality fell by this much between these two ratios, for this method on this task” as the claim, locate your own usable ceiling by measurement, and do not carry either the magnitude or any assumed shape onto a different task, model or compressor.

Fig. 10—The Compression KneeTask quality against compression ratio

26.3 A Worked Example: Caveman #

One tool repays close examination, since it is widely adopted, spans three of the four categories above, and, unusually, publishes its own negative results.

Caveman began as a Claude Code skill instructing the agent to answer in terse, telegraphic prose, and its early releases carried a fixed 65 percent output ratio into their savings reports. Its own documentation then does something rare: it withdrew the number. A correction committed on September 8, 2026 states that no reviewed aggregate output-reduction result is published, that the 65 percent ratio had never had a committed reviewed result behind it, that current reports ignore those historical estimates while preserving the original history rows, and that the input cost the skill adds is not measured there either.59 The standing caveats survive the withdrawal: the skill shrinks output tokens only, it adds input tokens of its own, and on already-terse workloads the result can be net negative.59

The withdrawal is the instructive event. A figure that had been running in released reports for months turned out to rest on no committed comparison, and the same correction removed the input-side number anyone would need to turn an output reduction into a net one. What the project now tells its users to do instead is to compare provider-billed totals on the same task with the tool on and off, and to switch it off for any workload where the billed total goes up.59

A second methodological point rides along with that one, and it generalizes well beyond one tool. A compression claim measured against an unprompted baseline tells you what the instruction bought, not what the tool did. The correct control is your current prompt plus “be concise,” which costs nothing and takes ten seconds.

The later releases move into structural compression and modality shift. A local proxy sits between the agent and the provider, typing each payload and routing it to a per-content-type compressor, including re-encoding JSON tool results as TOON (§3.1) when that measures smaller. Targets, which its documentation labels as inferred: roughly 70–90 percent on JSON, 85–95 percent on logs, 40–70 percent on code, 60–80 percent on diffs, 80–95 percent on search results, and 50–80 percent on text.59 A pinned 54-run benchmark reports 33.2 percent fewer provider-reported input tokens than direct Claude Code while passing all eighteen exact-answer checks, labeled as a benchmark counterfactual rather than a measurement of real traffic.59

The modality shift is the most striking mechanism, and the one most in need of scrutiny. Text is priced per token; a dense text slab rendered to a PNG is priced as image input. The project reports 55,413 estimated text tokens becoming 11,402 estimated image tokens on a dense payload, a 79 percent reduction, and applies the same trick to skill bodies—one skill going from 1,069 to 415 estimated tokens.59 It also states the limit plainly: the technique only pays on dense, long-line content, sparse code with short lines is not profitable because the page carries more overhead than the text it replaces, and it runs only for models with measured render legibility.59

To make the per-type targets concrete, consider what each compressor is actually doing. A log at 85–95 percent:

before (18 lines, ~240 tokens)
  2026-09-08T14:22:01Z INFO  starting worker pool size=8
  2026-09-08T14:22:01Z INFO  connected to postgres host=db-1
  2026-09-08T14:22:02Z INFO  processed batch id=1 rows=500
  ... 12 more INFO lines ...
  2026-09-08T14:22:09Z ERROR failed to commit batch id=14
    psycopg2.errors.SerializationFailure: could not serialize access
    at worker.py:212 in commit_batch

after (~30 tokens)
  [15 INFO lines elided]
  ERROR failed to commit batch id=14
    psycopg2.errors.SerializationFailure: could not serialize access
    at worker.py:212 in commit_batch

Nothing required to answer “why did this fail” was removed. That is the easy case, and it is why the log target is the most aggressive on the list.

Code at 40–70 percent is a different proposition:

before
  def commit_batch(conn, rows):
      cur = conn.cursor()
      try:
          cur.executemany(INSERT_SQL, rows)
          conn.commit()
      except SerializationFailure:
          conn.rollback()
          raise
      finally:
          cur.close()

after
  def commit_batch(conn, rows):   # body elided, 9 lines
      ...

The signature survives, while the behavior does not. If the agent’s task is “call this function,” the compression is free. If the task is “explain why batches fail intermittently,” the compressor just deleted the answer. This is why the code ratio is the conservative one, and why the recovery handle in §26.4 matters more here than anywhere else.

26.4 What These Tools Cost You #

Price is the least of it. Six further costs, ordered roughly by how often each is overlooked:

1
Cache interaction.

A compressor that rewrites content sitting in your cached prefix invalidates it (§23). Compression of a stable system prompt must produce byte-identical output across requests or you have traded a 0.1× read for a 1.25× write. Verify on the hit fraction, not on the absolute count: compression that works reduces cached reads in proportion to the prefix, so a deterministic 10,000-token prefix compressed to 5,000 should read 5,000 tokens from cache at an unchanged 100 percent hit rate. An unchanged absolute cache_read_input_tokens after compression is evidence the prefix did not shrink. What tells you the tool broke the cache is a write appearing on every request, a hit fraction that falls, or a total cost that rises—so compare warmed configurations on hit fraction, write incidence, prefix stability and total cost.

2
Recovery turns.

Lossy compression with a retrieval handle is safe in the sense that the original bytes survive, but retrieving them costs a full additional turn, and by §4.2’s arithmetic the marginal turn is the expensive one. A compressor that elides a function body the agent then needs has converted an input saving into a round trip.

3
Correctness risk by content type.

Dropping INFO lines from a log is nearly free. Eliding function bodies from source, or dropping middle rows from a result set, is not. The per-type targets in §26.3 should be read as a risk gradient, not a menu: the aggressive ratios are on content where loss is cheap, and the conservative ratio on code is conservative for a reason.

4
A new component in the request path.

A local proxy sees every request, including credentials passing through to the provider. That is a trust decision of the same class as an MCP server (§17.4), and it deserves the same review. Check the telemetry defaults—Caveman, for instance, sends anonymous usage statistics by default with a documented opt-out.59

5
Licensing.

Source-available is not open source. Caveman splits MIT for the skill and adoption surfaces against BSL-1.1 for the engine and proxy, with automatic conversion to Apache-2.0 on the earlier of a fixed date or four years after each version ships, and third-party hosted or embedded service use requiring a commercial license.59 That distinction matters to anyone embedding the tool in a product and is invisible from a README badge.

6
Evaluation debt.

Any lossy transform in the request path is a change requiring your eval suite to run (§34). If you do not have one, compression is the wrong place to start.

26.5 Evaluating a Savings Claim #

Six questions settle most claims, and asked in this order, most resolve after the second.

1
Which lane?

Output, uncached input, or cached input. On Sonnet 5 rates a 60 percent cut in cached input is worth a fiftieth of the same cut in output, and a tenth of the same cut in base input.

2
What was the baseline?

Unprompted, or the same prompt with a brevity instruction? If unprompted, the number is inflated by whatever “be concise” buys for free.

3
Is the overhead counted?

Skills, proxies, and classifiers add tokens or latency. Net or gross?

4
How was quality measured?

Exact-answer checks on a fixed suite are evidence. Absence of complaints is not.

5
Inferred or measured?

Local estimates from a token counter are not a provider invoice. Caveman’s own labeling scheme (inferred, benchmark counterfactual, verified) is a model other tools should copy.

6
Does it survive your traffic?

Compression ratios are a property of your content. The tool only exposes what is already there. JSON-heavy workloads compress well. Terse curated context has little left in it to remove.

26.6 The Order of Operations #

Compression is the last thing to reach for rather than the first, since the free levers are considerably larger.

1. Delete what should not be there.        §14, §15  — free, improves quality
2. Move always-on content on-demand.       §14.1-2   — free, improves quality
3. Cut tool surface area.                  §17.1     — free, improves quality
4. Make the cache work.                    §22       — ~10× on input, no quality cost
5. Route by task class.                    §27       — 5-20× on the routed traffic
6. Add a brevity instruction.              §8.1      — free, immediate
7. Only now: compression tooling.          this §    — real, but bounded and lossy

Steps one through three reduce dilution and therefore improve output quality (§12.1). Step four is a pure win. Compression sits below all of them because it is the only step on the list that can make the model’s input worse.

A team that has not done steps one through six and is evaluating a compression proxy is optimizing the wrong term.

26.7 Failure Modes #

  • Compressing a cached prefix. Trading a 0.1× read for a 1.25× write.
  • Believing an output-token headline is a bill. Output is one lane, and on cache-heavy agentic workloads it is not the largest one by volume.
  • No terse control arm. You measured the instruction, not the tool.
  • Aggressive ratios on agentic code work. Compression research is dominated by question answering and summarization. It is not silent on code—LLMLingua-2 and EFPC both report a LongBench Code column, and LongLLMLingua carries code completion in its task inventory—but code completion over a fixed context is not editing and debugging a repository over a long horizon, and nothing cited here measures the second. Treat aggressive ratios on agentic code work as unmeasured rather than as the QA knee transferred. [unmeasured]
  • Adding a proxy without reviewing it. It sees every request and every credential.
  • Compressing before curating. The cheapest tokens to remove are the ones that should never have been in the window.

References cited in this section

4 of 81 · numbering matches the PDF

  1. 56Huiqiang Jiang, Qianhui Wu, Chin-Yew Lin, Yuqing Yang, and Lili Qiu, "LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models," EMNLP 2023, arXiv:2310.05736; Zhuoshi Pan et al., "LLMLingua-2: Data Distillation for Efficient and Faithful Task-Agnostic Prompt Compression," ACL 2024 Findings, arXiv:2403.12968; and Huiqiang Jiang et al., "LongLLMLingua," ACL 2024. Peer-reviewed. Source of the up-to-20× compression claim, the GSM8K exact-match deltas of 1.44 and 1.52 points at 14× and 20×, the token-classification reformulation in LLMLingua-2, and LongLLMLingua's reported performance improvement of up to 21.4 percent at roughly 4× fewer tokens on NaturalQuestions. The improvement result is the important one: it shows compression can raise quality by removing distractors, consistent with the signal-to-noise finding in reference 35. Limitation to carry, stated precisely: these are LongBench-family and question-answering benchmarks, not agentic coding trajectories, and the compressors were evaluated against models older than the current generation. They are not, however, silent on code—LLMLingua-2's Table 2 and LongLLMLingua's results both carry a LongBench Code column, and LongLLMLingua includes an LCC code-completion example—so an earlier revision of this document overstated the gap by saying nothing cited measured code. What remains unmeasured is long-horizon agentic repository editing and debugging, which is the workload §26.7 warns against transferring these trade-offs to.arxiv.org/abs/2310.05736 ↗
  2. 57Yun-Hao Cao, Yangsong Wang, Shuzheng Hao, Zhenxing Li, Chengjun Zhan, Sichao Liu, and Yi-Qi Hu, "EFPC: Towards Efficient and Flexible Prompt Compression," arXiv:2503.07956v1, 2025. Preprint. Cited for the re-tabulated LLMLingua-2 single-document QA baseline—35.5 at 3× against 29.8 at 5×, which is 5.7 score points and a 16.1 percent relative decline. Its Table 3 marks those rows as taken from Pan et al. (2024), which is reference 56, rather than re-measured here; an earlier revision of this document called them an independent measurement, and they are not one. What the two numbers support is a 5.7-point drop between those two ratios on that task. They do not support a steepening: two points fit a line as well as a curve, and an earlier revision of this entry claimed the curve steepens between 3× and 5× on their strength, which does not follow. The paper also proposes a competing method, so its choice of baseline is not disinterested, and neither the magnitude nor any knee location generalizes to other tasks, models and compressors.arxiv.org/abs/2503.07956 ↗
  3. 58Practitioner surveys of the prompt-compression tool landscape, including PointFive, "Top 10 Prompt Compression Solutions (2026)," and NeuralTrust, "Prompt Compression: Cut Token Costs Without Losing Quality," https://neuraltrust.ai/blog/prompt-compression-guide, both verified September 8, 2026. Vendor-adjacent commercial content; both publishers sell related products, which should be stated. Cited only for the qualitative shape of practitioner opinion—useful reduction at moderate ratios, a quality cost at aggressive ones—and for the observation that extreme published ratios carry large accuracy drops. An earlier revision of this entry attributed a 30-to-70-percent input-reduction band with minimal accuracy loss, and a tradeoff threshold at roughly 80 percent, to these two pages; neither establishes that band, and the PointFive page now carries a different title and is an evaluation guide. The numeric band in §26.2 is the author's own heuristic and is labeled as one. Treat these pages as a summary of practice rather than as measurement.www.pointfive.co/guides/top-prompt-compression-solutions-2026 ↗
  4. 59Julius Brussee, Caveman with docs/HONEST-NUMBERS.md and docs/WRAP-BENCHMARK.md, verified September 8, 2026, and the pipeline diagram docs/assets/pixel-pipeline.svg, pinned at commit bb1cfc8b5ce5 of the same date. Project documentation, self-reported. Source of the withdrawn 65 percent output ratio, the 33.2 percent fewer provider-reported input tokens in a pinned 54-run Claude Code benchmark passing 18 of 18 exact-answer checks, the per-content-type compression targets, the pixel-mode figures (55,413 estimated text tokens to 11,402 estimated image tokens on a dense payload of minified JSON and long-line logs, which the diagram labels −79 percent and inferred; a skill body from 1,069 to 415, which the README carries), the split MIT / BSL-1.1 licensing with automatic Apache-2.0 conversion, and the default-on anonymous telemetry with a documented opt-out. Cited here as a worked example rather than as an endorsement. The project is unusually rigorous about its own limits. The first of those two files, as corrected on September 8, 2026, states that the skill shrinks output tokens only, that no reviewed aggregate output-reduction result is published and the earlier fixed 65 percent ratio had no committed reviewed result behind it, that the input cost the skill adds is not measured there, that whole-session savings can go net negative on terse workloads, and that local results are labeled inferred rather than verified. That self-reporting discipline is the reason it is usable as an example; the numbers themselves remain vendor-reported and unreplicated. Two notes on retrieval, because the dense-payload pair is easy to conclude is missing. It sits in the text of a diagram rather than in prose, so reading the Markdown files above will not find it. And the project is actively developed: the pair is absent from the current revision of every file here, which is why the diagram is pinned to a commit rather than cited at main.github.com/JuliusBrussee/caveman ↗
PDF↓