Section 9 of 45 5 min read

Examples, and When They Stop Helping

What examples actually do, the decision rule for using them, how many, and where they start costing more than they return.

Objective

Give the decision rule for when examples earn their tokens, how to construct them, and the specific ways they backfire.

9.1 What Examples Actually Do #

An example does not teach the model a capability at all. Instead, it selects among capabilities the model already possesses by demonstrating the mapping from input shape to output shape. That framing explains both when examples work spectacularly and when they accomplish nothing.

They work where the task is easier to demonstrate than to describe. Formatting conventions, edge-case handling, tone, classification boundaries in an ambiguous space—all easier to show. They do nothing when the task is well-specified in words and the model already knows how, in which case the tokens are wasted and the examples can narrow the model’s approach unhelpfully.

9.2 The Decision Rule #

The question is whether a two-sentence description would make the output unambiguous.

If yes, describe it. If you find yourself writing “well, except when…” three times, show it.

Describable, so describe:
  "Return the result as a JSON object with keys `severity` (one of
   low/medium/high) and `summary` (under 20 words)."

Not describable, so demonstrate:
  Classifying support tickets into an internal taxonomy where the
  boundary between "billing" and "account" depends on institutional
  convention that no rule captures.

Describing costs a sentence; demonstrating costs a block of tokens on every request that carries it. So the default is description, and examples have to justify the cost by doing something a description cannot, which in practice means encoding a boundary that exists only by convention.

9.3 Constructing Examples #

Five rules govern this, ordered by how often violating them causes a problem.

1
Cover the edges, not the center.

The model handles the typical case. Spend your examples on the boundary: the empty input, the ambiguous case, the one where the right answer is to decline. An example set of five typical cases teaches less than two typical and three hard.

2
Keep the format identical.

Every example must use the exact structure you want back. The model is learning the shape as much as the content, and an inconsistency in the demonstration produces an inconsistency in the output.

3
Balance the labels.

In classification, an unbalanced example set biases the output distribution toward the majority class. If you show four positives and one negative, expect more positives than the data warrants.

4
Include a negative case.

One example of correct refusal or correct empty output is worth several positive examples because it establishes that not-answering is in the space.

5
Order matters and recency wins.

The last example has more influence than the first. Put the most representative case last.

A well-formed few-shot block:

<examples>
  <example>
    <input>The charge on my card doesn't match my invoice.</input>
    <output>{"category": "billing", "confidence": "high"}</output>
  </example>

  <example>
    <input>I can't log in and my email changed.</input>
    <output>{"category": "account", "confidence": "high"}</output>
  </example>

  <example>
    <!-- Edge case: sits on the billing/account boundary -->
    <input>My subscription renewed but my seat count is wrong.</input>
    <output>{"category": "billing", "confidence": "medium"}</output>
  </example>

  <example>
    <!-- Negative case: the model must be able to decline -->
    <input>hi</input>
    <output>{"category": "unclassifiable", "confidence": "high"}</output>
  </example>
</examples>

One property of that block stands out against the five rules: half of the four examples are not ordinary cases. That ratio is the whole technique. The model already handles the typical input; what it cannot infer is where your boundaries sit and what you want when the input does not belong in the taxonomy at all. The refusal sits last against rule five for the same reason: recency is spent on the behavior most likely to go missing, and a classifier that cannot decline is the failure this block exists to prevent.

9.4 How Many #

The curve is steep and then flat. Moving from zero examples to one or two produces the large gain, while moving from two to five produces a modest one. Beyond five, the additions are usually noise and, on reasoning tasks, they can be actively negative.

The exception arises when your context is cached. Anthropic’s own guidance notes that with prompt caching you can include twenty or more diverse high-quality examples, since the cost of carrying them collapses to a tenth after the first request.12 That significantly changes the calculus: if the example block is in a stable cached prefix and your session makes many calls, more examples are close to free. If you are making one call, they are not.

The economics, examined on Sonnet 5 at $2.00/MTok base input, $2.50 cache write, $0.20 cache read (verified September 8, 2026):26

Scenario4,000-token example block, 50 requests
Uncached50 × 4,000 × $2.00/M = $0.40
Cached (1 write, 49 reads)(4,000 × $2.50/M) + (49 × 4,000 × $0.20/M) = $0.05

That is an eight-fold reduction, and it turns “we cannot afford twenty examples” into “we can afford twenty examples when rightly leveraged.” Section 22 covers how to make sure they actually land in the cached prefix, which is where most people get this wrong.

9.5 Where Examples Backfire #

They anchor too hard

If all your examples use the same structure and a real input does not fit it, the model will force the input into that structure rather than handling it correctly. Diversity in the example set is a defense.

They leak

Specific values from examples appear in outputs. If an example contains user_id: 12345, expect 12345 to appear in responses about entirely different users. Use obviously synthetic placeholders.

They contradict the instruction

When the prose says one thing and the examples demonstrate another, the examples win. The consequence is that an example set that drifted from an updated instruction silently overrides the update.

They are stale

An example set built against a schema that has since changed teaches the old schema with full authority.

9.6 The Verification Habit #

The cheapest useful test on any few-shot prompt is to remove the examples and run it again. If the output is materially the same, delete the examples and reclaim the tokens. If it degrades, you have confirmed they are load-bearing and you now know by how much.

Almost nobody performs this test, and consequently a substantial fraction of production example blocks amount to pure cost.

References cited in this section

2 of 81 · numbering matches the PDF

  1. 12Anthropic, "Prompt Caching," Claude Platform documentation verified September 8, 2026. Vendor documentation; cited as product fact for mechanism, pricing multipliers, minimum cacheable lengths, invalidation behavior, the 20-block lookback window, pre-warming, and data retention. The pricing table in this reference is the primary source for all Anthropic rates quoted in this document.platform.claude.com/docs/en/build-with-claude/prompt-caching ↗
  2. 26Anthropic model pricing, published in the pricing table of reference 12 and verified September 8, 2026. Source for all Claude per-model rates: Fable 5.1 $10/$50, Opus 5 $5/$25, Sonnet 5 $2/$10, Sonnet 4.6 $3/$15, Haiku 4.5 $1/$5 per MTok, with cache multipliers of 1.25× (5m write), 2× (1h write), and 0.1× read (0.025× on Fable 5.1 and Mythos 5.1).platform.claude.com/docs/en/about-claude/pricing ↗
PDF↓