Section 5 of 45 6 min read

Why the Same Prompt Gives You a Different Answer

The sampler, why temperature zero is not determinism, and how to design around variance rather than argue with it.

Objective

Explain nondeterminism accurately because “set temperature to zero” is the most commonly given and most misleading advice in this field.

5.1 The Sampler #

A model produces, at each position, a probability distribution over its entire vocabulary, on the order of 100,000 to 200,000 entries. A sampler then selects one token from that distribution, appends it, and the process repeats.

The distribution is deterministic given identical inputs and identical weights. The selection, however, is not, unless it is made so deliberately.

Section 31 covers the sampling parameters in detail. For the moment, the summary is this: temperature scales the distribution before selection, top_p truncates it to a cumulative-probability mass, and top_k truncates it to a count. Setting temperature to zero makes selection greedy—always the highest-probability token.

5.2 Why Zero Is Not Deterministic #

Greedy decoding removes sampling randomness. However, it does not produce reproducible output, and the reasons have nothing whatever to do with your request.

Floating-point non-associativity

Inference runs across many GPUs using batched matrix operations. The order in which partial sums are accumulated depends on batch composition, which depends on what other traffic the provider is serving at that instant. Floating-point addition is not associative, so (a + b) + c and a + (b + c) can differ in the last bits. When two candidate tokens have nearly identical probabilities, a last-bit difference flips the argmax, and the trajectories diverge from that point.

Mixture-of-experts routing

In an MoE architecture, tokens are routed to a subset of expert networks. Routing decisions can depend on batch composition. Different batch, different experts, different logits.

Silent infrastructure changes

Providers update serving stacks, quantization, and kernel implementations without ever changing the model identifier. A model string promises capability, and nothing about bit-exact behavior.

Deployment geography

Each of the three above has a spatial version. A provider serves from more than one fleet, a rollout reaches those fleets in sequence rather than all at once, and local traffic at a local hour shapes the batch your request joins. The same model string can therefore be served by a different build, on different hardware, inside a different batch, depending on where the request lands—and where it lands is not fixed unless you fix it. On the Claude API the default is explicit about this: inference_geo defaults to "global", documented as running inference in any available geography.11

Greedy decoding gives you low variance, not zero variance. Anthropic states this directly for prompt caching, noting that identical requests are not guaranteed to produce identical outputs, and the same holds without caching.12 The corollary is the part to act on: a model string names a model, a dated snapshot names a version of it, and neither one names the fleet that served you.

5.3 Designing Around It Rather Than Against It #

Since variance cannot be eliminated, the productive question becomes how much of it a system can absorb.

Fail loudly. A prompt whose output is validated against a schema turns a variance failure into a caught error. A prompt whose output is scraped with a permissive regular expression invites a silent corruption instead—not inevitably, since a strict pattern anchored end to end can reject malformed output as explicitly as a schema does, but that is a discipline the regex has to be written with and a schema gives you by construction. Section 11 is entirely about this.

Make the target region wide. If several phrasings of an answer are acceptable, variance does not hurt you. If exactly one string is acceptable, you have built a system that will fail intermittently. Where you genuinely need one string, constrain the decode (Section 32) rather than asking politely.

Test distributionally. Running a prompt once and declaring it good is the most common evaluation error in practice. Run it enough times to see the tail. Section 34 covers how many.

The output below is what a test has to accommodate:

from collections import Counter
import anthropic

client = anthropic.Anthropic()

PROMPT = ("Classify the sentiment of this review as positive, negative, or "
          "neutral. Reply with one word.\n\nReview: It arrived on time.")

def run(temperature: float, n: int = 30) -> Counter:
    out = Counter()
    for _ in range(n):
        r = client.messages.create(
            model="claude-haiku-4-5", max_tokens=8,
            temperature=temperature,
            messages=[{"role": "user", "content": PROMPT}],
        )
        out[r.content[0].text.strip().lower()] += 1
    return out

print("temp 0.0:", run(0.0))    # e.g. Counter({'neutral': 28, 'positive': 2})
print("temp 1.0:", run(1.0))    # e.g. Counter({'neutral': 19, 'positive': 11})

Two things usually surprise people on the first run. Temperature zero is not unanimous, for the reasons above. And a split at temperature zero is a signal to chase rather than a reading to trust. Prompt ambiguity produces it; so do the numerical and serving sources above; so does plain model error on a case you specified completely. Treat repeated disagreement as instability requiring diagnosis. It points at the prompt often enough to be the first thing to check, and it does not on its own establish that the prompt is the cause.

Pick a prompt you already trust and run it thirty times before reading further. The result is the argument for Section 34.

Pin what can be pinned, and log the rest. Variance you cannot remove you can still attribute. Pin the dated snapshot so a version change becomes an event rather than a surprise, and where the provider exposes the serving geography, pin that and record what it reports. On the Claude API, inference_geo takes "us" or "global" and usage.inference_geo comes back on every response saying where inference actually ran; pinning to "us" costs 1.1x the standard rate on Claude 4.6 and later, applied to input tokens, output tokens, cache writes, and cache reads alike.11 It is sold as a data-residency control, and it is also the only supported way to hold one of §5.2’s variables still. A log line carrying model, snapshot, and geo is what turns “it has been behaving differently this week” into a question with an answer.

Use the variance. Sampling several candidates and selecting among them with an external check is a well-measured technique. Work combining serial iteration with parallel candidate generation, where each trajectory produces a test script alongside its draft edit, reached 57.4 percent on a benchmark at roughly $4.60 per instance; selecting across edits drawn from top existing submissions reached 66.2 percent, beating the best individual member of that ensemble.13 The candidate pool often contains a correct answer the selector fails to pick, which makes selection, not generation, the under-invested engineering problem.

5.4 When to Use Which Temperature #

This question comes up constantly, and the answer is less interesting than people hope:

TaskSettingReason
Extraction, classification, routingGreedy (0)One correct answer; you want the mode
Code generation0 to 0.3Correctness dominates; some variety helps escape bad starts
Refactoring against testsGreedy, with candidate samplingVerify externally, sample for coverage
Drafting prose0.7 to 1.0Variety is the point
Brainstorming1.0+ with high top_pYou want the tail

Two qualifications belong with that table. Many reasoning models ignore or restrict temperature because the reasoning process is itself trained behavior and providers do not want it perturbed. And changing temperature does not change the ranking of tokens, only the probability of departing from it. A prompt that produces a wrong answer at temperature 1 usually produces the same wrong answer, more consistently, at temperature 0.

5.5 Failure Modes #

  • “We set temperature to zero so it’s deterministic.” It is not, and building a system on that assumption produces failures that are impossible to reproduce.
  • Judging a prompt on one run. You have observed one sample from a distribution you have not characterized.
  • Pinning to an undated model alias. Aliases move. Where reproducibility matters, pin the dated snapshot and treat a version change as a change requiring re-evaluation.
  • Assuming one model string means one serving fleet. It names a model. It does not name a build, a hardware generation, or a geography. Pin the geography where the provider exposes it and log what served the request; where it does not, treat location as an uncontrolled variable in any comparison you draw.

References cited in this section

3 of 81 · numbering matches the PDF

  1. 11Anthropic, "Data residency," Claude Platform documentation verified September 20, 2026. Vendor documentation; cited as product fact for the inference_geo request parameter and its two accepted values—"global", the default, documented as running inference in any available geography, and "us", which confines it to US-based infrastructure—for the usage.inference_geo response field reporting where inference actually ran, for the workspace-level default_inference_geo and allowed_inference_geos settings, and for the 1.1x multiplier applied to US-only inference on Claude 4.6 and later across input tokens, output tokens, cache writes, and cache reads. Scope matters for §5.3's advice: the parameter exists on the Claude API and Claude Platform on AWS only, and on Claude 4.6 and later models only, returning a 400 on earlier ones. Bedrock and Google Cloud determine the inference region from the endpoint URL or inference profile instead, and Microsoft Foundry from the deployment type, which means the geography is pinnable on those platforms but is not reported back in the response. Cited here because it is the one point in the stack where a serving-infrastructure variable is exposed to the caller rather than inferred from behavior. The 1.1x multiplier is a billing fact and should be re-checked alongside the price tables.platform.claude.com/docs/en/manage-claude/data-residency ↗
  2. 12Anthropic, "Prompt Caching," Claude Platform documentation verified September 8, 2026. Vendor documentation; cited as product fact for mechanism, pricing multipliers, minimum cacheable lengths, invalidation behavior, the 20-block lookback window, pre-warming, and data retention. The pricing table in this reference is the primary source for all Anthropic rates quoted in this document.platform.claude.com/docs/en/build-with-claude/prompt-caching ↗
  3. 13Work combining serial iteration with parallel candidate generation, where each trajectory generates a test script alongside its draft edit, reaching 57.4 percent at roughly $4.60 per instance; selection across edits drawn from top existing submissions reached 66.2 percent, outperforming the best individual ensemble member. Establishes that selection, not generation, is the under-invested problem. Cited via reference 1.
PDF↓