Section 31 of 45 6 min read

Sampling, Logprobs, and Decoding Control

From logits to a token, what each sampling parameter does, reading logprobs, calibration, stop sequences, and seeds.

Objective

Cover the decoding layer in enough depth to use it as an instrument rather than as a set of knobs with folk meanings.

31.1 From Logits to a Token #

At each position, the model produces a logit vector over the vocabulary, and the sampler transforms it and selects. The function below is that sampler, written out in full so the order of operations is visible.

import numpy as np

def sample(logits, temperature=1.0, top_p=1.0, top_k=0, rng=None):
    rng = rng or np.random.default_rng()
    logits = np.asarray(logits, dtype=np.float64)

    if temperature == 0:
        return int(np.argmax(logits))                    # greedy

    logits = logits / temperature                        # 1. scale

    if top_k:                                            # 2. truncate by count
        keep = np.argpartition(logits, -top_k)[-top_k:]  # exactly k indices
        mask = np.zeros(logits.shape, dtype=bool); mask[keep] = True
        logits = np.where(mask, logits, -np.inf)

    probs = np.exp(logits - logits.max())
    probs /= probs.sum()

    if top_p < 1.0:                                      # 3. truncate by mass
        order  = np.argsort(-probs)
        cum    = np.cumsum(probs[order])
        keep   = order[:np.searchsorted(cum, top_p) + 1]
        mask   = np.zeros_like(probs); mask[keep] = 1
        probs *= mask; probs /= probs.sum()

    return int(rng.choice(len(probs), p=probs))          # 4. sample

The order matters. Temperature is applied before normalization, so it changes the shape of the distribution. Top-k and top-p truncate after, so they remove candidates without reshaping the survivors’ relative probabilities.

31.2 What Each Parameter Actually Does #

Temperature

divides the logits. Below 1 it sharpens—the gap between the top candidate and the rest widens, making the mode more likely. Above 1 it flattens, making the tail accessible. At 0 it degenerates to argmax.

The important property, and the one that surprises practitioners, is this: temperature does not change the ranking. A prompt whose highest-probability token is wrong produces that wrong token more reliably at temperature 0.2 than at temperature 1.0. Lowering temperature makes a model more consistent, not more correct.

Top-p (nucleus)

keeps the smallest set of tokens whose cumulative probability reaches p. It adapts to the distribution’s shape: where the model is confident, the nucleus is one or two tokens; where it is uncertain, the nucleus is wide. This adaptivity is why top-p is generally preferred over top-k.

Top-k

keeps a fixed count. It does not adapt, which means it is either too permissive on confident positions or too restrictive on uncertain ones. Use it only where you want a hard cap.

The tie policy is the part implementations get wrong, and it is why the sampler above selects k indices rather than thresholding on the kth logit. Thresholding keeps everything tied with the cut value: on logits [3, 2, 2, 1] with top_k=2, it leaves three candidates alive at roughly [0.576, 0.212, 0.212, 0], which is not the fixed cap the parameter promises. Selecting indices breaks the tie arbitrarily; if your application needs a defined tie order, impose one.

Repetition, frequency and presence penalties

discourage tokens that have already appeared, and they are three different mechanisms wearing one name. The count-based pair subtract from the logit and differ in exactly what they multiply: a frequency penalty scales with how many times the token has already occurred, while a presence penalty applies once at full strength as soon as it has occurred at all. vLLM’s implementation makes the distinction plain—the frequency term multiplies an occurrence-count tensor and the presence term a binary occurrence mask—so at coefficient 1 and five prior occurrences the frequency penalty subtracts 5 and the presence penalty subtracts 1.78 One escalates with repetition; the other is a flat tax on having appeared. The classic repetition penalty is sign-aware and multiplicative—Hugging Face’s implementation divides positive logits by the penalty and multiplies negative ones, and the tokens it considers include the prompt by default rather than only what has been generated. Conventions are runtime-specific, so read the one you are using. Either way they are blunt: applied hard enough to stop a loop, they also suppress legitimately repeated tokens—variable names, syntax keywords, the same word in a list. On code generation they cause more problems than they solve. Prefer fixing the prompt or the stop condition.

31.3 Reading Logprobs #

Logprobs convert the model from an opaque generator into an instrument. Where a provider exposes them, they are the most useful diagnostic available.

from openai import OpenAI
client = OpenAI()

resp = client.chat.completions.create(
    model="gpt-5.6-terra",
    messages=[{"role": "user",
               "content": "Classify sentiment as positive/negative/neutral. "
                          "Reply with one word.\n\nText: It arrived on time."}],
    logprobs=True,
    top_logprobs=5,
    max_completion_tokens=4,
)

import math
for tok in resp.choices[0].logprobs.content:
    print(f"{tok.token!r}  p={math.exp(tok.logprob):.4f}")
    for alt in tok.top_logprobs:
        print(f"    {alt.token!r:12} {math.exp(alt.logprob):.4f}")

# 'neutral'  p=0.5120
#     'neutral'     0.5120
#     'positive'    0.4610      ← the model is nearly split
#     'negative'    0.0180

That output reveals something no amount of asking ever would. The model reported “neutral” with the confidence of a coin flip between two categories. A system that logs only the answer sees a confident classification; a system that logs the distribution sees a case that needs review.

31.4 Calibration, Properly #

Verbalized confidence—asking the model “how sure are you?”—and token logprobs are two candidate uncertainty signals, and which is better is a question about your task rather than a settled ranking. They measure different quantities: a logprob is a genuine next-token likelihood, while what you usually want is the probability that the answer is correct, and the provenance of a number does not decide whether it predicts that. The assumption that the self-reported one must be worse because it is “just generated text” does not survive the published comparison—Tian et al. report GPT-4’s verbalized confidence on TriviaQA at an expected calibration error of 0.024 against 0.078 for the model’s own label probability, and title the result accordingly. Evaluate both on your data. Whichever you pick needs calibration against ground truth before you can act on a threshold, and that requirement is the part that does not vary.

import math

def token_confidence(logprobs_content, positions=None) -> float:
    """Geometric mean probability across selected token positions.
    Use positions to restrict to the tokens that carry the answer."""
    toks = logprobs_content if positions is None else [logprobs_content[i] for i in positions]
    if not toks:
        return 0.0
    return math.exp(sum(t.logprob for t in toks) / len(toks))

def route_by_confidence(conf: float, low=0.55, high=0.90) -> str:
    if conf >= high: return "auto_accept"
    if conf >= low:  return "human_review"
    return "escalate_to_stronger_model"

Two design notes apply. Restrict the calculation to the tokens that carry the answer—averaging over boilerplate dilutes the signal toward 1.0 and makes everything look confident. And the thresholds must be fitted on labeled data from your own task; the numbers above are placeholders, and using them unfitted is worse than not routing at all.

The calibration procedure runs as follows: collect a few hundred labeled cases, compute confidence for each, bin by confidence decile, and plot observed accuracy per bin. A well-calibrated model produces a diagonal. Most do not, and the deviation is what your thresholds must correct for.

31.5 Stop Sequences and Structural Control #

Stop sequences terminate generation the moment a given string appears. They are underused and they are a genuine cost lever because output is the expensive lane (§21.4).

response = client.messages.create(
    model="claude-sonnet-5",
    max_tokens=4096,
    stop_sequences=["</analysis>", "\n\n---"],
    messages=[{"role": "user", "content": prompt}],
)
# stop_reason == "stop_sequence" says a stop sequence fired;
# the stop_sequence field says which one.

Two patterns cover most uses. Using a closing tag as a stop sequence lets you generate exactly one structured block and stop, without paying for a trailing explanation—though the API omits the matched sequence from the returned text, so re-append the closing tag if a parser downstream expects it. And prefilling the assistant turn constrains the opening:

messages = [
    {"role": "user", "content": "Extract the fields as JSON."},
    {"role": "assistant", "content": "{"},        # prefill: forces JSON start
]

Prefilling is cheap and effective where it is supported, and support is narrowing—including on the model in the call just above it. Anthropic’s current generation rejects a trailing assistant turn with HTTP 400 on Sonnet 5, Opus 5 and the Fable and 4.6-through-4.8 families, and the documented replacement is a structured output or a system instruction.76 Where prefill does work, it eliminates the “Sure! Here’s the JSON:” preamble entirely rather than asking the model not to produce it; where it does not, reach for §32 instead.

31.6 Seeds #

Some providers accept a seed for reproducibility. It fixes the pseudorandom stream used by the sampler and nothing else. It does not address floating-point non-associativity, batch-dependent kernel selection, MoE routing variability, silent infrastructure changes, or which fleet in which geography served the request (§5.2).

A seed reduces variance without producing reproducibility. Treating it as though it did will build systems whose failures cannot be reproduced.

References cited in this section

2 of 81 · numbering matches the PDF

  1. 78vLLM documentation: "Batch invariance," "Reproducibility," https://docs.vllm.ai/en/latest/usage/reproducibility/, and "Offline inference," https://docs.vllm.ai/en/latest/serving/offline_inference/, all verified September 25, 2026. Project documentation. Cited for two corrections. Batch-invariant serving is an opt-in mode with an enumerated list of tested models, DeepSeek and Qwen MoE families among them, which is why §30.4 describes hosted nondeterminism as a property of default serving rather than an impossibility in principle; the project also tracks configurations where the guarantee does not yet hold, so the mode is scoped rather than universal, and reproducibility additionally requires pinning kernels, sampling, hardware and versions. And vLLM exposes offline in-process inference through LLM.generate() and LLM.chat(), so it is not API-only (§37.4). Also cited, against the penalty implementation in vllm/model_executor/layers/utils.py, for the §31.2 distinction between the two count-based penalties: the frequency term multiplies an occurrence-count tensor and the presence term a binary occurrence mask, so only the former scales with repetition. An earlier revision of this document described both as scaled by how often the token occurred.docs.vllm.ai/en/latest/features/batch_invariance ↗
  2. 76Anthropic, "Structured outputs," Claude Platform documentation together with the Claude Sonnet 5 migration guide, https://platform.claude.com/docs/en/models/sonnet-5/migration-guide, and "Extended thinking," https://platform.claude.com/docs/en/build-with-claude/thinking, all verified September 25, 2026. Vendor documentation. Cited as product fact for four things the surrounding chapters depend on: that the structured-output API is constrained decoding and is documented in those words (§11.1); the documented capitalization caveat on enum, under which a normally completed response can return a casing the schema does not allow, with no special stop reason to signal it (§32.2); that thinking is on by default on Opus 5 and Sonnet 5, so a thinking block can precede the first text block and response parsing must select by block type rather than by index (§10.4); and that assistant prefill returns HTTP 400 on Sonnet 5, Opus 5 and the Fable and 4.6-through-4.8 families, with structured outputs or a system instruction as the documented replacement (§31.5). Numbered after reference 75 because these pages were added during a later revision; the reference numbering is append-only and no longer strictly follows first citation.platform.claude.com/docs/en/build-with-claude/structured-outputs ↗
PDF↓