Constrained Decoding and Grammar Masking
How grammar masking works, what it guarantees and what it does not, the measured reasoning cost, and the provider surfaces.
Explain how format guarantees are actually implemented, what they cost, and when to reach for them.
32.1 The Mechanism #
Constrained decoding enforces a format by masking the logit vector prior to sampling. At each position, a state machine derived from a grammar determines which tokens could legally continue the output; every other token’s logit is set to negative infinity.
The result is that invalid output is not merely discouraged but rendered unreachable. Reduced to its core, the mechanism is three lines.
class GrammarStuck(RuntimeError):
"""No legal continuation exists. Raised so a failed transition is
catchable and distinguishable from a completed parse."""
def masked_sample(logits, allowed_token_ids, temperature=0.7):
"""Grammar masking, at its core: everything not allowed becomes impossible."""
allowed = list(allowed_token_ids)
if not allowed:
# A dead grammar state is a caller error, not a token. Without this the
# temperature-zero path returns argmax over all -inf, handing back
# token 0 — a token the grammar just forbade.
raise GrammarStuck("no legal continuation in this state")
masked = np.full_like(logits, -np.inf)
masked[allowed] = logits[allowed]
return sample(masked, temperature=temperature)
That guard is not decoration. sample returns argmax at temperature zero without asking whether anything survived the mask, so an empty allowed set silently yields token 0—a forbidden token, returned by the function whose whole promise is that forbidden tokens are unreachable. At nonzero temperature the same input surfaces as a NaN inside the normalization instead, which is no better. Distinguish a completed grammar state, where the right move is to stop, from a failed transition, where the right move is to raise; neither should resolve to whichever token sits at index zero.
The engineering difficulty lies in computing allowed_token_ids efficiently at every position because the grammar operates on characters while the model operates on tokens, and the mapping is many-to-many. A token like ":" may advance the JSON state machine by one character; a token like "name": advances it by seven and only in a specific state. Production implementations precompute token-level transition tables from character-level grammars, which is why they are fast enough to use.
32.2 What It Guarantees, and What It Does Not #
Guaranteed: a completed generation parses. A JSON schema constraint produces valid JSON conforming to the schema with no retry, provided the generation finishes. The vendor documents two exits that do not: a refusal returns stop_reason: "refusal" and output that need not match the schema, and a max_tokens cut returns output that is valid so far and incomplete. Both are stop reasons, which is the second reason to preserve them rather than collapse them into success (§4.2).
A third exception is not a stop reason at all, and it is the one that catches people who have stopped validating. Anthropic documents a capitalization caveat on enum: a normally completed response can return a value whose casing differs from the one the schema allows, with nothing in the response to flag it.76 Any enum field therefore needs a documented normalization or rejection policy on your side, and §32.4’s typed contract needs its serializer configured to match.
Not guaranteed: that the content is correct, that required fields are populated meaningfully, or that enum selections are appropriate. The grammar constrains form and says nothing about truth.
This is the same layered-validation point as §11.4, now with a sharper edge: constrained decoding does most of the work of layers one and two and none of layers three and four. It does not retire the first two, because of the three exceptions above—keep a cheap parse-and-contract check and decide what it does when it fires, rather than deleting it on the strength of the guarantee. The identifiers can still be hallucinated, and the combinations incoherent.
32.3 The Reasoning Cost, Measured #
This is the finding that should govern any use of the technique, and it needs stating carefully, because the intuitive version of it is wrong. Format restrictions measurably degrade reasoning: a peer-reviewed study across multiple models found a significant decline under format restriction, with the study’s JSON-mode condition worst and natural-language-then-convert least harmful. The same study found format restriction improved classification accuracy.34 The tempting inference—that strictness itself is the mechanism, so constrained decoding must be the worst offender—does not survive the same table. Its JSON-Schema condition is the tighter constraint and it outscores format-restricting instructions on every reasoning task tabulated for GPT-4o-mini, and beats even unrestricted natural language on one of them. Poor JSON-mode results are not a result about constrained decoding in general. What follows for practice is that the penalty holds, its size is a property of the specific method, model and task, and the only way to know yours is to measure it.
The mechanism is plausible, even though it has not been measured directly. Reasoning benefits from the model generating intermediate content freely, and a grammar requiring the very next token to be { forecloses that path at position zero: the model commits to a structure before it has worked anything out.
Two scope notes belong with it, because the finding is being asked to carry more than it measured. The study tested particular models and methods from its moment.34 And the mechanism as stated assumes the constrained output is the whole generation—where a model reasons in a separate channel ahead of the constrained text, the room the grammar closes is not the room the reasoning was using, and the argument does not transfer without re-measurement. Measure it on the model you are shipping before treating the extra extraction call in §11.2 as almost always worth it.
The mitigation preserving both properties is to constrain a field within the structure that gives reasoning room.
{
"type": "object",
"properties": {
"reasoning": {
"type": "string",
"description": "Work through the analysis here before answering. Multiple sentences."
},
"answer": { "type": "string", "enum": ["approve", "reject", "escalate"] },
"confidence": { "type": "string", "enum": ["low", "medium", "high"] }
},
"required": ["reasoning", "answer", "confidence"],
"additionalProperties": false
}
Field order is load-bearing here. reasoning must come first because JSON generates left to right and a model that emits answer before reasoning has already committed. Putting the reasoning field last produces post-hoc rationalization, which is worse than no reasoning at all because it looks like justification.
32.4 Provider Surfaces #
# Anthropic — schema-guided structured output
response = client.messages.create(
model="claude-sonnet-5",
max_tokens=2048,
output_config={"format": {"type": "json_schema", "schema": SCHEMA}},
messages=[{"role": "user", "content": prompt}],
)
// OpenAI — structured output on the Responses API
const response = await client.responses.create({
model: "gpt-5.6-terra",
text: { format: { type: "json_schema", name: "decision", schema, strict: true } },
input: [{ role: "user", content: prompt }],
});
// C# — schema as a typed contract, with the reasoning field first
[JsonUnmappedMemberHandling(JsonUnmappedMemberHandling.Disallow)]
public sealed record Decision
{
[JsonPropertyOrder(0)]
[JsonPropertyName("reasoning")] public required string Reasoning { get; init; }
[JsonPropertyOrder(1)]
[JsonPropertyName("answer")] public required DecisionOutcome Answer { get; init; }
[JsonPropertyOrder(2)]
[JsonPropertyName("confidence")] public required Confidence Confidence { get; init; }
}
// Strings only, both directions. The parameterless converter defaults
// `allowIntegerValues` to true, so `"answer": 999` deserializes to an
// undefined member and serializes back out as 999 — a number, against a
// schema that permits three strings.
public sealed class StrictEnumConverter<TEnum> : JsonStringEnumConverter<TEnum>
where TEnum : struct, Enum
{
public StrictEnumConverter() : base(namingPolicy: null, allowIntegerValues: false) { }
}
// The wire names are pinned to the schema's lowercase enum values. The
// parameterless converter emits the member names — "Approve", "Low" — which
// the schema above does not allow, and case-insensitive reading hides it:
// deserialization works while serialization violates the contract.
[JsonConverter(typeof(StrictEnumConverter<DecisionOutcome>))]
public enum DecisionOutcome
{
[JsonStringEnumMemberName("approve")] Approve,
[JsonStringEnumMemberName("reject")] Reject,
[JsonStringEnumMemberName("escalate")] Escalate,
}
[JsonConverter(typeof(StrictEnumConverter<Confidence>))]
public enum Confidence
{
[JsonStringEnumMemberName("low")] Low,
[JsonStringEnumMemberName("medium")] Medium,
[JsonStringEnumMemberName("high")] High,
}
That record is a round-trip contract, so the enum naming has to be stated rather than inherited. It is also where §32.2’s capitalization caveat lands in practice: the schema permits approve, the default converter writes Approve, and a case-insensitive reader on the way back in makes the mismatch invisible until something stricter reads the same payload.
Naming is only half of it, and the other half is the direction the defaults lean. Four of them are permissive in ways a contract is not, and all four were reproduced against this record: integer enum values are accepted and re-emitted as integers unless the converter is constructed with allowIntegerValues: false; required rejects a missing property and admits an explicit null, so a non-null check is still needed on Reasoning (§11.4); unmapped members are skipped unless the type disallows them, which silently drops a field the model invented rather than reporting it; and enum strings outside the set are not reliably rejected, which is the one that looks safest and is not.
That last one repays the detail because allowIntegerValues: false reads like it closes the enum and it does not. The converter’s name-parsing path splits the incoming string on commas and ORs the members it resolves, without requiring a [Flags] enum and without checking that the result is a defined value. So "approve, reject" is accepted and arrives as Reject—a value the schema permits, reached through a string it does not—and "reject, escalate" arrives as the undefined value 3. A follow-up Enum.IsDefined check catches the second and misses the first, which is the more dangerous of the two: one valid decision silently substituted for another. Genuine string membership needs the raw JSON validated against the schema, or a converter that matches the wire names exactly. A typed record narrows the contract; it does not enforce the schema on its own, and where the payload matters, validate against the schema itself.
One caching note catches practitioners regularly: the structured-output schema is rendered into the prompt as instructions, so changing a schema invalidates the cache in the same way changing a system prompt does (§23.1). Version your schemas and change them at boundaries.
32.5 Beyond JSON #
Grammar masking generalizes to any context-free grammar whatsoever. Local inference stacks expose this directly, which is where it becomes powerful: output guaranteed to conform to a grammar you supply—a DSL, a restricted query subset, a wire format.
# GBNF-style grammar constraining output to a SELECT-only SQL subset.
root ::= select-stmt
select-stmt ::= "SELECT " columns " FROM " table where-clause? ";"
columns ::= column (", " column)*
column ::= [a-z_]+
table ::= [a-z_]+
where-clause::= " WHERE " condition (" AND " condition)*
condition ::= column " " op " " value
op ::= "=" | ">" | "<" | ">=" | "<=" | "!="
value ::= "'" [^']* "'" | [0-9]+
A model decoding under that grammar cannot emit DROP TABLE. Not because it was instructed not to, but because the tokens are masked. This is the strongest available form of the §19.1 distinction between asking and enforcing, applied to generation itself.
Be exact about the property it delivers, which is conformance to that grammar rather than validity in a SQL dialect. column and table are both [a-z_]+ here, so SELECT from FROM select; is a well-formed sentence of this grammar and a syntax error in SQLite. If dialect validity is what you need, quote identifiers, exclude reserved words, and run the result through the dialect’s own parser before execution. Authorization is a third and separate concern: masking DROP says nothing about which rows this caller may read.
The limitation is narrower than it was, and its precise shape matters. Hosted APIs do expose grammar constraint on tool input: OpenAI’s custom tools accept a Lark grammar or a regex, subject to a complexity limit the API enforces by rejecting the grammar.77 Hosted constraint over the final assistant text—which is what the SQL example above is doing—is rarer, but it exists: Fireworks accepts a GBNF grammar as response_format={"type": "grammar", ...} on an ordinary serverless chat completion and returns the constrained string in the message content.81 So the surface is uneven across providers rather than absent, and OpenAI in particular exposes it for tool input and not for final text. What self-hosting buys is control over the masking implementation, the runtime, and the absence of a provider-side complexity ceiling—a different operational commitment, and no longer the only way to get a custom grammar.
References cited in this section
4 of 81 · numbering matches the PDF
- 76Anthropic, "Structured outputs," Claude Platform documentation together with the Claude Sonnet 5 migration guide, https://platform.claude.com/docs/en/models/sonnet-5/migration-guide, and "Extended thinking," https://platform.claude.com/docs/en/build-with-claude/thinking, all verified September 25, 2026. Vendor documentation. Cited as product fact for four things the surrounding chapters depend on: that the structured-output API is constrained decoding and is documented in those words (§11.1); the documented capitalization caveat on enum, under which a normally completed response can return a casing the schema does not allow, with no special stop reason to signal it (§32.2); that thinking is on by default on Opus 5 and Sonnet 5, so a thinking block can precede the first text block and response parsing must select by block type rather than by index (§10.4); and that assistant prefill returns HTTP 400 on Sonnet 5, Opus 5 and the Fable and 4.6-through-4.8 families, with structured outputs or a system instruction as the documented replacement (§31.5). Numbered after reference 75 because these pages were added during a later revision; the reference numbering is append-only and no longer strictly follows first citation.platform.claude.com/docs/en/build-with-claude/structured-outputs ↗
- 34Zhi Rui Tam, Cheng-Kuang Wu, Yi-Lin Tsai, Chieh-Yen Lin, Hung-yi Lee, and Yun-Nung Chen, "Let Me Speak Freely? A Study on the Impact of Format Restrictions on Performance of Large Language Models," Proceedings of EMNLP 2024, Industry Track, pages 1218–1236, DOI 10.18653/v1/2024.emnlp-industry.91, arXiv:2408.02442. Peer-reviewed. Establishes that format restrictions degrade reasoning while improving classification accuracy. Cited for that direction and not for a monotonic strictness ordering, which an earlier revision of this document reported as constrained-decoding > format-restricting-instructions > natural-language-then-convert. The published Table 2 contradicts it: on gpt-4o-mini the JSON-Schema condition, which is the stricter constraint, scores 91.71 / 81.77 / 86.07 on GSM8K, Shuffled Objects and Last Letter against the format-restricting instruction's 87.17 / 81.46 / 84.73, and on Last Letter exceeds the natural-language mean of 83.11. The JSON-mode condition is the worst in all three rows, so the study's poor JSON-mode results are what the ordering was built on, and they do not generalize to schema-constrained decoding. The reported standard deviations are wide and these comparisons are not tests of pairwise significance; they are sufficient to refute a universal ordering, not to establish the reverse one. Two further limitations are load-bearing: the study predates current-generation reasoning models, and the authors have published updates in response to methodological critique, which do not establish an intrinsic monotonic penalty either. The direction is well established; the effect size on current models is unconfirmed.arxiv.org/abs/2408.02442 ↗
- 77OpenAI, "Function calling," OpenAI API documentation verified September 25, 2026. Vendor documentation. Cited for the custom-tools section, which accepts a context-free grammar in lark syntax or a regex to constrain a tool's input, and which documents that the API rejects a grammar it considers too complex. This is the hosted counterexample §32.5 needed: grammar constraint over tool input is available without self-hosting, while an arbitrary grammar over the final assistant text is not.developers.openai.com/api/docs/guides/function-calling ↗
- 81Fireworks AI, "Grammar mode," Fireworks documentation verified September 25, 2026. Vendor documentation. Cited as product fact for §32.5's hosted-grammar claim: grammar mode accepts a GBNF grammar—an extension of BNF with regex-like features, following llama.cpp's implementation—passed in the request's response-format field on any Fireworks model, and it constrains the generated assistant text rather than only a tool's input. This is the entry that claim needs; reference 77 covers OpenAI's tool-input grammars only, and an earlier revision of this document attached the Fireworks sentence to it.docs.fireworks.ai/structured-responses/structured-output-grammar-based ↗