Output Contracts
Three levels of enforcement, the measured cost of a schema on reasoning, schema design, validation in depth, and retrying well.
Cover getting reliably parseable output, the real cost of imposing a schema, and the validation architecture that makes the whole thing safe.
11.1 Three Levels of Enforcement #
There are exactly three ways to obtain structured output, and they differ in guarantee, cost, and effect on quality.
| Level | Mechanism | Guarantee | Quality effect |
|---|---|---|---|
| Asking | ”Return JSON with keys x, y” in the prompt | None | Mildly negative on reasoning34 |
| JSON mode | Output constrained to valid JSON, fields unchecked | It parses | Negative on reasoning34 |
| Schema-constrained | Grammar-masked sampling, which is how the provider structured-output APIs are built | Conforms to the schema on a completed generation, with documented exceptions (§32.2) | Strongest effect, positive or negative |
The third row merges two things routinely sold as separate products: a provider’s structured-output API is constrained decoding, and Anthropic documents it in those words.76 Choosing the API rather than a grammar is a choice of interface, not of enforcement strength.
Most practitioners reach for that third row by default. That is frequently wrong, and the reason why is the purpose of the following subsection.
11.2 The Cost of a Schema #
Format restrictions measurably degrade reasoning. A peer-reviewed study across multiple models and task families found a significant decline in reasoning ability under format restriction, with natural-language-then-convert the least harmful condition and its JSON-mode condition the most. The same work found format restrictions improved classification accuracy.34 What that study does not establish is that severity tracks strictness monotonically: its schema-constrained condition scores above its format-restricting-instruction condition on all three of the reasoning tasks it tabulates for GPT-4o-mini, despite being the tighter constraint. Treat the effect as dependent on the method, model and task, and measure the configuration you actually deploy rather than ordering conditions by how strict they sound.34
That asymmetry is the actionable finding, and it yields a clean rule:
- Classification, extraction, routing—impose the schema. It helps.
- Reasoning, analysis, debugging, design—do not impose a schema on the reasoning. Let the model work in prose, then extract.
The two-stage pattern that follows:
# Stage 1: reason freely. No schema, no format pressure.
analysis = text_of(client.messages.create( # §10.4: by type, not by index
model="claude-opus-5",
max_tokens=4096,
messages=[{"role": "user", "content": f"Analyze this incident:\n\n{report}"}],
))
# Stage 2: extract structure from the finished reasoning. Cheap model,
# schema imposed, no reasoning required — this is now a parsing task.
structured = client.messages.create(
model="claude-haiku-4-5",
max_tokens=1024,
output_config={"format": {"type": "json_schema", "schema": INCIDENT_SCHEMA}},
messages=[{"role": "user",
"content": f"Extract the fields defined by the schema from "
f"this analysis. Do not infer anything not stated.\n\n"
f"{analysis}"}],
)
This costs one extra call on a cheap model. On Haiku 4.5 at $1/$5 per MTok,26 the extraction call on a 2,000-token analysis costs roughly a fifth of a cent, so price is rarely the deciding factor. What it buys back is the reasoning quality the schema would otherwise have taken on the models and methods where that effect was measured—and §32.3 sets out why the scope matters, and why a model that reasons in a separate channel ahead of the constrained text may not pay the same cost. Treat the two-call split as the option to beat rather than the settled answer, and run both against your eval suite (§34) before adopting either as a default.
The instruction “do not infer anything not stated” in stage two is not decoration. Without it, the extractor will fill required fields with plausible values when the analysis did not supply them, which converts a missing-data problem into a fabrication problem.
11.3 Schema Design #
The schema is itself a prompt. Models read field names and descriptions, and a well-named schema needs less instruction around it.
INCIDENT_SCHEMA = {
"type": "object",
"properties": {
"severity": {
"type": "string",
"enum": ["sev1", "sev2", "sev3", "sev4"],
"description": "sev1 = full outage; sev2 = major degradation; "
"sev3 = partial; sev4 = cosmetic.",
},
"root_cause": {
"type": ["string", "null"],
"description": "The cause if the analysis identifies one. "
"Null if the analysis is inconclusive. Do not guess.",
},
"affected_services": {
"type": "array",
"items": {"type": "string"},
"description": "Service names exactly as they appear in the analysis.",
},
"customer_impact_confirmed": {"type": "boolean"},
},
"required": ["severity", "root_cause", "affected_services",
"customer_impact_confirmed"],
"additionalProperties": False,
}
Four design decisions appear there, and each prevents a specific failure:
- Enums over free strings wherever the value space is closed. This eliminates an entire class of downstream normalization.
- Nullable, with an explicit meaning.
"type": ["string", "null"]plus “Null if inconclusive. Do not guess” is how you get honest missing data. A required non-nullable string with no route for “unknown” invites fabrication: the model can still write a truthful marker into it, and generation can still fail outright, but you have made the plausible guess the path of least resistance and left yourself no way to tell one from the other. - Descriptions carrying the decision rule. The severity description is the classification rubric, positioned where the model reads it while filling the field.
additionalProperties: False. Without it, models add fields, and your consumer either ignores them or breaks.
Field ordering matters as well. Models generate JSON left to right, so any field depending on reasoning should follow the fields that establish it. Put severity after root_cause if severity follows from cause.
11.4 Validating in Depth #
Schema validity is not correctness, since a response can satisfy every constraint and still be wrong. Structure your validation in layers.
using System.Text.Json;
using System.Text.Json.Serialization;
// `Disallow` is the analog of the schema's `additionalProperties: false`.
// Without it the deserializer ignores unmapped fields, and a model that invents
// one passes a validator whose whole job is refusing what the schema forbids.
[JsonUnmappedMemberHandling(JsonUnmappedMemberHandling.Disallow)]
public sealed record IncidentReport
{
[JsonPropertyName("severity")] public required string Severity { get; init; }
// `required string?` — the schema lists root_cause in `required` AND allows
// null. Those are two different constraints: the property must be PRESENT,
// and its value may be null. Dropping `required` accepts a model that
// simply omitted the field, which is exactly the silence §34.3 grades for.
[JsonPropertyName("root_cause")] public required string? RootCause { get; init; }
[JsonPropertyName("affected_services")] public required string[] AffectedServices { get; init; }
[JsonPropertyName("customer_impact_confirmed")] public required bool CustomerImpactConfirmed { get; init; }
}
public static class IncidentValidator
{
private static readonly string[] Severities = ["sev1", "sev2", "sev3", "sev4"];
public static Result<IncidentReport> Validate(string raw, IReadOnlySet<string> knownServices)
{
// Layer 1 — syntactic. Does it parse at all?
IncidentReport? report;
try { report = JsonSerializer.Deserialize<IncidentReport>(raw); }
catch (JsonException ex) { return Result.Fail<IncidentReport>($"malformed json: {ex.Message}"); }
if (report is null) return Result.Fail<IncidentReport>("null document");
// `required` enforces presence, not non-null. An explicit JSON null
// satisfies it and arrives as a null reference, which would throw out
// of layer three rather than returning a failure.
if (report.AffectedServices is null || report.AffectedServices.Any(s => s is null))
return Result.Fail<IncidentReport>("affected_services must be a list of strings");
// Layer 2 — schematic. Are the values in range?
if (!Severities.Contains(report.Severity))
return Result.Fail<IncidentReport>($"unknown severity '{report.Severity}'");
// Layer 3 — semantic. Do the values refer to things that exist?
// This is the layer that catches hallucinated service names, and
// it is the layer almost nobody builds.
var unknown = report.AffectedServices.Where(s => !knownServices.Contains(s)).ToArray();
if (unknown.Length > 0)
return Result.Fail<IncidentReport>($"unknown services: {string.Join(", ", unknown)}");
// Layer 4 — business. Is this combination coherent?
if (report.Severity == "sev1" && !report.CustomerImpactConfirmed)
return Result.Fail<IncidentReport>("sev1 requires confirmed customer impact");
return Result.Ok(report);
}
}
Layer three is the one that matters most, and it is the one usually missing. A model asked for service names will produce service names; whether those services exist is a question only your system can answer. Every identifier a model emits—table names, endpoints, package names, file paths, ticket IDs—should be checked against a real registry before it is used. Hallucinated package names in particular are an established supply-chain attack surface.
Layer one carries three traps, and the deserializer’s defaults supply all of them. The required modifier makes the deserializer reject a missing property and says nothing about an explicit null, and nullable annotations are not enforced by default. Without the affected_services guard, "affected_services": null passes both early layers and throws out of a function whose whole contract is returning failures as values.
The other two are omissions rather than guards, and they make a validator quietly weaker than the schema it claims to enforce. A property left un-required because it is nullable accepts a response that never mentioned it, so root_cause needs required string?—“must be present, may be null”—rather than string?, which is “may be absent”. And unmapped members are skipped by default, so additionalProperties: false buys nothing on the consuming side until the type disallows them. Reproduced under .NET 8: without those two, a document missing root_cause and a document carrying an invented field both reach layer four and return Ok.
11.5 Retrying Well #
When validation fails, the retry must carry the specific failure, and it must not invite a rewrite.
def request_validated(client, prompt, schema, validate, max_attempts=3):
messages = [{"role": "user", "content": prompt}]
for attempt in range(max_attempts):
response = client.messages.create(
model="claude-sonnet-5",
max_tokens=2048,
output_config={"format": {"type": "json_schema", "schema": schema}},
messages=messages,
)
raw = text_of(response) # §10.4: by type, not by index
ok, err, value = validate(raw)
if ok:
return value, attempt + 1
messages += [
{"role": "assistant", "content": response.content},
{"role": "user", "content":
f"Validation failed: {err}\n\n"
f"Return the corrected object. Change only what the error "
f"names; leave every other field exactly as you had it."},
]
raise ValidationExhausted(err)
The last clause is doing real work. Unscoped retries produce rewrites, and rewrites break fields that were already correct. This is the same principle as the oracle loop in §10.4 and it generalizes into a rule: when correcting a model, name the defect and fence the change.
Instrument retry-exhaustion separately from other failures. A rising exhaustion rate is an early signal of prompt drift, schema drift, or a model change, and it is invisible if you fold it into a generic error counter.
11.6 Failure Modes #
- Schema on a reasoning task, unmeasured. Degraded reasoning in the tested models and methods; re-measure before assuming it on yours (§32.3).34
- Required non-nullable fields with no “unknown” path. Invites fabrication rather than preventing it, and leaves you unable to distinguish a guess from a finding.
- Regex parsing of prose output. Works until it does not, and fails silently.
- Validating structure and stopping. Well-formed and wrong is the dangerous case.
- Trusting model-emitted identifiers. Check them against something real.
References cited in this section
3 of 81 · numbering matches the PDF
- 34Zhi Rui Tam, Cheng-Kuang Wu, Yi-Lin Tsai, Chieh-Yen Lin, Hung-yi Lee, and Yun-Nung Chen, "Let Me Speak Freely? A Study on the Impact of Format Restrictions on Performance of Large Language Models," Proceedings of EMNLP 2024, Industry Track, pages 1218–1236, DOI 10.18653/v1/2024.emnlp-industry.91, arXiv:2408.02442. Peer-reviewed. Establishes that format restrictions degrade reasoning while improving classification accuracy. Cited for that direction and not for a monotonic strictness ordering, which an earlier revision of this document reported as constrained-decoding > format-restricting-instructions > natural-language-then-convert. The published Table 2 contradicts it: on gpt-4o-mini the JSON-Schema condition, which is the stricter constraint, scores 91.71 / 81.77 / 86.07 on GSM8K, Shuffled Objects and Last Letter against the format-restricting instruction's 87.17 / 81.46 / 84.73, and on Last Letter exceeds the natural-language mean of 83.11. The JSON-mode condition is the worst in all three rows, so the study's poor JSON-mode results are what the ordering was built on, and they do not generalize to schema-constrained decoding. The reported standard deviations are wide and these comparisons are not tests of pairwise significance; they are sufficient to refute a universal ordering, not to establish the reverse one. Two further limitations are load-bearing: the study predates current-generation reasoning models, and the authors have published updates in response to methodological critique, which do not establish an intrinsic monotonic penalty either. The direction is well established; the effect size on current models is unconfirmed.arxiv.org/abs/2408.02442 ↗
- 76Anthropic, "Structured outputs," Claude Platform documentation together with the Claude Sonnet 5 migration guide, https://platform.claude.com/docs/en/models/sonnet-5/migration-guide, and "Extended thinking," https://platform.claude.com/docs/en/build-with-claude/thinking, all verified September 25, 2026. Vendor documentation. Cited as product fact for four things the surrounding chapters depend on: that the structured-output API is constrained decoding and is documented in those words (§11.1); the documented capitalization caveat on enum, under which a normally completed response can return a casing the schema does not allow, with no special stop reason to signal it (§32.2); that thinking is on by default on Opus 5 and Sonnet 5, so a thinking block can precede the first text block and response parsing must select by block type rather than by index (§10.4); and that assistant prefill returns HTTP 400 on Sonnet 5, Opus 5 and the Fable and 4.6-through-4.8 families, with structured outputs or a system instruction as the documented replacement (§31.5). Numbered after reference 75 because these pages were added during a later revision; the reference numbering is append-only and no longer strictly follows first citation.platform.claude.com/docs/en/build-with-claude/structured-outputs ↗
- 26Anthropic model pricing, published in the pricing table of reference 12 and verified September 8, 2026. Source for all Claude per-model rates: Fable 5.1 $10/$50, Opus 5 $5/$25, Sonnet 5 $2/$10, Sonnet 4.6 $3/$15, Haiku 4.5 $1/$5 per MTok, with cache multipliers of 1.25× (5m write), 2× (1h write), and 0.1× read (0.025× on Fable 5.1 and Mythos 5.1).platform.claude.com/docs/en/about-claude/pricing ↗