Model Routing
The routing premise, four strategies, what the numbers actually say, why agentic workloads are the hard case, and where to start.
Cover routing as an architecture rather than a config line: the strategies, the published savings and what they were measured on, and why agentic workloads are the hardest case for the technique.
27.1 The Premise #
Model tiers differ by roughly twenty to one on output price, and most traffic does not require the top tier. Routing is the practice of matching each request to the cheapest model that will handle it acceptably.
Verified September 8, 2026, per million tokens:26,52
| Tier | Example | Input / Output | Use for |
|---|---|---|---|
| Small | Claude Haiku 4.5 | $1.00 / $5.00 | Extraction, classification, formatting, summarization |
| Small | gpt-5.6-luna | $0.20 / $1.20 | Same |
| Mid | Claude Sonnet 5 | $2.00 / $10.00 | Most coding work |
| Mid | gpt-5.6-terra | $2.00 / $12.00 | Same |
| Frontier | Claude Opus 5 | $5.00 / $25.00 | Architecture, hard debugging, long-horizon |
| Frontier | gpt-5.6-sol | $4.00 / $20.00 | Same |
The ratio between the cheapest small model and the frontier tier on output is just under twenty-one to one. That spread, rather than the routing algorithm, is where the saving comes from.
The premise is sound and the savings arrive. The difficulty is deciding which requests the cheap tier handles well enough, and the published numbers are measured on workloads where that question answers itself more easily than yours will.
27.2 Four Strategies #
| Strategy | Decision point | Overhead | Best for |
|---|---|---|---|
| Static | Before the request, by rule | None | Task classes you can name |
| Classifier | Before the request, by model | A model call, typically sub-second | High-volume, varied difficulty |
| Cascade | After a cheap attempt | A full cheap call, always | High latency tolerance |
| Semantic | Before the request, by embedding | An embedding call | Topic-bounded domains |
Static routing is a lookup table from task class to model. It requires no infrastructure, and it captures most of the available saving in agentic systems because the task class is usually known before the work begins. Start here.
Classifier routing trains a small model to predict difficulty and picks a tier from that prediction. It adds a few hundred milliseconds before the main call and requires labeled data, though less of it hand-made than expected: the published router learned from roughly 65,000 public preference comparisons, and the golden-labeled set added on top to lift benchmark performance was about 1,500 examples, under two percent of the training data.60
Cascade routing sends everything to the cheap model first and escalates when the result fails a quality check. It needs no difficulty prediction because the cheap model’s own output supplies the signal. It also pays for the cheap call on every request, including the ones that were always going to escalate.
Semantic routing embeds the request and routes by proximity to labeled clusters. It works where the domain is bounded and the mapping from topic to difficulty is stable.
27.3 What the Numbers Actually Say #
RouteLLM learned routing policies from human preference data and reported over 85 percent cost reduction on MT-Bench while retaining 95 percent of the strong model’s performance, sending only 14 percent of queries to the frontier tier.60 Cascade approaches have reported reductions up to 98 percent on some benchmarks.61 A useful secondary finding is that routers transfer: policies trained on one model pair generalize to a different pair without retraining, which suggests the classifier is learning something about query difficulty rather than about specific models.60
Three significant qualifications apply, and they have considerable implications..
The 85 percent figure is MT-Bench with a specific model pair. On MMLU, a classifier-based router reached roughly 45 percent savings at comparable quality.60 Neither benchmark involves tool calls or structured outputs. MT-Bench does have a multi-turn component—it is eighty two-turn conversations, which is where its name comes from, and the routing paper’s own token-cost arithmetic sets that aside by assuming a single-turn setting rather than by the benchmark lacking one. Two turns of open-ended chat is still the wrong shape for a fifty-turn tool-using loop, and that gap, not an absence of multi-turn state, is what limits the transfer.
The number depends on how skewed your distribution is toward easy requests. A workload that is uniformly hard saves nothing because every request escalates.
Start conservative, measure savings and quality against your own eval suite, and tighten once a quality monitor has earned trust. Earlier revisions of this document quoted specific production savings, quality-retention and classifier-latency figures here; they came from accounts that cannot be named or dated, so they have been removed rather than dressed in a disclaimer.62 Your numbers are a property of your traffic mix in any case, which is why the sequence matters more than anyone else’s starting point. [unmeasured]
27.4 Why Agentic Workloads Are the Hard Case #
This is the part most routing material omits, and it is precisely the part that matters for anyone reading this document.
A routing failure in a chat product produces a worse answer, which the user re-asks. A routing failure inside an agent loop produces a malformed tool call: the downstream step parses invalid input, throws or produces a wrong result, and the error propagates. The retry then costs a frontier call anyway, on top of the cheap call, the failed tool execution, and the wasted turns.63
The published routing benchmarks make no claim about structured agent task completion.63 Extrapolating an MT-Bench conversation result to a fifty-turn agentic loop is not supported by the evidence, and the failure mode is asymmetric: a wrong route costs far more than a right one saves.
That asymmetry has three consequences.
Once a session commits to a model, keep it. Switching mid-session invalidates the entire cache (§23.1) and pays a cold start on top of whatever the routing saved.
In a system with sub-agents (§18), assigning cheap models to mechanical roles—extraction, summarization, classification, formatting—captures most of the available saving with no classifier, no routing failure mode, and no added attack surface.
Escalation after a failed test run is grounded. Escalation on a model’s self-reported confidence is the calibration failure of §31.4 wearing a cost-optimization hat.
27.5 Routing You Did Not Build #
Everything above assumes you are the one routing. On several platforms you need not be, and on at least one you are highly encouraged away from it.
Automatic model selection is the recommended path on GitHub Copilot rather than an advanced option. GitHub’s own material for the feature establishes the orchestration mechanism and its vendor-measured results; what it does not carry is adoption data. Neither the orchestration announcement nor its research-preview discussion supplies a request count or a share of paying users,64 so treat figures of that shape as unsourced until someone names the window and the denominator.
The recommendation matters more than the volume in any case. A default shapes what most teams end up running, and a recommended default shapes what they are asked to justify departing from.
That inverts the decision this chapter has been describing. The question is no longer merely which strategy to implement, but whether to delegate the choice at all, and on what terms.
What delegation gains and costs. It gains a routing policy maintained by someone with visibility you do not have: per-model capability measurements across a large traffic population, updated as models change. No individual team can replicate that. It costs the two things §27.7 says to instrument. You do not choose the tiers, so escalation rate is the platform’s to observe rather than yours. And unless the platform exposes its decisions, cost per accepted outcome cannot be attributed to a tier because you do not know which tier ran.
The category is broader than tier selection. Platform routing now extends past “which model” into “which workflow,” and the patterns being productized are the ones this chapter already named. One system currently in research preview composes three: a single model solving directly; a cascade in which an efficient model drafts and a quality gate decides whether to escalate; and a critique pattern in which one model drafts, an independent read-only critic from a different model family reviews, and the drafting model revises once.64
Two observations about that set, both bearing on material established earlier.
The critique pattern is not the self-critique §10.3 warns against. The measurements there concern a model reviewing its own output, where iterative self-review raised recall while collapsing signal-to-noise below 1.0. Using an independent critic from a different model family is the mitigation §34.3 prescribes for model-as-judge, and the distinction is the whole difference between the two.
The same vendor states that first-turn, single-prompt tasks are the best place to begin and that multi-turn performance is the next target.64 That is an independent arrival at §27.4’s argument, from a team building the system rather than critiquing it, and it leaves compound workflows inside long loops as the unsolved case.
Evaluating a routing product. Four questions, and the first is the one vendors answer least clearly.
Quality at a cost ceiling, cost at a quality floor, and latency are three different objectives producing three different policies. A system tuned for one will disappoint anyone expecting another.
Which model drafted, which reviewed, whether an escalation fired, and why. Without that, §27.7’s instrumentation is unavailable and an incident review has no trail to follow.
Per task class, per repository, or not at all.
Drafting, critique, revision, escalation, retry, and fallback all bill. A comparison counting only the final successful call understates the true figure, which is the error figure 11 exists to prevent.
On the published numbers. Routing products are evaluated by their vendors, against baselines their vendors select, and the resulting figures deserve the treatment §18.3 gives multi-agent claims. The preview referenced above reports improved verified task quality at substantially reduced estimated cost against a frontier baseline on one public benchmark, and the same publication notes that benchmark’s relative saturation and includes an internal benchmark nobody outside the company can reproduce.64 Those disclosures are creditable and they do not make the numbers independent. Treat them as a reason to run your own comparison rather than as a substitute for one.
A research preview is also not an architecture decision. Availability, pricing, execution patterns, and the models composed are all subject to change before general availability, and some previews do not reach it.
Where that leaves you. Evaluate platform routing before building anything, since it replaces engineering you would otherwise do, under a policy maintained by someone measuring more traffic than you will. Evaluate it the way you would evaluate any routing change: on cost per accepted outcome for your own task classes, over your own traffic, with the eval suite running (§34). If the platform will not tell you what it decided, that is not a minor gap in reporting. It is §27.7’s second metric made uncomputable, and it should weigh accordingly.
27.6 A Reference Implementation #
from dataclasses import dataclass
from enum import Enum
class Tier(str, Enum):
SMALL = "claude-haiku-4-5"
MID = "claude-sonnet-5"
FRONTIER = "claude-opus-5"
@dataclass(frozen=True)
class RoutePolicy:
tier: Tier
max_cost_usd: float
max_turns: int
escalate_to: Tier | None = None
# Static, by task class. No classifier, no latency, no routing failure mode.
POLICY = {
"extract": RoutePolicy(Tier.SMALL, 0.05, 4),
"classify": RoutePolicy(Tier.SMALL, 0.05, 4),
"format": RoutePolicy(Tier.SMALL, 0.05, 4),
"lint-fix": RoutePolicy(Tier.SMALL, 0.10, 8, escalate_to=Tier.MID),
"test-repair": RoutePolicy(Tier.MID, 0.75, 20, escalate_to=Tier.FRONTIER),
"feature": RoutePolicy(Tier.MID, 3.00, 40, escalate_to=Tier.FRONTIER),
"architecture": RoutePolicy(Tier.FRONTIER, 8.00, 60),
"incident-debug": RoutePolicy(Tier.FRONTIER,15.00, 80),
}
def run_routed(task_class: str, task: str, oracle, run_session) -> dict:
"""Escalate only on an oracle verdict, and only once. Never mid-session."""
policy = POLICY[task_class]
first = run_session(policy.tier, task, policy.max_cost_usd, policy.max_turns)
ok, feedback = oracle(first) # tests, typecheck, schema — §10.4
if ok or policy.escalate_to is None:
return {**first, "tier": policy.tier, "escalated": False,
"accepted": ok, "total_cost_usd": first["cost_usd"]}
# Fresh session on the stronger tier, and a continuation of nothing: the
# failed attempt would contaminate the context (§4.3) and the model change
# invalidates the cache regardless (§23.1).
stronger = run_session(policy.escalate_to, task, policy.max_cost_usd * 3,
policy.max_turns)
accepted, second_feedback = oracle(stronger) # the escalated run is a
# candidate, not a verdict
return {**stronger, "tier": policy.escalate_to, "escalated": True,
"accepted": accepted,
# Both attempts were billed. Reporting only the second one makes
# cost per accepted outcome (§27.7) impossible to compute.
"total_cost_usd": first["cost_usd"] + stronger["cost_usd"],
"first_attempt_feedback": feedback,
"final_feedback": second_feedback}
Three design decisions carry the weight here. Escalation starts a fresh session rather than continuing the failed one because carrying the failure forward is the self-conditioning problem in §4.3 and the model change invalidates the cache anyway. Escalation happens once because a two-step escalation that fails twice has spent more than going straight to the frontier tier would have. And the return value carries the summed cost of both attempts along with a verdict on the escalated one because a router that reports only what the winning tier spent, and never checks whether the winner was right, has removed the two inputs the next subsection needs.
27.7 Instrumenting It #
Routing without measurement is merely a cost increase with extra steps. Four metrics apply.
| Metric | Watch for |
|---|---|
| Escalation rate by task class | Above ~40%, the cheap tier is wrong for that class |
| Cost per accepted outcome, by tier | The only number that settles whether routing helped |
| Quality delta at each tier | Requires an eval suite (§34) |
| Router latency | Sits on the request path; measure your own, since no recoverable published figure exists |
The first row is the alarm that matters. A cascade escalating most of its traffic has handed back most of its saving, and past the crossing in the figure below it costs more than no routing at all. The failure is silent—the outputs are fine, the bill is higher, and nobody looks.
The second row supplies the denominator. Routing that cuts per-attempt cost 40 percent while dropping first-pass acceptance from 85 to 45 percent has raised cost per accepted outcome by 13 percent even though the invoice fell. Denominate the cut in money, not tokens: a routing change moves the rate as well as the count, so a 40 percent reduction in tokens and a 40 percent reduction in spend are different quantities and only the second supports that conclusion (§25.6).
The arithmetic, worked on a month of a single task class:
1,000 tasks/month, previously all on the frontier tier
cost per task $2.40
first-pass acceptance 85%
accepted outcomes 850
cost per accepted outcome $2,400 / 850 = $2.82
After routing 70% to the mid tier
700 mid-tier tasks x $0.95 = $665
300 frontier tasks x $2.40 = $720
invoice $1,385 <- down 42%, and this is the
number that gets reported
mid-tier acceptance 62% -> 434 accepted
frontier acceptance 85% -> 255 accepted
accepted outcomes 689
cost per accepted outcome $1,385 / 689 = $2.01 <- genuinely better
That routing worked, and it worked by less than the invoice suggests: 29 percent on the metric that matters against 42 percent on the one that gets quoted. Now run the same arithmetic with mid-tier acceptance at 30 percent instead of 62:
accepted outcomes 210 + 255 = 465
cost per accepted outcome $1,385 / 465 = $2.98 <- worse than the baseline
Same 42 percent cut to the invoice, and the change is now a loss. The line between those two outcomes sits at 33.7 percent mid-tier acceptance for these costs, and it is further down than intuition puts it: even at 45 percent, which looks like a rout, routing is still 14 percent ahead of sending everything to the frontier tier. Nothing in the token metrics distinguishes any of those cases. Only acceptance does, which is why it is the first thing to instrument and the last thing most teams add.
27.8 Where to Start #
On several surfaces, routing is on by default (§27.5). Measure it against a fixed tier before assuming you need to build anything.
Most of the saving, none of the infrastructure.
before changing anything else.
for classes where you have a real check.
only above roughly a hundred thousand requests a day, where the engineering pays for itself.
, and only where latency tolerance is high and the quality check is cheap and reliable.
27.9 Failure Modes #
- Routing per turn inside a session. Cache invalidation plus cold start on every switch.
- Escalating on an uncalibrated confidence, self-reported or logprob. Calibrate it or use an oracle (§31.4).
- Quoting a benchmark percentage to finance. MT-Bench is not your traffic.60
- Continuing the failed session on the stronger model. Carries the failure into the new context.
- No escalation-rate alarm. The most common way a cascade quietly costs more than it saves.
- Routing structured agent work by classifier. A malformed tool call cascades; the retry costs the frontier call anyway.
- Delegating routing to a platform that will not report its decisions. Cost per accepted outcome cannot be attributed to a tier you cannot see (§27.5).
- Treating a vendor’s routing benchmark as a prediction. Vendor-selected baselines on vendor-selected benchmarks, with at least one that nobody outside the company can reproduce (§27.5).64
References cited in this section
7 of 81 · numbering matches the PDF
- 26Anthropic model pricing, published in the pricing table of reference 12 and verified September 8, 2026. Source for all Claude per-model rates: Fable 5.1 $10/$50, Opus 5 $5/$25, Sonnet 5 $2/$10, Sonnet 4.6 $3/$15, Haiku 4.5 $1/$5 per MTok, with cache multipliers of 1.25× (5m write), 2× (1h write), and 0.1× read (0.025× on Fable 5.1 and Mythos 5.1).platform.claude.com/docs/en/about-claude/pricing ↗
- 52OpenAI, "Pricing," OpenAI API documentation verified September 8, 2026, with the per-model pages under https://developers.openai.com/api/docs/models/ verified September 24, 2026. Vendor documentation; source for all OpenAI per-model rates including the short-context and long-context tiers. The 272K-token threshold and the 2× input, 2× cache and 1.5× output multipliers are stated on the model pages rather than on the pricing page, which lists the two tiers without naming the boundary. One rate in §21.3's table is promotional rather than standing: the gpt-5.6-sol model page documents its $4.00/$20.00 as available at least through November 21, 2026, verified September 25, 2026, and describes it as a reduction against the prior generation. Re-check it before carrying it into a forecast.developers.openai.com/api/docs/pricing ↗
- 60Isaac Ong, Amjad Almahairi, Vincent Wu, Wei-Lin Chiang, Tianhao Wu, Joseph E. Gonzalez, M. Waleed Kadous, and Ion Stoica, "RouteLLM: Learning to Route LLMs with Preference Data," ICLR 2025, arXiv:2406.18665. Peer-reviewed. Source of the over-85-percent cost reduction on MT-Bench at 95 percent of the strong model's performance with 14 percent of queries routed to the frontier tier, the roughly 45 percent saving on MMLU with a classifier router, the cross-model-pair transfer finding, and the augmentation result: roughly 1,500 golden-labeled MMLU examples, described by the authors as less than 2 percent of the overall training data, added on top of roughly 65,000 public preference comparisons. Limitation, and it is load-bearing: the benchmarks are conversational, use a specific model pair, and make no claim about structured agent task completion.arxiv.org/abs/2406.18665 ↗
- 61Lingjiao Chen, Matei Zaharia, and James Zou, "FrugalGPT: How to Use Large Language Models While Reducing Cost and Improving Performance," arXiv:2305.05176, 2023. Preprint. Source of the cascade-routing approach and reported reductions up to 98 percent on the evaluated benchmarks. The headline figure is benchmark-specific and depends heavily on how skewed the query distribution is toward easy requests.arxiv.org/abs/2305.05176 ↗
- 62Practitioner accounts of production routing deployments. Blog-level sources of varying rigor, verified September 8, 2026; cited for the shape of the cost-quality dial rather than as measurement. Unresolved attribution: an earlier revision of this entry reported a conservative band of 25 to 35 percent savings at roughly 99 percent quality retention, and classifier latency of 300 to 500 milliseconds on the request path. No author, deployment, URL, measurement window or denominator could be recovered for any of those three figures, so they have been removed from this entry and from the §27.7 latency row rather than retained as a typical setting. Measure routing savings, quality retention and router latency on your own deployment; nothing in this document's routing recommendations depends on the withdrawn numbers.
- 63Analysis of routing failure modes specific to agent pipelines verified September 8, 2026. Vendor-adjacent commentary rather than measurement, and cited as reasoning rather than evidence. The argument is mechanical and holds independently: a routing failure in a conversational product yields a worse answer the user re-asks, while a routing failure inside an agent loop yields a malformed tool call whose error propagates downstream and whose retry costs the frontier call anyway. The related point, that role-based structural routing captures most of the saving without classifier overhead or routing failure modes, is the basis for the recommendation in §27.4.www.openlegion.ai/en/learn/llm-routing ↗
- 64GitHub, "Project HydraFusion: Frontier quality via multi-model orchestration," The GitHub Blog, September 4, 2026 and the accompanying research-preview announcement at https://github.com/orgs/community/discussions/206492, both verified September 11, 2026. Vendor announcement and vendor-published evaluation. Cited for the three composed execution patterns (single, cascade, and a critique pattern using an independent read-only critic from a different model family), for the statement that first-turn single-prompt tasks are the recommended starting point with multi-turn performance identified as future work, and for the cost-accounting methodology counting every invoked leg including drafting, critique, revision, escalation, retry, and fallback. The reported results—improved verified task quality at substantially lower estimated cost against frontier baselines—are GitHub's own, measured against GitHub-selected baselines under GitHub-selected conditions, and have not been independently replicated. GitHub discloses that one of the three benchmarks is relatively saturated and that a second is internal and derived from its own Copilot sessions, which no external party can reproduce. Neither page carries adoption or volume statistics: an earlier revision of this document attributed a June 2026 request count and a share of paying users to them, and that attribution was wrong. Do not restore figures of that shape without a source that names the window and the denominator. The feature is an experimental research preview limited to the Copilot CLI at the time of writing; availability, pricing, execution patterns, and composed models are all subject to change.github.blog/ai-and-ml/github-copilot/project-hydrafusion-frontier-quality-via-multi-model-orchestration ↗