Budgeting and Instrumentation
Cost is stochastic rather than linear. The minimum instrumentation, alarms worth having, runaway detection, and a defensible budget.
Cover what to measure, what the numbers mean, and why the obvious response to a cost surprise is usually the wrong one.
25.1 Cost Is Stochastic, Not Linear #
One finding breaks naive capacity planning outright: runs on the same task can differ by up to thirtyfold in total tokens.6
The same work established two further findings that destroy the intuitive planning approach. Accuracy often peaks at intermediate cost and saturates above it, so spending more does not monotonically buy correctness. And two separate results say that cost is hard to anticipate. Models forecast their own token consumption poorly, with correlations up to 0.39 and a systematic bias toward underestimating. And expert-rated task difficulty—SWE-bench-Verified’s estimate of how long a professional developer would need to resolve the issue—is only a weak predictor of what an agent actually spends.6 Keep those distinct: the second measures whether human effort categories predict agent tokens, and no one in that study was asked to forecast an agent’s bill. It is a reason to distrust difficulty as a proxy, not evidence about how well an expert could forecast agent cost if asked, which was not tested. Weak is also not useless—the categories carry some signal, and coarse signal beats none when you are sizing a budget.
The consequence is that per-task budgeting must be distributional rather than a point estimate. A mean is not a plan. Track p50, p95, and p99, and set alarms on the tail.
Spending more also does not guarantee scoring better. Two entries recorded from the same public leaderboard make the point: one model-scaffold combination at $366.81 for a 27.2 percent resolve rate against another at $67.09 for 38.0 percent.55 Two points refute “the dearer system wins”; they say nothing about the correlation across the population, which two observations cannot estimate at all. The leaderboard route they were read from no longer resolves either, so treat the pair as an illustration and the live board as the thing to re-read. Paying more is not a strategy.
25.2 The Minimum Instrumentation #
Four things matter here that most organizations ignore completely.
Success, max turns, budget exceeded, execution error, validation exhausted. A population that is 100 percent “success” is usually a taxonomy that is not instrumented; check that before believing it, since a small batch of genuinely easy work can also come back clean.
With the p95, not the mean.
, computed from reads, writes and fresh input after the provider adapter has made those three disjoint (§22.8).
Expect these to differ—§4.3’s trajectory work puts failed recoveries at more than twice the length of successful ones. Identical distributions are a prompt to audit how the outcome label is assigned and how the sample was drawn, rather than proof on their own that failure is going undetected.
Those four become one record written at the end of every session. The type below carries them as termination, cost_usd, turns, and a computed cache_hit_rate, alongside four correlation keys that make the population queryable: the task class, so distributions can be compared like with like; the model snapshot, so a vendor-side change is separable from your own (§33.5); the prompt fingerprint, so a quality shift can be tied to a prompt edit (§13.3–13.4); and wall time, so latency regressions surface without a second pipeline.
from dataclasses import dataclass, field
from enum import Enum
class Termination(str, Enum):
COMPLETED = "completed"
MAX_TURNS = "max_turns_exceeded"
BUDGET = "budget_exceeded"
TOOL_ERROR = "execution_error"
VALIDATION = "validation_exhausted"
ABORTED = "aborted_on_divergence"
@dataclass
class SessionRecord:
session_id: str
task_class: str
provider: str # which usage shape `usage` carries (§22.8)
model: str
prompt_fingerprint: str
termination: Termination
turns: int
usage: dict = field(default_factory=dict)
cost_usd: float = 0.0
wall_ms: int = 0
@property
def cache_hit_rate(self) -> float:
# `usage` holds the provider's native fields; §22.8's adapter is what
# makes reads/writes/fresh disjoint, and it runs here. Never read the
# raw fields directly, since OpenAI's `input_tokens` already contains
# the cached ones and summing them double-counts.
reads, writes, fresh = normalize(self.usage, self.provider)
total = reads + writes + fresh
return reads / total if total else 0.0
None of those four fields is useful on a single session. They are population statistics, and that is the point: you are not debugging one run, you are characterizing a distribution. A single expensive session tells you nothing because §25.1 established that cost varies up to thirtyfold on identical work. The same session sitting at the ninety-ninth percentile of its task class tells you a great deal.
If you remember nothing else from this chapter, remember the typed termination reason. It costs one enum, it is the field organizations most reliably omit, and it converts “the agents seem worse this week” from an impression into a query.
25.3 Which Alarms to Set #
| Alarm | Threshold | What it catches |
|---|---|---|
| Session cost above p99 | Rolling p99 by task class | Runaway loop |
| Cache hit rate drop | −15 points week over week | Prefix change, gateway, routing |
| Context threshold crossed | 272K on OpenAI, 200K on Gemini | Rate increase (§21.3) |
| Validation exhaustion rate | +50% week over week | Prompt or model drift |
| Turn count above p95 | Rolling p95 | Agent is stuck |
| Termination distribution shift | Any category moving >10 pts | Something changed upstream |
The third one is provider-specific and frequently missed. On a model with a context threshold, crossing it raises the rate on every request that stays above the line. Alarming on the crossing rather than on the resulting bill gives you a chance to intervene.
25.4 Runaway Detection #
Mechanisms exist, and the evidence for these mechanisms does not. Turn caps, budget caps, concurrency and spawn-depth limits, execution-time ceilings—all real, all implementable, none with a published measurement of how often production agents actually enter non-terminating loops, no loop-specific detector with a published false-positive rate, and no published cost-anomaly threshold.
Be precise about what has been measured because something adjacent has. The trajectory work behind §4.3 reports a general failure monitor at roughly 2 to 3 percent false positives and about 82 percent precision at flagging a run that has already locked in, with recall under thirty percent and a median lead time of zero.10 Note which event that lead time is measured against: t_lock, the point after which the paper observes no successful recovery, and not the decisive error, which precedes lock-in. Only 3.7 to 8.7 percent of failures are flagged before lock-in at all. That is a confirmation instrument rather than a runaway alarm—it tells you a run has failed, not that one is about to—and it is the nearest thing to a calibrated number in this area.
Build it, instrument it, and do not claim a benchmark you do not have. [unmeasured]
class RunawayGuard:
"""Both dimensions. At least one major SDK defaults both to unlimited."""
def __init__(self, max_turns=40, max_cost_usd=5.00, max_wall_s=1800):
self.max_turns, self.max_cost, self.max_wall = max_turns, max_cost_usd, max_wall_s
def check(self, turns, cost_usd, elapsed_s) -> Termination | None:
if turns >= self.max_turns: return Termination.MAX_TURNS
if cost_usd >= self.max_cost: return Termination.BUDGET
if elapsed_s >= self.max_wall: return Termination.BUDGET
return None
The early-abort case has better support. Eighty-two percent of failed trajectories keep executing after the point at which the failure is already unrecoverable—repairing the wrong problem, retrying a strategy that cannot work, or verifying an outcome that can no longer change—and the first observable signal surfaces roughly ten steps after the decisive error, which is the earlier of the two events.10 An agent that is going to fail is cheapest to stop early, and a long trajectory is a negative signal.
25.5 Why Uniform Caps Are the Wrong Response #
The most common reaction to a cost surprise is a uniform per-developer cap. It is reactive by definition and it cuts the wrong spending first.
The reason is that the headroom is widest precisely where acceptance is worst. If a team’s first-pass acceptance rate is 30 percent, cost per accepted outcome is 3.3× the cost of an attempt; at 85 percent it is 1.18×. Closing that gap is the ceiling on what better context, a stronger model or more verification could return, and a uniform cap constrains both teams identically while leaving the larger ceiling untouched.
Headroom is not the same as marginal return, and conflating the two gets allocation backward. It says nothing about what the next dollar buys, because raising acceptance usually raises attempt cost too, and the ratio is what matters:
| Change | Before | After |
|---|---|---|
| Attempt cost 1.00 → 2.00, acceptance .30 → .31 | 3.33 | 6.45 |
| Attempt cost 1.00 → 1.05, acceptance .85 → .99 | 1.18 | 1.06 |
The wide-headroom team got materially worse and the narrow-headroom one got better. So treat the multiple as potential worth investigating where acceptance is poor, and allocate on the measured change in cost per accepted outcome rather than on the headroom figure.
The alternative is spend shaped by task class and measured by outcome rather than by seat:
BUDGETS = {
"lint-fix": {"model": "claude-haiku-4-5", "max_cost": 0.10, "max_turns": 8},
"test-repair": {"model": "claude-sonnet-5", "max_cost": 0.75, "max_turns": 20},
"feature": {"model": "claude-sonnet-5", "max_cost": 3.00, "max_turns": 40},
"architecture": {"model": "claude-opus-5", "max_cost": 8.00, "max_turns": 60},
"incident-debug": {"model": "claude-opus-5", "max_cost": 15.00, "max_turns": 80},
}
That is a policy someone can defend in a budget conversation because each line is a claim about the value of the work rather than an arbitrary ceiling on a person. Construct §25.4’s guard from these per task class rather than relying on its defaults, which are tighter than the two largest classes here allow.
25.6 What a Sound Budget Looks Like #
Everything in this chapter converges on a single artifact: a spend policy someone can defend in a budget review without reaching for a percentage they cannot source.
A budget that holds up has four properties, and each comes from a preceding subsection. It is distributional, expressed as p50 and p95 by task class rather than as a mean because cost varies up to thirtyfold on identical work (§25.1). It is bounded on two dimensions, turns and dollars because at least one major SDK leaves both unlimited by default (§25.4). It is shaped by task value rather than by seat because a uniform cap constrains a team with wide headroom and one with none identically, and because where additional spend actually returns is a measured question rather than one headroom answers (§25.5). And it is denominated in accepted outcomes, not attempted ones, because a 40 percent cut in per-attempt cost that drops first-pass acceptance from 85 to 45 percent raises the real cost by 13 percent while lowering the invoice (§21.5). Denominate that cut in money rather than tokens: the lanes price differently, so a 40 percent cut in token count that falls mostly on output can cut the bill by two-thirds and come out ahead even on those acceptance numbers.
The last of these is the property that changes conversations. A team asked to cut agent spend by a third can always do it. Whether they should is a question the invoice cannot answer and cost-per-accepted-outcome can.
One lever is large enough to warrant its own chapter. Model tiers differ by roughly twenty to one on output price, so routing mechanical work to a small model is a five- to tenfold saving on that traffic at no quality cost, and unlike every other control here, it requires no one to change how they work. §27 covers the strategies, the published savings and what they were actually measured on, why agentic workloads are the hard case, and how to instrument it.
References cited in this section
3 of 81 · numbering matches the PDF
- 6Original work on agentic iteration economics establishing that agentic tasks consume roughly a thousand times the tokens of code chat, that input rather than output drives that cost, that runs on the same task differ by up to thirtyfold in total tokens, that accuracy frequently peaks at intermediate cost, and two separate results about anticipating cost that an earlier revision of this document merged into a claim about human forecasting. The paper tests model self-prediction directly: correlations up to 0.39, with systematic underestimation. Its human data are SWE-bench-Verified's expert estimates of how long a professional developer would need to resolve each issue, compared against agent token consumption; it finds that difficulty category is a weak predictor of spend. No human was asked to forecast agent tokens or cost, so the study does not establish that expert humans forecast task cost badly—only that human-effort difficulty transfers weakly as a proxy for it. Cited via reference 1, which contains the full source annotation. The thirtyfold variance figure is the single most consequential number for capacity planning in this document.
- 55Public agentic-coding leaderboard cost comparison: one model-scaffold pair at $366.81 for a 27.2 percent resolve rate against another at $67.09 for 38.0 percent. A counterexample to the assumption that spending more buys accuracy, and nothing more: an earlier revision of this entry said the pair establishes that cost and accuracy are not correlated, which two observations cannot do—they have a sample correlation of exactly −1, and one broken monotonicity says nothing about a population relationship. The associated methodological critique should be read in full: accuracy-only evaluation, inadequate holdout sets producing shortcut-taking, and a lack of standardization, with the finding that state-of-the-art agents are needlessly complex and costly. Cited via reference 1. The two price/score pairs were recorded without model or scaffold names, dataset variant, cost denominator, or a dated snapshot, and the leaderboard route they were read from now returns 404 (the surviving view, /swebench_verified_mini, is a different 50-task board). They are reproduced in §25.1 as an illustration rather than as a citable record, and that is their settled status rather than a gap awaiting an archived table: the pair reached this document through reference 1's first edition, recorded without the board, model-scaffold pair, dataset variant or snapshot date that would let a particular table be retrieved again. The methodological critique is what the reference actually supports.
- 10Failure taxonomy across 1,794 complete agent trajectories and more than 63,000 execution steps, seven models and three scaffolds. Source of the false-premise rate (30.7 percent), the epistemic/competence/environment breakdown (57.9 / 32.8 / 9.4 percent), the finding that 82 percent of failed trajectories continue executing after the failure is empirically unrecoverable, that the first observable signal surfaces roughly ten steps after the decisive error, and that 71 percent of successful trajectories recover from at least one error. The paper separates three events, and an earlier revision of this document collapsed two of them: the decisive error, t_lock (the point after which no correct recovery is observed), and the first observable signal. The 82 percent continued-execution figure is measured from t_lock; the ten-step lag is measured from the decisive error. Also the source of the prefix-monitor results cited in §25.4: roughly 2 to 3 percent false positives and about 82 percent precision at recognizing a locked-in failure, against recall under thirty percent and a median lead time of zero relative to t_lock, with only 3.7 to 8.7 percent of failures flagged before lock-in. That is failure confirmation rather than loop detection, and §25.4 is scoped accordingly. The strongest published failure taxonomy for coding agents. Cited via reference 1.