The Turn, the Loop, and the Session
The three units of agentic work — turn, loop, session — and why history growing monotonically is the thing that sets the bill.
Define the three units of agentic work precisely because cost, quality, and failure all attach to different ones, and conflating them is why so many teams cannot diagnose their own agents.
4.1 The Turn #
A turn is one request-response pair. The model receives the fully rendered context and produces text, tool calls, or both, and that is the whole of it.
Turns constitute the billing unit, and every turn re-sends the entire context. There is no server-side conversation state on a raw API unless you opt into it, and that single fact is what makes caching load-bearing rather than optional.
4.2 The Loop #
A loop is a sequence of turns driven by tool use. The model requests a tool, the harness executes it, the result is appended to history, and the model is called again. The loop terminates when the model returns a response carrying no tool calls, or when a limit fires.
def run_loop(client, tools, tool_impls, messages, max_turns=25):
"""Minimal agentic loop. Everything a coding agent does is this,
plus context management, plus permissions."""
for turn in range(max_turns):
response = client.messages.create(
model="claude-opus-5",
max_tokens=4096,
tools=tools,
messages=messages,
)
messages.append({"role": "assistant", "content": response.content})
if response.stop_reason != "tool_use":
# end_turn is the only one of these that can carry a finished
# answer; max_tokens truncated, refusal declined, pause_turn wants
# continuation. It says how the turn ended, not whether the task
# succeeded — that is the oracle's job (§10.4). Collapsing all four
# into "completed" throws away both distinctions at once.
return response, response.stop_reason
results = []
for block in response.content:
if block.type == "tool_use":
results.append({
"type": "tool_result",
"tool_use_id": block.id,
"content": tool_impls[block.name](**block.input),
})
messages.append({"role": "user", "content": results})
return None, "max_turns_exceeded"
Two properties of that code deserve attention, since they are where real systems go wrong.
History grows monotonically. Every turn appends the assistant’s output and the tool results. Turn twenty carries everything from turns one through nineteen. Vendor modeling of a fifty-turn session puts input at roughly 5,000 tokens per turn for the first ten turns, 20,000 for turns eleven through thirty, and 35,000 for turns thirty-one through fifty, with input outnumbering output twenty to twenty-five times.7 Cost per turn is therefore roughly linear in turn index, and the marginal turn is invariably the expensive one. An agent that is going to fail is cheapest to stop early.
Termination is typed and most people discard the type. The sample function above returns the provider’s own stop reason or "max_turns_exceeded" because end_turn, max_tokens, refusal and pause_turn are four different outcomes and only the first one can carry a finished answer. Even then it is not an acceptance: end_turn means the assistant ended its turn naturally, which a wrong answer does too. Record provider termination and task acceptance as two fields, and let the oracle populate the second (§10.4). Real harnesses distinguish more still: success, max turns exceeded, budget exceeded, execution error, structured-output retry exhaustion. An organization recording only that “the agent finished” has discarded its most useful diagnostic signal. Instrument why the loop stopped.
4.3 The Session #
A session is everything from context initialization to context reset. It spans many loops, and it is therefore the unit that accumulates instruction files, tool schemas, cached prefixes, compaction events, and, critically, the model’s own errors.
The session is where quality degrades, and the mechanism is a specific one. Peer-reviewed work identifies self-conditioning: models become more likely to make mistakes when the context contains their errors from prior turns. This is not a long-context artifact: injecting artificial error histories reproduces the effect, and larger models are more susceptible despite better long-context handling.8 That result has a stated boundary: the reasoning-trained models in the same work eliminated the measured effect entirely, so self-conditioning is established on non-reasoning models and the two findings below are separate evidence that does not depend on it.8 Separately, multi-turn degradation: conversational performance falls by roughly 39 percent, decomposed into a small aptitude loss and a large increase in unreliability, driven by models making early assumptions and over-relying on them.9 And in analysis of real agent trajectories, the single largest failure category is false premise at 30.7 percent: the agent forms a wrong belief early and never revisits it.10
Those three independent findings converge on a single mechanism: early wrong commitment, left uncorrected.
What that looks like in a real trajectory, abbreviated from a failed session:
turn what happened input tokens
---- --------------------------------------------- -------------
1 read src/auth/session.py 4,100
2 concluded the bug is in token refresh 6,800
3 edited refresh_token(); ran tests; 3 failed 11,200
4 edited refresh_token() again; same 3 failures 15,900
5 widened the search: read 4 more files 34,600
6 edited refresh_token() a third time; 3 failures 39,100
7 "let me reconsider the approach" 41,800
8 edited refresh_token() a fourth time; 3 failures 45,300
... (nine further turns, same three tests) ~92,000
actual cause: a fixture in tests/conftest.py froze the clock
The decisive error is turn 2. Everything after it is competent work on the wrong premise, and each turn makes the premise harder to abandon because the context now contains eight pieces of evidence that the agent has been working on token refresh. Turn 5 is the tell: widening the search is a signal the agent is stuck, and it arrives three turns after the point where a reset would have been cheap.
Against that trace, the three findings above stop being abstract. Turn 2 is the false premise. Turns 3 through 8 are self-conditioning. The absence of any stop is the 82 percent below.
The consequence for how a session is run, and the practical heart of Section 24:
- Restart beats steer. Mid-task correction degrades outcomes because the failed attempt contaminates the context. Clearing and restarting from a clean state with better instructions is mechanistically justified. Nudging a confused agent is not.
- Early abort beats late detection. In the trajectory study, 82 percent of failed trajectories keep executing after the point at which the failure is already empirically unrecoverable, and the first observable signal surfaces roughly ten steps after the decisive error, which is the earlier of the two events.10 A long-running session signals trouble more often than diligence.
- Design for recovery, not avoidance. Seventy-one percent of successful trajectories recover from at least one error.10 Errors are normal; unrecovered errors are the problem.
4.4 What This Means When You Are Typing #
Translated into behavior at the keyboard, this yields four habits:
Front-load constraints into the first message of a session, not the fifth. They compete with less.
When the agent goes wrong twice on the same thing, clear the context rather than explaining again. The third explanation is competing with two failures.
Batch related work into one session and unrelated work into separate ones. This is a cache decision as well as a quality one—see Section 22.6.
Set both a turn ceiling and a spend ceiling. At least one major SDK defaults both to unlimited, and most people never change it.
4.5 Failure Modes #
- Governing the output and ignoring the loop. Everything expensive and everything diagnostic happens before the final answer exists.
- Unlimited loops by default. Two dimensions, both frequently unset.
- Steering a failing agent. It feels responsible and it makes outcomes worse.
- Treating session length as productivity. It correlates with failure, not thoroughness.
References cited in this section
4 of 81 · numbering matches the PDF
- 7Vendor modeling of token accumulation across a fifty-turn agentic session: roughly 5,000 input tokens per turn for turns one through ten, 20,000 for turns eleven through thirty, and 35,000 for turns thirty-one through fifty, with input outnumbering output twenty to twenty-five times. Vendor-published; the shape is the point rather than the specific figures. Cited via reference 1.
- 8Peer-reviewed work identifying self-conditioning: models become more likely to err when the context contains their own prior errors. Not a long-context artifact—injecting artificial error histories reproduces the effect, and larger models are more susceptible despite better long-context handling. The same work finds that reasoning-trained models eliminate self-conditioning entirely, which is a meaningful qualification on the practices in §24. Cited via reference 1.
- 9Multi-turn conversational degradation of approximately 39 percent, decomposed into a minor aptitude loss and a large increase in unreliability, driven by models making early assumptions and over-relying on them. The study tested conversational generation rather than agentic coding trajectories, so transfer to coding is plausible and unproven. Cited via reference 1.
- 10Failure taxonomy across 1,794 complete agent trajectories and more than 63,000 execution steps, seven models and three scaffolds. Source of the false-premise rate (30.7 percent), the epistemic/competence/environment breakdown (57.9 / 32.8 / 9.4 percent), the finding that 82 percent of failed trajectories continue executing after the failure is empirically unrecoverable, that the first observable signal surfaces roughly ten steps after the decisive error, and that 71 percent of successful trajectories recover from at least one error. The paper separates three events, and an earlier revision of this document collapsed two of them: the decisive error, t_lock (the point after which no correct recovery is observed), and the first observable signal. The 82 percent continued-execution figure is measured from t_lock; the ten-step lag is measured from the decisive error. Also the source of the prefix-monitor results cited in §25.4: roughly 2 to 3 percent false positives and about 82 percent precision at recognizing a locked-in failure, against recall under thirty percent and a median lead time of zero relative to t_lock, with only 3.7 to 8.7 percent of failures flagged before lock-in. That is failure confirmation rather than loop detection, and §25.4 is scoped accordingly. The strongest published failure taxonomy for coding agents. Cited via reference 1.