Section 6 of 45 12 min read

The Harness

An agent is a model plus a harness, the harness rewrites what you wrote, the effect is measured and large, and you can see inside it.

Objective

Establish that the model is only one component of what answers you, that the layer around it rewrites what you wrote before the model ever sees it, and that this layer measurably changes outcomes.

6.1 Agent Equals Model Plus Harness #

A language model produces text. It does not open files, run tests, execute commands, or remember yesterday. The system that turns generated text into action, feeds results back, and decides whether to continue is the harness.

Vendor documentation is also clear on this. VS Code’s account of its own Copilot harness gives it three responsibilities: context assembly (building the prompt from a system message, the user’s query, workspace structure, conversation history, tool results, custom instructions, and memory from earlier sessions), tool exposure (declaring which tools the model may call, with schemas and descriptions, where the available set can change per request), and tool execution (validating arguments, running the tool, handling errors, formatting the result, and feeding it back).14

None of those three responsibilities belongs to the model, and yet all three determine what it produces.

Carry this framing: you do not use a model, you use a model inside a harness. When output changes, the model is one of at least two suspects, and it is frequently the innocent one.

6.2 The Harness Rewrites What You Wrote #

This is the part that surprises practitioners, and it is why this chapter follows nondeterminism rather than sitting in the platform section.

Between your keystroke and the API call, the harness inserts a system prompt you did not write, injects workspace state you never mentioned, retrieves files you did not name, appends tool schemas you did not choose, and, on several products, selects a different system prompt depending on which model you picked.

VS Code documents the per-model divergence concretely. Claude models are given replace_string_in_file for edits while GPT models are given apply_patch. Gemini needs explicit reminders to use tool-calling rather than narrating it, and breaks on orphaned tool calls in history. Some models work best with a concise system prompt; others need verbose, structured instruction to stay on track. The harness selects a different system prompt per model, and the divergence is finer than per-vendor—one Claude version gets a different prompt than the next.14

That has an immediate consequence for anyone comparing models. Switching the model dropdown does not hold the prompt constant and simply change the model. It changes the model and the system prompt and the tool names and the conversation-management behavior, simultaneously. Any conclusion you draw from that comparison is confounded.

The same holds across products. The prompt you type is identical in Copilot, Cursor, and Claude Code; what arrives at the model is not remotely the same object.

6.3 How Large the Effect Is #

Two independent lines of evidence bear on this, one academic and one industrial.

Academic. The founding result on agent-computer interfaces established that the same model scores materially differently under different interfaces, and that constrained, purpose-built tools with structured feedback outperform raw shell access.3 The harness is not a delivery mechanism for the model’s ability; it is a determinant of it.

The counter-demonstration from the same lineage is equally important and cuts the other way. A roughly hundred-line harness with no tools but bash, a linear history, no context management, no sub-agents, and no memory scores within a few points of far more elaborate scaffolds.3 Harness sophistication is not monotonically good, and anyone selling it should be asked to beat that baseline.

Industrial. VS Code and OpenAI ran a two-week controlled experiment on live traffic, splitting agent sessions across one control and two treatment groups at 25/25/25, testing whether nudging the model to explore less and validate sooner would make it faster and cheaper without making it worse. Treatment A added a single compact instruction to limit exploration before acting; Treatment B restructured the prompt into explicit before-first-edit and after-first-edit sections.15

MetricTreatment ATreatment B
p50 time to first edit−2.88% (2.0s faster, p=0.027)−5.68% (3.9s faster, p=2e−5)
p95 time to first edit−1.93% (not significant)−9.30% (38.8s faster, p=1e−10)
p95 total tokens per turn−5.19% (p=0.016)−7.64% (p=0.0003)
Average tool calls per turn−3.19% (0.77 fewer, p=0.009)−8.54% (2.04 fewer, p=1e−12)
10-minute code survival−0.40% (not significant)−0.44% (p=0.049)
Commit survival−0.48% (not significant)+0.68% (not significant)

Treatment B shipped as the default.15

The table yields three findings, and most readers miss the third.

A system prompt change, with no model change and no tool change, moved tail latency by nine percent, tail token consumption by nearly eight percent, and tool calls per turn by more than eight percent, all at p-values that leave no room for argument, and the harness is the only thing that changed.

The published result includes an unfavorable movement. Ten-minute code survival fell slightly under both treatments, and under Treatment B the drop was just barely significant. The team shipped anyway, judging the efficiency gains larger and more robust than a lightly significant quality movement. Whether you agree with that trade is beside the point; that it is visible is the reason this is credible evidence rather than marketing.

Furthermore, the effect sizes are modest in absolute terms. A few percent on latency and tokens is a win at platform scale, and not the difference between a working agent and a broken one. Harness tuning is optimization, not transformation, consistent with the hundred-line-baseline result above.

6.4 A Vocabulary Collision #

Section 4 defined the turn as one request-response pair, the loop as a sequence of turns driven by tool use, and the session as everything up to a context reset. That is the API-level vocabulary and it is what this document uses throughout.

Product documentation frequently uses the same words differently. VS Code calls the turn the user-visible chat exchange, in which you send one message and the agent eventually replies, and calls one pass through the tool-calling loop a round, with the full set of rounds being the run.14

ConceptThis documentVS Code harness docs
One model request-responseTurnRound
The tool-calling cycleLoopRun
One user message to final reply(one loop)Turn
Everything to a context resetSession(session)

Neither usage is wrong. When you read a vendor’s cost or latency figures, check which unit they mean. “Tokens per turn” differs by an order of magnitude depending on whose definition is in force, and a “24 tool calls per turn” figure is a statement about a whole loop, not about one model call.

6.5 Who Owns What #

Section 2.4 split visibility across surfaces. The harness question, however, is narrower and considerably more useful: for each behavior, is this something you can change?

BehaviorOwnerYour lever
System prompt contentHarness, but often overridableClient-dependent: Claude Code exposes --system-prompt and --append-system-prompt, each with a file variant
Per-model system prompt selectionHarnessNone directly
Tool names, schemas, and descriptionsHarness (plus your MCP servers)Which servers you connect
Which files get retrieved and injectedHarnessNaming files explicitly in your prompt
Cache breakpoint placementHarnessNot editing the prefix mid-task
Compaction threshold and summarization promptHarness (threshold sometimes configurable)When you reset, and re-injected constraint files
Loop termination limitsHarness, usually configurableTurn and budget caps
Instruction files, skills, subagents, hooksYouAll of Part III
Model selectionYouRouting by task class

The left column is long, and very little of it belongs to the practitioner. That is the picture, and it explains why Part III is organized the way it is: the mechanisms in that part are precisely the surface the harness exposes for you to influence a system you do not control.

This also explains a common frustration. A prompt that works beautifully in one product and poorly in another has usually not encountered a worse model. It has encountered a different system prompt, a different retrieval strategy, and a different set of tool names.

6.6 Seeing Inside It #

One can see considerably more than most practitioners realize, and looking once settles what a great deal of speculation cannot.

Product inspection surfaces. VS Code ships a Chat Debug View that exposes the prompts, tool calls, and results behind an agent run, and a tools picker showing what is available for a given request.14 Most CLI agents expose a context command (§14.5 in the configuration chapter) and a verbose or debug flag. Use them on a fresh session before you type anything; whatever is reported is the fixed block you pay for on every turn.

Proxy inspection. Where a product supports bring-your-own-key or a custom base URL, you can route it through a local proxy and read the actual request. This is the most complete view available and it takes about twenty lines.

# A minimal logging proxy. Point the agent's base URL at http://localhost:8080
# and read exactly what it sends — system prompt, tools, the lot.
import json, os
from aiohttp import web, ClientSession

UPSTREAM = "https://api.anthropic.com"

async def forward(request: web.Request) -> web.StreamResponse:
    body = await request.read()
    try:
        parsed = json.loads(body)
        print(f"\n=== {request.method} {request.path} ===")
        print(f"model:  {parsed.get('model')}")
        print(f"tools:  {[t.get('name') for t in parsed.get('tools', [])]}")
        sys_blocks = parsed.get("system") or []
        if isinstance(sys_blocks, str):
            sys_blocks = [{"text": sys_blocks}]
        for i, b in enumerate(sys_blocks):
            print(f"system[{i}] ({len(b.get('text',''))} chars, "
                  f"cache={'yes' if b.get('cache_control') else 'no'}):")
            print(b.get("text", "")[:600])
        print(f"messages: {len(parsed.get('messages', []))}")
    except (ValueError, AttributeError):
        pass                                   # non-JSON body; forward untouched

    headers = {k: v for k, v in request.headers.items() if k.lower() != "host"}
    async with ClientSession() as s:
        async with s.request(request.method, UPSTREAM + str(request.rel_url),
                             data=body, headers=headers) as upstream:
            resp = web.StreamResponse(
                status=upstream.status,
                headers={k: v for k, v in upstream.headers.items()
                         if k.lower() not in
                         {"content-encoding", "content-length",
                          "transfer-encoding"}})
            await resp.prepare(request)
            async for chunk in upstream.content.iter_any():
                await resp.write(chunk)          # relay as it arrives; agents stream
            await resp.write_eof()
            return resp

app = web.Application()
app.router.add_route("*", "/{tail:.*}", forward)
web.run_app(app, host="127.0.0.1", port=8080)      # loopback only; see below

Three cautions apply. It relays whatever credentials the caller sent and authenticates nobody, so anything that can reach the port has an open forwarding path to the upstream API, and run_app binds every interface unless you pass host—which is why the call above pins loopback explicitly rather than trusting the default. Never run it in shared infrastructure. It prints system-prompt content to stdout, so keep it away from production traffic and off any shared terminal. And check your vendor agreement before intercepting a commercial product’s traffic, which is permitted for bring-your-own-key configurations and frequently not otherwise.

What you learn repays the effort. Most people are surprised by the size of the system prompt, by how many tools are declared, and by how much workspace state is injected without being asked for.

6.7 What This Changes About How You Work #

Six consequences follow, ordered roughly by how often each matters.

1
Attribute failures correctly.

When an agent behaves badly, the candidate causes are the model, the harness, your configuration, and your prompt—in roughly that order of how often people blame them and the reverse order of how often they are responsible. Check your own configuration first, then the harness’s behavior, then the model.

2
Do not benchmark models by switching the dropdown.

You are changing several variables at once (§6.2). A model comparison means something runs the same prompts through the same harness path, or through a raw API where you control the whole request.

3
Treat the harness version as an eval variable.

Products ship harness changes continuously—every VS Code release ships harness improvements alongside model updates.14 A quality change with no change on your side is at least as likely to be a harness update as a model update. Record the product version alongside the model snapshot in your telemetry (§13.4), or you will be unable to correlate.

4
Prompt to the intersection when portability matters.

Explicit structure, explicit constraints, and an explicit output contract survive translation across harnesses. Reliance on a particular tool name, a particular retrieval behavior, or a particular default does not.

5
Do not fight the harness.

Instructing an agent to “ignore your system prompt” or to use a tool it has not been given produces confusion, not compliance. Where a harness behavior genuinely blocks you, the answer is a different surface or the raw API, not a more clever sentence.

6
Use the raw API when you need determinism about the request.

Anything you intend to evaluate rigorously, version, or ship as a product belongs on the API, where the request is yours. Coding agents are for interactive work; they are a poor substrate for a system whose inputs must be known.

6.8 Building One #

Most readers should not build a harness at all. Those who should are building a product atop a model, automating a workflow that no existing agent fits, or running evaluations that demand full control of the request.

If you do, two results ought to govern the design. Constrained, purpose-built tools with structured feedback beat raw shell access.3 And the hundred-line baseline—bash only, linear history, no context management, no sub-agents, no memory—comes within a few points of elaborate scaffolds.3

The synthesis follows: start at the baseline, add one mechanism at a time, and measure each addition against the version without it. Treat each added mechanism as a hypothesis, and drop the ones that do not pay.

# The minimum viable harness. Everything a product harness does is this
# plus context management, permissions, retrieval, and instrumentation —
# each of which should have to earn its place.
class Harness:
    def __init__(self, client, model, system, tools, impls, price,
                 max_turns=40, max_cost=5.00):
        self.client, self.model = client, model
        self.system, self.tools, self.impls = system, tools, impls
        self.price = price                      # usage -> dollars; see §21.4
        self.max_turns, self.max_cost = max_turns, max_cost

    def run(self, task: str):
        messages = [{"role": "user", "content": task}]
        cost, transcript = 0.0, []

        for turn in range(self.max_turns):
            resp = self.client.messages.create(
                model=self.model, max_tokens=4096,
                system=self.system, tools=self.tools, messages=messages,
            )
            cost += self.price(resp.usage)
            transcript.append({"turn": turn, "usage": dict(resp.usage),
                               "stop": resp.stop_reason})
            messages.append({"role": "assistant", "content": resp.content})

            if resp.stop_reason != "tool_use":
                # The provider's own stop reason, not a flattened "completed":
                # max_tokens and refusal are failures wearing a success label.
                return {"status": resp.stop_reason, "cost": cost,
                        "turns": turn + 1, "transcript": transcript}
            if cost >= self.max_cost:
                return {"status": "budget_exceeded", "cost": cost,
                        "turns": turn + 1, "transcript": transcript}

            results = []
            for block in resp.content:
                if block.type == "tool_use":
                    try:
                        out = self.impls[block.name](**block.input)
                        results.append({"type": "tool_result",
                                        "tool_use_id": block.id, "content": out})
                    except Exception as exc:
                        # Structured failure beats a crash: the model can
                        # recover from an error it can read (§10.4).
                        results.append({"type": "tool_result",
                                        "tool_use_id": block.id,
                                        "is_error": True,
                                        "content": f"{type(exc).__name__}: {exc}"})
            messages.append({"role": "user", "content": results})

        return {"status": "max_turns_exceeded", "cost": cost,
                "turns": self.max_turns, "transcript": transcript}

Two things make that a harness rather than a loop: it returns a typed termination reason rather than just an answer (§4.2), and a tool that raises is converted into a readable error the model can act on rather than an exception that ends the run. Those two decisions carry more weight than any amount of orchestration built on top.

6.9 Failure Modes #

  • Treating the harness as a thin wrapper. It rewrites your prompt, chooses your tools, and picks a different system prompt depending on the model you selected.
  • Comparing models by switching the picker. Confounded by design.14
  • Blaming the model for harness behavior. The most common misattribution in this field.
  • Assuming vendor turn-based figures use your definition of a turn. They frequently do not (§6.4).
  • Building an elaborate harness before establishing that a simple one fails. The baseline is within a few points.3
  • Never once looking at what is actually sent. Twenty minutes with a debug view or a proxy replaces months of speculation.

References cited in this section

3 of 81 · numbering matches the PDF

  1. 14Julia Kasper, Megan Rogge, and Aaron Munger, "The Coding Harness Behind GitHub Copilot in VS Code," Visual Studio Code blog, May 15, 2026 Vendor-published engineering account; cited as product fact for the three harness responsibilities (context assembly, tool exposure, tool execution), the turn/round/run vocabulary, and the per-model divergence in system prompts, tool sets, and conversation management—including that Claude models are given replace_string_in_file while GPT models are given apply_patch, that Gemini requires reminders to use tool-calling and breaks on orphaned tool calls, and that the harness selects a different system prompt per model version. The same post describes VSC-Bench and reports an effort setting that consumed more tokens while resolving slightly fewer tasks, which independently corroborates the saturation finding in reference 6. As a vendor account of its own product it is authoritative on mechanism and self-interested on quality; the mechanism is what is cited here.code.visualstudio.com/blogs/2026/05/15/agent-harnesses-github-copilot-vscode ↗
  2. 3John Yang, Carlos E. Jimenez, Alexander Wettig, Kilian Lieret, Shunyu Yao, Karthik Narasimhan, and Ofir Press, "SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering," NeurIPS 2024, arXiv:2405.15793. Peer-reviewed. The founding result establishing that the same model scores materially differently under different agent-computer interfaces, and that constrained purpose-built tools with structured feedback outperform raw shell access. The counter-demonstration is equally important, and it comes from the same group's later mini-swe-agent ( rather than from the 2024 paper: a roughly hundred-line harness with no tools but bash, a linear history, no context management, no sub-agents, and no memory scores within a few points of elaborate scaffolds. Anyone selling harness complexity should be asked to beat that baseline.github.com/SWE-agent/mini-swe-agent ↗
  3. 15Visual Studio Code Team, "How Prompt Tuning Improved GPT-5.5 in VS Code," Visual Studio Code blog, July 6, 2026 Vendor-published controlled experiment: a two-week online A/B test on live agent traffic, one control and two system-prompt treatments at a 25/25/25 split, with per-metric effect sizes and p-values. Source of the figures in §6.3. Unusually credible for vendor research because it reports an unfavorable movement (a small, marginally significant decline in ten-minute code survival under the shipped treatment) alongside the favorable ones, and because the prompt text of both treatments is published in the open-source repository. Limitations to carry: a single model on a single product, quality proxied by code-survival rather than correctness, and no independent replication.code.visualstudio.com/blogs/2026/07/06/optimizing-vscode-coding-harness-model-providers ↗
PDF↓