Section 2 of 45 6 min read

What a Prompt Actually Is

A prompt is a rendered document assembled by someone else. Roles are not permissions, and the part you typed is often under two percent of it.

Objective

Replace the mental model of “a prompt is the thing I type” with the accurate one because almost every prompting mistake at intermediate level is a consequence of the wrong model.

2.1 The Rendered Context #

A model never receives the message one composes. It receives a single flat sequence of tokens that the provider has assembled from several structured inputs. That assembly is called rendering, and understanding its order is the whole of Section 22’s caching material.

The canonical order, which holds with minor variation across every major provider:

[ provider-internal system content ]    ← you did not write this, and often cannot see it
[ tool definitions ]                    ← every schema, whether or not a tool is called
[ system / developer instructions ]     ← your standing instructions
[ conversation history ]                ← every prior user, assistant, tool_use, tool_result block
[ current user message ]                ← the thing you typed

Everything above the user message constitutes prefix. Cost, cache behavior, and a considerable share of quality behavior all follow from that single fact.

Here is the same structure as an actual request. Anthropic’s Messages API:

import anthropic

client = anthropic.Anthropic()

response = client.messages.create(
    model="claude-opus-5",
    max_tokens=2048,
    system=[                                   # rendered second, after tools
        {"type": "text", "text": "You are a code reviewer for a Python service."}
    ],
    tools=[                                    # rendered FIRST, before system
        {
            "name": "read_file",
            "description": "Read a file from the repository.",
            "input_schema": {
                "type": "object",
                "properties": {"path": {"type": "string"}},
                "required": ["path"],
            },
        }
    ],
    messages=[                                 # rendered last
        {"role": "user", "content": "Review the diff in src/auth.py."}
    ],
)

The inversion matters: tools is the last argument you think about and the first thing the model reads. That ordering has a direct consequence: where definitions load into the prefix, adding an MCP server mid-session invalidates the entire cache. Where tool search defers them instead, the same change appends rather than invalidating. Section 23 works through both.

The same shape in C#:

using System.Text.Json;
using Anthropic;
using Anthropic.Models.Messages;

var client = new AnthropicClient();

var response = await client.Messages.Create(new MessageCreateParams
{
    Model     = "claude-opus-5",
    MaxTokens = 2048,
    System    = new List<TextBlockParam>
    {
        new() { Text = "You are a code reviewer for a Python service." }
    },
    Tools     = [ new Tool
    {
        Name        = "read_file",
        Description = "Read a file from the repository.",
        InputSchema = new()
        {
            Properties = new Dictionary<string, JsonElement>
            {
                ["path"] = JsonSerializer.SerializeToElement(new { type = "string" }),
            },
            Required = ["path"],
        }
    }],
    Messages  = [ new() { Role = Role.User,
                          Content = "Review the diff in src/auth.py." } ]
});

InputSchema.Type is set to "object" by the constructor, and System needs the concrete List<TextBlockParam> because a collection expression will not convert to the union the parameter takes.

OpenAI’s Responses API uses different names for the same structure, calling them instructions and input rather than system and messages, and adds a developer role that the documentation places ahead of user messages.2

import OpenAI from "openai";

const client = new OpenAI();

const response = await client.responses.create({
  model: "gpt-5.6-sol",
  input: [
    { role: "developer", content: "You are a code reviewer for a Python service." },
    { role: "user", content: "Review the diff in src/auth.py." },
  ],
  tools: [
    {
      type: "function",
      name: "read_file",
      description: "Read a file from the repository.",
      strict: false,
      parameters: {
        type: "object",
        properties: { path: { type: "string" } },
        required: ["path"],
      },
    },
  ],
});

The structure is almost identical across all three sources, even where the vocabulary changes. As such, the following three facts are the foundation for the rest of this document.

You are assembling a document

Every field in those calls becomes a region of one token sequence, and the model reads the whole thing as continuous text. There is no structural separation between your instruction and the data you paste below it—a fact that is a quality problem in Part II and a security problem in §36.

The API determines position

You choose what goes in each field. The order those fields render in—tools before system, history before your message—is fixed. §22 is largely about working with that order, and §23 about what happens when you fight it.

The cheapest thing to change is the thing at the end

Everything before your message is prefix, and prefix is where both cost and cached state live. This is why “add a follow-up question” and “go back and reword the original” have costs that differ by orders of magnitude, despite feeling like the same act.

2.2 Roles Are Not Permissions #

A persistent misconception deserves correction here, and it carries a security consequence covered properly in Section 36. The system or developer role confers no privilege whatsoever. It marks a position in the token sequence that carries some post-training weight, nothing more. Models are trained to grant system content more authority than user content, and they generally do. However, nothing in the architecture renders them incapable of ignoring it.

The practical version follows directly. An instruction placed in a system prompt amounts to a strong preference rather than a constraint. If a rule must hold, it belongs in code that runs outside the model. A hook, a validator, or a schema will hold under adversarial input; a sentence asking the model to comply remains only a request. This distinction recurs throughout Part III and is the entire argument of Section 19.

2.3 What You Control Most Matters Least #

A realistic agentic session on a mid-sized repository breaks down like this:

ComponentTokensShare
Provider system content~2,0001.9%
Tool definitions (12 tools, 3 MCP servers)~14,00013.6%
Repository instruction file~1,8001.7%
Conversation history (turn 22)~85,00082.6%
Retrieved file contents in history(included above)
Your current message~1200.1%
Total~102,920100%

Of those figures, the current-message row is the one that surprises people. On turn twenty-two of a real session, the words one types amount to roughly one part in a thousand of what the model reads. One may compose the most carefully constructed sentence imaginable, and it will still compete for attention against eighty-five thousand tokens of accumulated history, some portion of which contains one’s own earlier mistakes.

Prompting still matters. What changes is where the leverage sits, and it sits overwhelmingly in what surrounds the message. Reducing that 85,000 to 30,000 by clearing context at a task boundary will alter the model’s behavior far more than any rewording of the 120 ever could. Section 24 covers how and when.

2.4 What You Can and Cannot See #

The boundaries differ by surface, and practitioners routinely assume access they do not possess:

LayerDirect APICoding agent (CLI/IDE)Chat product
Provider system contentNot visibleNot visibleNot visible
Tool definitionsYou author themInspectable, usually via a commandNot visible
System/developer instructionsYou author themPartly yours, partly the harness’sNot visible
Conversation historyYou own it entirelyManaged for you; inspectableManaged for you
Compaction behaviorYours to trigger, or server-side on requestAutomatic, threshold configurableAutomatic, opaque

The middle column is therefore where most professional work happens, and where most confusion lives, since responsibility is split between two parties. The practitioner owns the instruction file and the tool grants. The harness owns the system prompt, the compaction threshold, the cache breakpoint placement, and the repository context injection. Therefore, when output degrades, the first diagnostic question is which side of that line the cause sits on.

2.5 Failure Modes #

  • Writing to the model as though the message is all it sees. It produces prompts that repeat context already present and omit constraints the model genuinely lacks.
  • Treating the system prompt as enforcement. It is instruction rather than a control, and the first adversarial input can talk past it. What survives a compaction is a separate question with a different answer for each mechanism—a request-level system field is re-sent on every call, a root instruction file may be re-read from disk, and something said once in conversation can vanish into the summary (§24.3).
  • Assuming the harness is a thin wrapper. Harness design measurably changes outcomes independent of model choice, in both directions.3 Section 6 is entirely about this.

References cited in this section

2 of 81 · numbering matches the PDF

  1. 2OpenAI, "Text generation" and "Counting tokens," OpenAI API documentation and https://developers.openai.com/api/docs/guides/token-counting, verified September 25, 2026. Vendor documentation; cited as product fact for the Responses API's instructions and input parameters and its developer role, which the documentation prioritizes ahead of user messages, and for the POST /v1/responses/input_tokens endpoint. OpenAI describes that count as the exact number the model will receive, including the formatting tokens a local tokenizer cannot see, which is a stronger claim than Anthropic makes for its own counting endpoint.developers.openai.com/api/docs/guides/text ↗
  2. 3John Yang, Carlos E. Jimenez, Alexander Wettig, Kilian Lieret, Shunyu Yao, Karthik Narasimhan, and Ofir Press, "SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering," NeurIPS 2024, arXiv:2405.15793. Peer-reviewed. The founding result establishing that the same model scores materially differently under different agent-computer interfaces, and that constrained purpose-built tools with structured feedback outperform raw shell access. The counter-demonstration is equally important, and it comes from the same group's later mini-swe-agent ( rather than from the 2024 paper: a roughly hundred-line harness with no tools but bash, a linear history, no context management, no sub-agents, and no memory scores within a few points of elaborate scaffolds. Anyone selling harness complexity should be asked to beat that baseline.github.com/SWE-agent/mini-swe-agent ↗
PDF↓