Section 17 of 45 9 min read

Tools, Schemas, and MCP

Eagerly loaded tool definitions sit in the window. Writing one the model uses correctly, return values as context, authorization, and vetting a server.

Objective

Cover the largest fixed-cost item in most agentic sessions, how to write a tool the model uses correctly, and the data-flow question nobody asks until it is too late.

17.1 Eagerly Loaded Tool Definitions Are in the Window #

By default, every connected tool’s full schema rides in the request on every turn, whether or not that tool is ever called. This is the surprise that catches teams adopting MCP, and it is usually the largest line item in the fixed block. “By default” is load-bearing: three things that are easy to treat as one concept—what the client transmits, what enters the model’s context, and what you are billed for—come apart once schemas are deferred rather than sent up front, and the mitigations below are exactly that.

Measured in practice, a moderately equipped session:

SourceToolsApproximate tokens
Built-in agent tools83,500
A code-hosting MCP server3411,000
A ticketing MCP server217,500
A database MCP server124,000
Fixed block75~26,000

That is twenty-six thousand tokens consumed before anything happens. On a fifty-turn session that is 1.3 million tokens of tool definitions, and if the cache is working they cost a tenth of base input, but they still occupy the window, and window pressure is a quality problem independent of price (§12.1).

There are three mitigations, and each one helps.

The first arrived in the protocol itself. As of the 2026-07-28 revision, results from tools/list and the other list endpoints carry ttlMs and cacheScope fields, modeled on HTTP cache control, so a client knows how long a tool listing stays fresh and whether it is safe to share across users.42 That does not shrink the schemas rendered into a given request, but it does remove the need to re-fetch them, and it gives a harness a sanctioned way to hold a stable tool block across a session rather than rebuilding one.

The second is deferred tool loading, where schemas are fetched on demand rather than sent up front:

# Anthropic: defer tool schemas out of the cached prefix
tools = [
    {"type": "mcp_toolset", "mcp_server_name": "github",
     "default_config": {"defer_loading": True}},
    {"type": "mcp_toolset", "mcp_server_name": "jira",
     "default_config": {"defer_loading": True}},
]
// OpenAI: tool search with deferred loading; discovered tools append at the
// end of context, which preserves the earlier reusable prefix.
const response = await client.responses.create({
  model: "gpt-5.6-sol",
  tools: [
    { type: "tool_search" },
    ...ALL_TOOLS.map((t) => ({ ...t, defer_loading: true })),
  ],
  input,
});

And on the request side, disabling a tool without removing its definition preserves the cache:

// Wrong: removing the tool rewrites the tool block and invalidates everything.
// Right: keep definitions stable, restrict what is callable.
const response = await client.responses.create({
  model: "gpt-5.6-sol",
  tools: ALL_TOOLS,                          // unchanged, stays cached
  tool_choice: "none",                       // or:
  // tool_choice: { type: "allowed_tools", mode: "auto",
  //                tools: [{ type: "function", name: "read_file" },
  //                        { type: "function", name: "run_tests" }] },
  input,
});

The rule those two mechanisms share is counterintuitive, so look at it closely: to reduce what the model can do, restrict it; to reduce what the model reads, defer it. Never delete. Deleting a tool definition rewrites the block that renders first and invalidates everything after it.

And the decision that precedes both: every server you connect eagerly is a per-turn tax on every session, paid whether or not a single one of its tools is ever called. Twenty-six thousand tokens is not an abstraction—it is a fifth of a 128K window occupied before the work starts, which is a quality cost (§12.1) on top of the price. Deferring changes the shape of that bill rather than only its size: a deferred definition need not enter the model’s context at all until a search discovers it, so the tax becomes proportional to what the task actually reaches for. What deferral does not remove is the discovery step, the catalog entry the search reads, and the schema’s arrival mid-context once it is found. So connect what the task needs rather than what the catalog offers, and defer the rest instead of sending it.

17.2 Writing a Tool the Model Uses Correctly #

The tool schema is itself a prompt. It is read every turn and it is the only thing the model knows about what the tool does.

{
    "name": "search_incidents",
    "description": (
        "Search the incident database by service, severity, and time range. "
        "Returns at most 50 incidents, newest first. "
        "Use this before creating a new incident to check for duplicates. "
        "Does NOT return incident timelines — use get_incident_timeline for that."
    ),
    "input_schema": {
        "type": "object",
        "properties": {
            "service": {
                "type": "string",
                "description": "Exact service name as registered in the catalog. "
                               "Use list_services if unsure — a wrong name returns "
                               "empty rather than an error.",
            },
            "severity": {
                "type": "string",
                "enum": ["sev1", "sev2", "sev3", "sev4"],
                "description": "Omit to search all severities.",
            },
            "since": {
                "type": "string",
                "format": "date-time",
                "description": "ISO 8601. Defaults to 30 days ago if omitted.",
            },
        },
        "required": ["service"],
    },
}

That description accomplishes five things:

  • States what comes back, including the limit. A model that does not know the result cap will assume completeness.
  • States when to use it. “Before creating a new incident” is workflow guidance the model cannot infer.
  • States what it does not do, and names the alternative. This prevents the most common tool misuse, which is calling the wrong tool for an adjacent need.
  • Warns about a silent failure mode. “A wrong name returns empty rather than an error” prevents the model from concluding there are no incidents when it simply misspelled the service.
  • Documents defaults. Otherwise the model supplies its own, arbitrarily.

Tool descriptions are also where token discipline pays twice, since they are in the cached prefix on every turn, so verbosity is a subscription, which argues for density over length.

17.3 Return Values Are Context #

A tool result enters the window and remains there for the rest of the session. A tool that returns 40,000 tokens of JSON has permanently consumed a fifth of a 200K window, and every subsequent turn pays to re-read it.

# Bad: returns everything, forever.
def search_incidents(service, severity=None, since=None):
    rows = db.query(...)
    return json.dumps([dict(r) for r in rows])     # full records, all columns

# Good: return the minimum that lets the model decide what to fetch next.
def search_incidents(service, severity=None, since=None):
    rows = db.query(...)[:51]                      # one extra, to detect truncation
    summary = [
        {"id": r.id, "severity": r.severity, "title": r.title[:80],
         "opened": r.opened_at.isoformat()}
        for r in rows[:50]
    ]
    return json.dumps({
        "count": len(summary),
        "truncated": len(rows) == 51,
        "incidents": summary,
        "note": "Call get_incident(id) for full detail.",
    }, separators=(",", ":"))                       # compact separators, ~8% saved

The truncated flag is not optional here. Without it, a model receiving exactly fifty results has no way to know whether there were fifty-one, and will reason as though it has seen everything.

The corresponding SQL discipline, since database tools are among the most common and most costly:

-- Return the projection the tool needs, never SELECT *.
-- Cap in the query, not in application code, so the database does the work.
SELECT
    i.id,
    i.severity,
    LEFT(i.title, 80) AS title,
    i.opened_at
FROM incidents i
WHERE i.service_name = %(service)s
  AND (%(severity)s IS NULL OR i.severity = %(severity)s)
  AND i.opened_at >= COALESCE(%(since)s, NOW() - INTERVAL '30 days')
ORDER BY i.opened_at DESC
LIMIT 51;                    -- 51 so the caller can detect truncation at 50

The principle behind all three versions: a tool result is not a return value, it is a permanent addition to the context window. A function that returns forty thousand tokens has not been slow or wasteful; it has consumed a fifth of the model’s working memory for the remainder of the session, and every subsequent turn pays to re-read it.

Design tool outputs the way you would design a paginated API response for an expensive client: the minimum fields that support the next decision, an explicit truncation signal, and a documented way to fetch more. The model is a perfectly capable caller. It will fetch detail when it needs detail, but only if you told it that detail exists.

17.4 Where the Data Goes #

A tool call constitutes an outbound data transfer. The arguments land in a model request, the results land in the context window, and both land in the session log. Three deployment shapes give three different answers to “who can read this?”

ShapeData pathRead by
Local, over stdioProcess on your machine → model providerYou, the model provider
Remote, self-hostedYour network → your server → model providerYou, the model provider
Remote, third-party SaaSYour network → vendor → model providerYou, the vendor, the model provider

The third row is the one that surprises practitioners. Connecting a third-party MCP server means every argument you send it and every result it returns has passed through that vendor. For a server that reads customer data, that is a data-processing relationship requiring the same review as any other.

Two rules must be followed at all times. Never send credentials, secrets, or full customer records as tool arguments. And never assume a tool result is safe to act on—it is untrusted input that arrives inside your trusted context, which is the exact shape of the injection problem in Section 36.

17.5 Authorization #

Authorization is optional in the protocol, and the flow is specified for HTTP-based transports; a stdio server takes credentials from its environment instead. Where it applies, a protected MCP server acts as an OAuth 2.1 resource server and advertises its authorization server through protected-resource metadata. That indirection is the point, and it is what makes enterprise deployment possible. What it does not require: the authorization server is a separate role, which the specification permits you to host alongside the resource server or as a wholly separate entity.42

While two postures exist, only one of them survives scale:

Per-user OAuth 2.1Identity-managed access
SetupEach developer consents individuallyAdmin grants once, IdP-mediated
RevocationPer user, manuallyDeprovision in the IdP
AuditFragmentedCentralized
At 4,000 developersUnmanageableWorks

If you are piloting, per-user consent will serve. If you are rolling out, however, the identity-managed path is the only one carrying a working offboarding story. The two columns are also not exclusive: tokens can stay user-scoped while consent, audit and revocation run through the IdP. What does not scale is per-user consent with no central place to revoke it.

17.6 Vetting a Server #

Eight checks, requiring roughly fifteen minutes, cover it.

1
Read the manifest.

How many tools, what does each write, what are the schemas?

2
Count the tokens.

Connect it, run the context command, measure the delta. A 34-tool server is 11,000 tokens on every turn forever.

3
Enumerate write operations.

Read-only servers are a different risk class. Know which tools mutate.

4
Check the transport and the spec revision.

Streamable HTTP is current, and the legacy HTTP-with-SSE transport is now formally deprecated rather than merely discouraged. A server built against the 2026-07-28 revision differs materially from one built against its predecessor: the protocol core became stateless, the initialization handshake and session header were removed, and Mcp-Method and Mcp-Name headers became mandatory so that gateways can route without parsing bodies.42

5
Check the auth model.

Per-user OAuth or IdP-managed? Where do tokens live?

6
Determine where data goes.

Local, your infrastructure, or a vendor’s.

7
Check the publisher and update cadence.

Who can change this server’s behavior, and how would you find out?

8
Test in isolation first.

A scratch repository, not your monorepo.

17.7 Failure Modes #

  • Connecting servers because they are available. Each is a permanent per-turn tax.
  • Tools that return everything. The window is not a scratch buffer.
  • Tool descriptions that describe implementation. The model needs to know when to call it, not how it works.
  • Adding or removing a server mid-session where schemas load into the prefix. Rewrites the tool block, invalidates the whole cache; deferred tools append instead (§23).
  • Treating tool results as trusted. They are the most common injection vector.

References cited in this section

1 of 81 · numbering matches the PDF

  1. 42Model Context Protocol specification, revision 2026-07-28 verified September 11, 2026. Specification. Cited for transport (Streamable HTTP current, HTTP-with-SSE formally deprecated under the feature lifecycle policy), the OAuth resource-server authorization model, and the primitive set. This revision made the protocol core stateless, removing the initialize handshake and the Mcp-Session-Id header; required Mcp-Method and Mcp-Name headers on Streamable HTTP requests; added ttlMs and cacheScope to list and resource-read results; and deprecated OAuth 2.0 Dynamic Client Registration in favor of Client ID Metadata Documents. Servers written against 2025-11-25 or earlier will not be conformant without transport work. The community registry remains in preview with the API frozen at v0.1.modelcontextprotocol.io/specification/2026-07-28/changelog ↗
PDF↓