Section 13 of 45 3 min read

Turning a Prompt Into an Artifact

The threshold at which a prompt becomes an artifact, how it sits on disk, templating without injection, what to log, and who owns it.

Objective

Move prompting from something typed to something versioned, tested, and owned—the transition that separates a demo from a system.

13.1 The Threshold #

The moment a prompt is used more than once by more than one person, it becomes code. It has a version, a test, an owner, and a changelog, or it has silent drift.

The failure this prevents is both specific and common: someone improves a prompt for their case, ships it, and degrades three other cases nobody was tracking. Without a versioned artifact and a regression suite, that failure is undetectable until a user reports it.

13.2 Structure on Disk #

prompts/
  incident-triage/
    v3/
      system.md            # the instruction, version-controlled
      schema.json          # the output contract
      examples.jsonl       # few-shot cases, one JSON object per line
      eval.jsonl           # test cases with expected properties
      CHANGELOG.md         # what changed, why, and what the eval showed
    v2/                    # kept — you will need to roll back
    current -> v3          # symlink or config pointer

Versioning as directories rather than through git history alone matters, since you will frequently need two versions live at once: v3 for new traffic, v2 for a customer who has validated against it. A pointer file makes that a config change rather than a deploy.

13.3 Templating Without Injection #

Never build a prompt through string concatenation on untrusted input. The reasons are the whole of Section 36, but the discipline belongs here.

from dataclasses import dataclass
from pathlib import Path
import json, hashlib

@dataclass(frozen=True)
class PromptTemplate:
    name: str
    version: str
    system: str
    schema: dict

    @classmethod
    def load(cls, root: Path, name: str, version: str) -> "PromptTemplate":
        d = root / name / version
        return cls(
            name=name,
            version=version,
            system=(d / "system.md").read_text(),
            schema=json.loads((d / "schema.json").read_text()),
        )

    @property
    def fingerprint(self) -> str:
        """Stable hash of the system text and schema — the template, not the
        rendered prompt. That is deliberate: it is a template-version
        identifier, so substituting different user data must NOT change it.
        Log it with every call; it is how you correlate a quality change
        with a prompt change."""
        body = self.system + json.dumps(self.schema, sort_keys=True)
        return hashlib.sha256(body.encode()).hexdigest()[:12]

    def render(self, **fields) -> list[dict]:
        """User content is placed in its own delimited block. It is never
        interpolated into the instruction text."""
        parts = [{"type": "text",
                  "text": f"<{k}>\n{v}\n</{k}>"} for k, v in fields.items()]
        return [{"role": "user", "content": parts}]

The fingerprint is the piece most teams skip, and the one they most regret skipping. When quality changes, the first question is always “did the prompt change?” A twelve-character hash on every logged call answers it in seconds.

The rendering approach, with user data in named blocks and never interpolated into instruction prose, is the structural defense described in §8.4 and §36.3. It does not make injection impossible. It makes the boundary explicit enough that the model can be told about it.

13.4 What to Log #

Per call, at minimum:

{
  "prompt_name": "incident-triage",
  "prompt_version": "v3",
  "prompt_fingerprint": "a1b2c3d4e5f6",
  "model": "claude-sonnet-5",
  "model_snapshot": "claude-sonnet-5-20260812",
  "usage": {
    "input_tokens": 412,
    "cache_read_input_tokens": 18400,
    "cache_creation_input_tokens": 0,
    "output_tokens": 1180
  },
  "stop_reason": "end_turn",
  "validation": { "passed": true, "attempts": 1 },
  "latency_ms": 3820
}

Four of those fields carry disproportionate weight. The fingerprint correlates quality with prompt changes. The model snapshot correlates quality with silent vendor changes. The three-way token split is the only way to compute real cost and cache hit rate (Section 25). And the validation attempt count is a leading indicator of drift.

13.5 The Improvement Loop #

The discipline comes down to this:

1
Collect failures.

Every validation failure, every thumbs-down, every escalation becomes a case.

2
Add the case to the eval set before fixing it.

This is the step people skip, and skipping it is why prompt fixes regress.

3
Change one thing.

Multiple simultaneous changes make attribution impossible.

4
Run the full eval, not the new case.

The new case will pass; the question is what else moved.

5
Record the delta in the changelog.

“v4: added negative example for empty input. Eval: 0.87 → 0.91 overall, no regressions.”

6
Ship behind the pointer, keep the previous version live.

Use a symbolic link or configuration file to point to the current version; leave the previous version for validation.

Section 34 covers how to build the eval set so that step four means something statistically.

13.6 Ownership #

A prompt without a named owner will rot. The owner is responsible for the eval set, the changelog, and the decision to accept a regression in one dimension for a gain in another. Treating prompts as ownerless shared configuration produces exactly the outcome you would expect from ownerless shared configuration.

PDF↓