Turning a Prompt Into an Artifact
The threshold at which a prompt becomes an artifact, how it sits on disk, templating without injection, what to log, and who owns it.
Move prompting from something typed to something versioned, tested, and owned—the transition that separates a demo from a system.
13.1 The Threshold #
The moment a prompt is used more than once by more than one person, it becomes code. It has a version, a test, an owner, and a changelog, or it has silent drift.
The failure this prevents is both specific and common: someone improves a prompt for their case, ships it, and degrades three other cases nobody was tracking. Without a versioned artifact and a regression suite, that failure is undetectable until a user reports it.
13.2 Structure on Disk #
prompts/
incident-triage/
v3/
system.md # the instruction, version-controlled
schema.json # the output contract
examples.jsonl # few-shot cases, one JSON object per line
eval.jsonl # test cases with expected properties
CHANGELOG.md # what changed, why, and what the eval showed
v2/ # kept — you will need to roll back
current -> v3 # symlink or config pointer
Versioning as directories rather than through git history alone matters, since you will frequently need two versions live at once: v3 for new traffic, v2 for a customer who has validated against it. A pointer file makes that a config change rather than a deploy.
13.3 Templating Without Injection #
Never build a prompt through string concatenation on untrusted input. The reasons are the whole of Section 36, but the discipline belongs here.
from dataclasses import dataclass
from pathlib import Path
import json, hashlib
@dataclass(frozen=True)
class PromptTemplate:
name: str
version: str
system: str
schema: dict
@classmethod
def load(cls, root: Path, name: str, version: str) -> "PromptTemplate":
d = root / name / version
return cls(
name=name,
version=version,
system=(d / "system.md").read_text(),
schema=json.loads((d / "schema.json").read_text()),
)
@property
def fingerprint(self) -> str:
"""Stable hash of the system text and schema — the template, not the
rendered prompt. That is deliberate: it is a template-version
identifier, so substituting different user data must NOT change it.
Log it with every call; it is how you correlate a quality change
with a prompt change."""
body = self.system + json.dumps(self.schema, sort_keys=True)
return hashlib.sha256(body.encode()).hexdigest()[:12]
def render(self, **fields) -> list[dict]:
"""User content is placed in its own delimited block. It is never
interpolated into the instruction text."""
parts = [{"type": "text",
"text": f"<{k}>\n{v}\n</{k}>"} for k, v in fields.items()]
return [{"role": "user", "content": parts}]
The fingerprint is the piece most teams skip, and the one they most regret skipping. When quality changes, the first question is always “did the prompt change?” A twelve-character hash on every logged call answers it in seconds.
The rendering approach, with user data in named blocks and never interpolated into instruction prose, is the structural defense described in §8.4 and §36.3. It does not make injection impossible. It makes the boundary explicit enough that the model can be told about it.
13.4 What to Log #
Per call, at minimum:
{
"prompt_name": "incident-triage",
"prompt_version": "v3",
"prompt_fingerprint": "a1b2c3d4e5f6",
"model": "claude-sonnet-5",
"model_snapshot": "claude-sonnet-5-20260812",
"usage": {
"input_tokens": 412,
"cache_read_input_tokens": 18400,
"cache_creation_input_tokens": 0,
"output_tokens": 1180
},
"stop_reason": "end_turn",
"validation": { "passed": true, "attempts": 1 },
"latency_ms": 3820
}
Four of those fields carry disproportionate weight. The fingerprint correlates quality with prompt changes. The model snapshot correlates quality with silent vendor changes. The three-way token split is the only way to compute real cost and cache hit rate (Section 25). And the validation attempt count is a leading indicator of drift.
13.5 The Improvement Loop #
The discipline comes down to this:
Every validation failure, every thumbs-down, every escalation becomes a case.
This is the step people skip, and skipping it is why prompt fixes regress.
Multiple simultaneous changes make attribution impossible.
The new case will pass; the question is what else moved.
“v4: added negative example for empty input. Eval: 0.87 → 0.91 overall, no regressions.”
Use a symbolic link or configuration file to point to the current version; leave the previous version for validation.
Section 34 covers how to build the eval set so that step four means something statistically.
13.6 Ownership #
A prompt without a named owner will rot. The owner is responsible for the eval set, the changelog, and the decision to accept a regression in one dimension for a gain in another. Treating prompts as ownerless shared configuration produces exactly the outcome you would expect from ownerless shared configuration.