Skills and Progressive Disclosure
Progressive disclosure as a mechanism, why the description is the whole game, the trust problem, and measuring whether a skill fires.
Explain the mechanism that makes on-demand guidance cheap, how to write one that fires correctly, and what to inspect before running someone else’s.
16.1 The Mechanism #
A skill is simply a directory containing a Markdown file with frontmatter. The frontmatter—name and description—is loaded into context always. The body is loaded only when the model judges the skill relevant to the current task.
.github/skills/release-a-service/
SKILL.md # frontmatter always loaded; body loaded on demand
checklist.md # loaded only if SKILL.md references it and the model follows
scripts/verify.sh # executed, not loaded
That split constitutes the entire value proposition. A 1,400-token procedure costs you roughly 30 tokens per turn as metadata until something invokes it, against 1,400 on every turn in an instruction file—a forty-fold reduction in the standing cost of identical guidance.
Be exact about what happens after it fires because the saving is on loading, not on residency. In ordinary inline invocation the rendered body enters the conversation as a message and stays there on subsequent turns; Claude Code does not re-read the file, re-invoking with identical rendered content adds a note rather than a second copy, and auto-compaction re-attaches the most recent invocation of each skill after the summary within a budget.41 So the body is not billed on the invoking turn alone—it is billed on every turn that retains it, at whatever rate that position is cached at, until it ages out or the session resets. A skill that runs in a forked subagent context is the case where the body genuinely does not persist in the parent, and that is opt-in. The forty-fold figure is the right comparison for a skill that never fires; the comparison for one that fires early in a long session is much closer to the instruction file.
The format is an open specification, portable across hosts.41
16.2 The Description Is the Whole Game #
Skills fail in two directions, and both are description problems. Either the skill never fires or it fires on everything, and in both cases the description rather than the body is the cause.
---
name: release-a-service
description: >
Use when releasing a service to staging or production, cutting a release
branch, or preparing release notes. Covers the version bump, changelog
generation, tag format, staging soak requirements, and the production
promotion gate. Do not use for hotfixes — see hotfix-procedure.
---
Four properties of that description each do real work:
- It names triggers in the user’s vocabulary. “Releasing,” “cutting a release branch,” “release notes”—the words someone would actually type.
- It enumerates contents. The model can judge relevance from what the skill covers without loading it.
- It carries an explicit exclusion. The “do not use for hotfixes” clause is what stops over-firing, and it is the clause most often missing.
- It points at the alternative. Which makes the exclusion actionable rather than a dead end.
A description that will fail looks like this: description: Release process. Two words, no triggers, no scope, no exclusion. It will either never fire or fire on any mention of the word “release,” including “release the lock.”
16.3 Writing the Body #
Three rules govern the body.
A skill is a runbook. If the content explains why rather than how, it belongs in documentation.
make release VERSION=1.4.0 beats “run the release target with the version.”
A skill can point at checklist.md or a script, which the model loads or executes only if the procedure reaches that point, nesting progressive disclosure inside itself.
---
name: add-a-database-migration
description: >
Use when adding, modifying, or reverting a database migration; when a
schema change is needed; or when a PR requires a data backfill. Covers
file naming, the up/down requirement, index strategy on large tables,
and the review gate. Not for seed data — see seed-data-management.
---
# Adding a migration
## 1. Generate the file
make migration NAME=add_user_preferences
This creates `migrations/NNNN_add_user_preferences.sql` with the correct
sequence number. Do not create the file by hand — the sequence number is
derived from the current head and hand-numbering causes collisions.
## 2. Write both directions
Every migration has `-- +migrate Up` and `-- +migrate Down` sections. A
migration without a working `Down` will fail CI.
## 3. Large tables
If the target table exceeds 1M rows, index creation must use
`CREATE INDEX CONCURRENTLY` and must be in its own migration with no other
statements. `CONCURRENTLY` cannot run inside a transaction block.
Check row counts first:
make db-rowcount TABLE=users
## 4. Verify before committing
./scripts/verify-migration.sh migrations/NNNN_*.sql
This applies up, applies down, applies up again, and diffs the schema. If
it does not exit 0, the migration is not ready.
## 5. Review gate
Migrations touching `accounts`, `billing_*`, or `audit_log` require review
from the data team. Add the `needs-data-review` label.
See @checklist.md for the full pre-merge checklist.
What that skill does not contain matters as much: any explanation of why reversible migrations are the policy, any description of the database, and any style guidance; what remains is instructions for doing a thing.
16.4 The Trust Problem #
A skill is, in effect, executable guidance authored by a third party. It can instruct the model to run commands, read files, and call tools, and it enters the context with the same standing as guidance you wrote yourself. Vendor documentation carries this as an explicit warning, and it is the essential point of this section.
There are four ways it goes wrong, but only one of them requires malice:
| Failure | Requires a bad actor? |
|---|---|
| A skill instructs the agent to run a destructive command as part of a legitimate procedure | No |
| A skill’s description is broad enough that it fires constantly and pollutes context | No |
| A skill silently updates upstream and its behavior changes under you | No |
| A skill contains instructions designed to exfiltrate data or install a backdoor | Yes |
The third is the underrated case. A skill tracked against a moving upstream reference is a standing grant to a third party to change what your agent does, without review, at any time.
16.5 Before You Install #
Eight checks follow. They take roughly two minutes, and they are the difference between guidance and an incident.
The description is marketing; the body is the instruction.
Enumerate every one. Ask whether you would run it yourself.
Any curl, wget, or fetch is data leaving your environment.
Writes to ~, /etc, or CI configuration are a red flag.
A benign SKILL.md can reference a malicious script.
A skill with one commit from an account created last week is not a dependency.
Install at a specific commit or tag. Tracking upstream is a standing permission grant.
An over-broad description will fire on unrelated work and cost you on every one.
16.6 The Unattended Case #
A skill committed to a repository’s skills directory becomes available to any agent operating on that repository, including automated code review and cloud agents that run with no human watching. That is a materially different risk posture from a skill installed on one developer’s machine.
The governance consequence, stated once: repository-committed skills are a change to the execution environment of every unattended agent on that repository, and they should go through the review that implies. This document does not develop the control framework for that; the SDLC framework does.1
16.7 Measuring Whether It Works #
While three layers exist here, most teams instrument none of them. The commands below cover the first two, reporting what is installed and what the installation costs in fixed-block tokens.
# What is installed
/skills
# What fired this session — usually surfaced in the transcript or a debug flag
# What it cost — the delta in fixed-block tokens before and after installation
/context
The most useful single measurement is this: run /context on a fresh session, install the skill, run it again. The delta is your recurring metadata cost. Then run a task that should trigger it and confirm it fired. A skill that never fires is pure overhead, and a substantial fraction of installed skills never fire.
References cited in this section
2 of 81 · numbering matches the PDF
- 41Agent Skills specification and vendor documentation for skill directories on Claude Code, GitHub Copilot, and OpenAI Codex. Open specification plus vendor documentation; cited as product fact for the progressive-disclosure mechanism.agentskills.io ↗
- 1Joshua Davis, The AI SDLC: An Operating Model, Control Framework, and Maturity Progression for Engineering Organizations Building With Agents, v1.0, September 2026 The companion framework covering governance, controls, and organizational absorption. Cited here for scope boundaries rather than for evidence.jdav.is/ai-sdlc ↗