Metrics and Executive Reporting
What to report, what to keep separate, and the numbers that mislead most when combined.
Report non-human and human identity governance separately, always. A combined figure is the single most misleading number in this domain, because good human hygiene conceals a much worse machine posture.
Loop
- Termination reason distribution, trended—a population that is entirely “success” means the taxonomy is not instrumented
- Cost per task as a distribution, with the ninety-fifth percentile, not a mean
- Trajectory length for successful versus failed sessions
- Early aborts fired, and cost recovered
- Escalation rate per agent-week—suppression of escalation is a defect, not efficiency
Work class
- Merge, revert, and defect-escape rate by class, against a per-language and per-repository baseline
- Declared review cost against measured review hours, by class
- Classes throttled or tier-reduced on evidence in the trailing year
- Visible-suite to held-out-suite pass rate gap, by class—the reward-hacking indicator
Coverage
- Registered AI participants versus discovered AI participants; the gap is the metric
- Share of the agent estate emitting action telemetry
- Share of A2-and-above deployments with a named accountable human
- Share of repositories with enforced merge-gate configuration
Flow and absorption
- Change volume, split by agent-authored and human-authored
- Median and 95th-percentile time in review, trended
- Share of changes merged without human review
- Review hours required versus senior engineering hours available—the absorption ratio
- Merge conflict rate, with cross-agent concurrency broken out
Quality
- Change failure rate and incidents per unit of change, trended against AI adoption
- Defect escape rate by authorship class
- Security findings generated versus adjudicated; suppression and threshold changes reported explicitly
- Duplication and reuse trend
- Non-functional conformance: accessibility, performance budget adherence
Verification integrity
- Share of model-backed capabilities with a current, refreshed evaluation corpus
- Judge validation currency and chance-corrected agreement
- Mutation score for agent-authored modules, reported alongside line coverage
- Toolchain red team attack success rate, trended
Control efficacy
- Median credential lifetime in agent execution contexts
- Share of agents with zero standing credentials
- Share of releases with verified provenance
- Separation-of-duties test results, by quarter
- Kill-switch time-to-effect by agent class, from live test
Platform
- Context substrate coverage: share of repositories indexed and cataloged, entitlement enforcement verified
- Share of tool access flowing through the governed gateway rather than point integrations
- Agent telemetry generated against telemetry analyzed
- Guardrails graduated from team-local to platform capability
Workforce
- Entry-level hiring and internal promotion rate, trended over eight quarters
- Senior engineer tenure and attrition, read alongside review load
- Seeded-defect probe catch rate and median review time
- Comprehension: subsystems with a named owner who can explain them
Economics
- Cost per task and cost per merged change, trended
- AI consumption against forecast, with variance explained
- Total cost per engineer including consumption, against the license line
Governance
- Deployments by autonomy tier, with trend
- Deployments blocked or tier-reduced at review
- Active exceptional-access grants and their expiry dates
- Approval rate and median review duration for human-in-the-loop controls
TWO LEADING INDICATORS WORTH WATCHING ABOVE THE OTHERS
First, agent population growth rate against registry coverage growth rate: when the first exceeds the second, governance is degrading regardless of what any other metric says. Second, the absorption ratio: when required review hours exceed available senior engineering hours, every downstream quality metric will follow within two quarters, and the correct response is to reduce the merge rate rather than to exhort the team.