The Control Maturity Model
Five levels, scored as the minimum across dimensions rather than the average.
Section 4 gives the organizational progression—Assisted, Delegated, Owned. This section gives the control maturity that gates it. They are different axes, and the relationship between them is the framework’s central operating rule:
An organization may not occupy a stage its control maturity does not support. Level 2 across all dimensions is the entry condition for Stage 2. Level 3 is the entry condition for Stage 3.
Five levels across fifteen dimensions: the four mechanics domains in Part II, the seven lifecycle stages, and the four cross-cutting domains.
An organization’s overall level is the minimum across dimensions, not the average. Agentic delivery fails at its weakest boundary, and an arithmetic mean conceals precisely the gap that matters—an organization with excellent development controls and no oracle architecture is not at Level 3, it is at Level 1 with a good development story.
18.1 Level Definitions #
Realistic targets. Level 2 is the entry condition for Stage 2 and the minimum defensible bar for any AI participation at tier A2 or above in a repository that reaches production. Level 3 is the entry condition for Stage 3 and the target for organizations operating at scale or in a regulated sector. Level 4 is where autonomy expansion becomes genuinely data-driven.
18.2 The Grid #
| Dimension | L0 | L1 | L2 · MINIMUM BAR | L3 | L4 |
|---|---|---|---|---|---|
| M1Loop Governance | Loops unbounded and uninstrumented | Limits set per team, inconsistently | Turn and budget ceilings on every A2+ deployment; termination typed | Early-abort detection; constraints verified to survive compaction; cost distributional | Loop telemetry drives autonomy decisions; divergence halted automatically |
| M2Work Class Governance | Delegation is per-task habit | Classes named informally | Register with owners, tiers, and declared oracles | Oracles immutable and demonstrated; stop conditions halt; review cost funded | Class outcomes trended per language and repository; tiers adjust on evidence |
| M3Validation Architecture | Tests only, generated alongside code | Independence stated as policy | Layered checks; oracle independent of generator | Oracle outside agent write scope; held-out signal; hardened evaluation boundaries | Mutation-gated; differential verification; visible-to-held-out gap monitored |
| M4Platform | Licenses and guidance | Shared tooling, no substrate | Isolated runtime, governed tool gateway, agent identity registry | Entitlement-preserving context substrate; eval infrastructure; cost per task enforced | Fleet observability joined to outcomes; guardrails graduate from teams to platform |
| S1Planning | AI unpriced and unplanned | Tooling cost budgeted as licenses | Participation statement, tier, and consumption forecast per initiative | Absorption capacity modeled from telemetry and governing staffing | Autonomy and investment decisions driven by measured control efficacy |
| S2Requirements | Prose requirements only | Acceptance criteria present, inconsistently | Machine-legible criteria; NFRs numeric; probabilistic thresholds defined | Criteria drive automated verification; traceability enforced end to end | Requirements quality measured against downstream defect and rework rates |
| S3Design | Architecture tacit | Documented, not machine-surfaced | Constraints in a source-controlled artifact; blast radius documented; models pinned | Constraints enforced by blocking checks; environment separation verified in the running system | Architectural drift detected continuously; reuse and duplication actively managed |
| S4Development | Agents on human credentials, unscoped | Distinct accounts, manual scoping | Unique agent identity; configuration reviewed as code; dependencies allowlisted; self-approval blocked | Ephemeral credentials; provenance signed at authorship; review risk-tiered and integrity-measured | Review integrity governs merge rate automatically; agent scope adapts to measured behavior |
| S5Verification | Generated tests, generated code, same session | Tests reviewed by humans | Evaluation harness per model-backed capability; independence policy; qualification before write access | Evaluation gates merge and release; judges validated; toolchain red-teamed on a schedule | Statistical acceptance throughout; contamination-resistant corpora; qualification re-run on version change |
| S6Release | Authorship unknowable from the record | Manifest lists components | Release record distinguishes agent-authored change; SBOM produced; artifacts signed | Provenance verified as a gate; separation of duties tested; model changes gated as releases | Rollback on quality distribution automated with measured time-to-effect |
| S7Operations | Aggregate billing, no agent telemetry | Logs collected, unanalyzed | Agent telemetry centralized; cost attributed; kill switches documented | Continuous validation against production; kill switches tested with measured time-to-effect; NHI review separate | Behavioral baselining; comprehension and maintainability funded from measured signal |
| X1Identity | Shared or human credentials | Distinct accounts, manual lifecycle | One identity per deployment; documented lifecycle; registry with named owners | Task-scoped delegation; inspectable chains; automated deprovisioning | Continuous authorization; risk-adaptive scope |
| X2Provenance | None | Attribution by convention, unverified | Signed commits with session linkage; SBOM at release | Attestation chain from source to release, verified at consumption | Attestation queryable across the estate; forensic reconstruction rehearsed |
| X3Evaluation | None | Ad hoc, per team | Registry of corpora, metrics, thresholds, owners | Gating, validated judges, statistical acceptance, results as telemetry | Continuous validation; corpora derived from live traffic; efficacy measured |
| X4Governance | None | Policy written | Review body with authority to block; named owners at A2+; exception register | Independent assurance; approval integrity measured; separate NHI reporting | Autonomy decisions data-driven; workforce and absorption managed as capacity |
18.3 Using the Grid Honestly #
Two instructions that determine whether this is a management tool or a slide.
- Assess against evidence, not intent. Every cell above corresponds to an artifact named in the relevant section’s evidence list. A dimension is at a level when the artifacts exist and can be produced, not when the practice is believed to be in place. The gap between those two states is the single most common finding in independent assurance.
- Report the minimum and the distribution. Report the overall level as the minimum, and separately show the profile across all fifteen dimensions. The profile is what tells leadership where the next investment goes; the minimum is what tells them what they can defend.