The AI SDLC / Part IV / §18
Section 18 of 44 5 min read

The Control Maturity Model

Five levels, scored as the minimum across dimensions rather than the average.

Section 4 gives the organizational progression—Assisted, Delegated, Owned. This section gives the control maturity that gates it. They are different axes, and the relationship between them is the framework’s central operating rule:

An organization may not occupy a stage its control maturity does not support. Level 2 across all dimensions is the entry condition for Stage 2. Level 3 is the entry condition for Stage 3.

Five levels across fifteen dimensions: the four mechanics domains in Part II, the seven lifecycle stages, and the four cross-cutting domains.

An organization’s overall level is the minimum across dimensions, not the average. Agentic delivery fails at its weakest boundary, and an arithmetic mean conceals precisely the gap that matters—an organization with excellent development controls and no oracle architecture is not at Level 3, it is at Level 1 with a good development story.

18.1 Level Definitions #

L0Ad hoc
AI is in use; nobody knows where, by whom, or at what tier. No AI-specific controls. Policy, if it exists, is aspirational.
L1Aware
Partial inventory. Policy drafted. Controls manual, inconsistent, applied only to centrally procured tooling. Measurement is anecdotal.
L2Governed
Complete inventory of managed AI participation. Tiering applied. Identity, provenance, review, and evaluation baselines defined and enforced at deployment. Measurement exists and is human-paced.
L3Enforced
Controls enforced by the pipeline rather than by policy. Ephemeral credentials standard. Evaluation gates merge and release. Provenance emitted and verified. Absorption capacity measured and governing.
L4Adaptive
Continuous validation in production. Autonomy expansion driven by measured control efficacy. Independent assurance against evidence. Behavioral baselining per agent. Comprehension and maintainability managed as first-class operational concerns.

Realistic targets. Level 2 is the entry condition for Stage 2 and the minimum defensible bar for any AI participation at tier A2 or above in a repository that reaches production. Level 3 is the entry condition for Stage 3 and the target for organizations operating at scale or in a regulated sector. Level 4 is where autonomy expansion becomes genuinely data-driven.

18.2 The Grid #

DimensionL0L1L2 ·
MINIMUM BAR
L3L4
M1Loop GovernanceLoops unbounded and uninstrumentedLimits set per team, inconsistentlyTurn and budget ceilings on every A2+ deployment; termination typedEarly-abort detection; constraints verified to survive compaction; cost distributionalLoop telemetry drives autonomy decisions; divergence halted automatically
M2Work Class GovernanceDelegation is per-task habitClasses named informallyRegister with owners, tiers, and declared oraclesOracles immutable and demonstrated; stop conditions halt; review cost fundedClass outcomes trended per language and repository; tiers adjust on evidence
M3Validation ArchitectureTests only, generated alongside codeIndependence stated as policyLayered checks; oracle independent of generatorOracle outside agent write scope; held-out signal; hardened evaluation boundariesMutation-gated; differential verification; visible-to-held-out gap monitored
M4PlatformLicenses and guidanceShared tooling, no substrateIsolated runtime, governed tool gateway, agent identity registryEntitlement-preserving context substrate; eval infrastructure; cost per task enforcedFleet observability joined to outcomes; guardrails graduate from teams to platform
S1PlanningAI unpriced and unplannedTooling cost budgeted as licensesParticipation statement, tier, and consumption forecast per initiativeAbsorption capacity modeled from telemetry and governing staffingAutonomy and investment decisions driven by measured control efficacy
S2RequirementsProse requirements onlyAcceptance criteria present, inconsistentlyMachine-legible criteria; NFRs numeric; probabilistic thresholds definedCriteria drive automated verification; traceability enforced end to endRequirements quality measured against downstream defect and rework rates
S3DesignArchitecture tacitDocumented, not machine-surfacedConstraints in a source-controlled artifact; blast radius documented; models pinnedConstraints enforced by blocking checks; environment separation verified in the running systemArchitectural drift detected continuously; reuse and duplication actively managed
S4DevelopmentAgents on human credentials, unscopedDistinct accounts, manual scopingUnique agent identity; configuration reviewed as code; dependencies allowlisted; self-approval blockedEphemeral credentials; provenance signed at authorship; review risk-tiered and integrity-measuredReview integrity governs merge rate automatically; agent scope adapts to measured behavior
S5VerificationGenerated tests, generated code, same sessionTests reviewed by humansEvaluation harness per model-backed capability; independence policy; qualification before write accessEvaluation gates merge and release; judges validated; toolchain red-teamed on a scheduleStatistical acceptance throughout; contamination-resistant corpora; qualification re-run on version change
S6ReleaseAuthorship unknowable from the recordManifest lists componentsRelease record distinguishes agent-authored change; SBOM produced; artifacts signedProvenance verified as a gate; separation of duties tested; model changes gated as releasesRollback on quality distribution automated with measured time-to-effect
S7OperationsAggregate billing, no agent telemetryLogs collected, unanalyzedAgent telemetry centralized; cost attributed; kill switches documentedContinuous validation against production; kill switches tested with measured time-to-effect; NHI review separateBehavioral baselining; comprehension and maintainability funded from measured signal
X1IdentityShared or human credentialsDistinct accounts, manual lifecycleOne identity per deployment; documented lifecycle; registry with named ownersTask-scoped delegation; inspectable chains; automated deprovisioningContinuous authorization; risk-adaptive scope
X2ProvenanceNoneAttribution by convention, unverifiedSigned commits with session linkage; SBOM at releaseAttestation chain from source to release, verified at consumptionAttestation queryable across the estate; forensic reconstruction rehearsed
X3EvaluationNoneAd hoc, per teamRegistry of corpora, metrics, thresholds, ownersGating, validated judges, statistical acceptance, results as telemetryContinuous validation; corpora derived from live traffic; efficacy measured
X4GovernanceNonePolicy writtenReview body with authority to block; named owners at A2+; exception registerIndependent assurance; approval integrity measured; separate NHI reportingAutonomy decisions data-driven; workforce and absorption managed as capacity

18.3 Using the Grid Honestly #

Two instructions that determine whether this is a management tool or a slide.

  • Assess against evidence, not intent. Every cell above corresponds to an artifact named in the relevant section’s evidence list. A dimension is at a level when the artifacts exist and can be produced, not when the practice is believed to be in place. The gap between those two states is the single most common finding in independent assurance.
  • Report the minimum and the distribution. Report the overall level as the minimum, and separately show the profile across all fifteen dimensions. The profile is what tells leadership where the next investment goes; the minimum is what tells them what they can defend.
PDF