The AI SDLC / Part VI
Part VI · §34–42

Case Studies

Eight cases, reconstructed entirely from published primary sources. Nothing here is a composite, an anonymized client engagement, or a scenario written to illustrate a point.

9 sections 55 min read

Eight cases, reconstructed entirely from published primary sources. Nothing here is a composite, an anonymized client engagement, or a scenario written to illustrate a point. Where a figure is derived rather than reported, the text says so. Where the published account is first-party and unaudited, the text says that too.

Each case is read through the same five questions. What was the context, including the denominators that make a percentage mean anything. What was actually built. What decided correctness—the oracle, in the sense Section 24 gives the word. What was measured, and what was conspicuously not. What transfers to an organization that is not the one in the case.

Two of the eight are failures. They are here because the successes share a property that only the failures make legible: in every case that worked, the thing that made it work was a machine-checkable decision procedure sitting outside the agent's reach, and in both cases that failed, that procedure was either absent or was the thing the attacker used.

§CaseOrganizationWork classOutcome
34Backlog work at repository scaleMicrosoft, dotnet/runtimeIssue resolution878 agent pull requests, 67.9% merged
35Flaky-test repair as a deployed work classUberTest repair1,115 tests attempted, 17.7% end-to-end
36Two migrations, two oraclesGoogle and UberLarge-scale migrationPaired comparison
37Assured test generationMetaTest generation73% acceptance, 34-point inter-team spread
38Fleet maintenance at scaleSpotifyFleet-wide change2.5M+ automated pull requests
39Review at volumeUberCode review90%+ of ~65,000 weekly changes
40Agent tooling as attack surfaceNx ecosystemSupply chainCompromise
41An agent outside its authorized scopeReplit and a customerProduction dataDestruction
PDF