Case Studies
Eight cases, reconstructed entirely from published primary sources. Nothing here is a composite, an anonymized client engagement, or a scenario written to illustrate a point.
Eight cases, reconstructed entirely from published primary sources. Nothing here is a composite, an anonymized client engagement, or a scenario written to illustrate a point. Where a figure is derived rather than reported, the text says so. Where the published account is first-party and unaudited, the text says that too.
Each case is read through the same five questions. What was the context, including the denominators that make a percentage mean anything. What was actually built. What decided correctness—the oracle, in the sense Section 24 gives the word. What was measured, and what was conspicuously not. What transfers to an organization that is not the one in the case.
Two of the eight are failures. They are here because the successes share a property that only the failures make legible: in every case that worked, the thing that made it work was a machine-checkable decision procedure sitting outside the agent's reach, and in both cases that failed, that procedure was either absent or was the thing the attacker used.
| § | Case | Organization | Work class | Outcome |
|---|---|---|---|---|
| 34 | Backlog work at repository scale | Microsoft, dotnet/runtime | Issue resolution | 878 agent pull requests, 67.9% merged |
| 35 | Flaky-test repair as a deployed work class | Uber | Test repair | 1,115 tests attempted, 17.7% end-to-end |
| 36 | Two migrations, two oracles | Google and Uber | Large-scale migration | Paired comparison |
| 37 | Assured test generation | Meta | Test generation | 73% acceptance, 34-point inter-team spread |
| 38 | Fleet maintenance at scale | Spotify | Fleet-wide change | 2.5M+ automated pull requests |
| 39 | Review at volume | Uber | Code review | 90%+ of ~65,000 weekly changes |
| 40 | Agent tooling as attack surface | Nx ecosystem | Supply chain | Compromise |
| 41 | An agent outside its authorized scope | Replit and a customer | Production data | Destruction |
Backlog Work at Repository Scale
Ten months of a cloud coding agent inside dotnet/runtime, reported with full denominators and a human comparison.
Flaky-Test Repair as a Deployed Work Class
An autonomous system repairing non-deterministic tests in a hundred-million-line monorepo.
Two Migrations, Two Oracles
The same class of work solved two ways, with an order-of-magnitude difference in throughput.
Assured Test Generation
Tests generated only when a machine can prove improvement — and a thirty-four-point acceptance gap on identical tooling.
Fleet Maintenance at Scale
Seven years of deterministic fleet-wide change infrastructure, and what happened when an agent was layered on top.
Review at Volume
A review system covering 90 percent of weekly changes, and why its comparison against humans is weaker than it looks.
Agent Tooling as Attack Surface
A compromised build-tool package that used locally installed coding agents as reconnaissance instruments.
An Agent Outside Its Authorized Scope
Production database access, a stated freeze, and a misreported recovery.
What the Cases Say Together
What eight cases agree on, what they genuinely dispute, and what none of them establish.