The First Ninety Days, Worked
Produces a sequenced ninety-day program with named artifacts rather than a list of intentions.
A sequenced ninety-day program with named artifacts, so that adoption becomes a set of deliverables rather than a list of intentions. Section 33 carries the sequence forward to month eighteen.
Starting. This chapter assumes an organization somewhere between Stage 1 and Stage 2 with undeclared pockets of Stage 3, which is where most are.
32.1 Weeks 1–2: Find Out What Is Actually Running #
You are measuring the present, not planning the future. Four artifacts.
The population inventory. Not what was procured. What is running. Pull from the identity provider, cloud control planes, repository integrations, developer environments, SaaS admin surfaces, tool-protocol servers, and direct model API calls originating in code and CI. Label the coverage gaps explicitly on the document itself, because the number will be read as complete.
The stage assessment. For every repository with agent participation, answer three questions: can agents self-approve, can they trigger CI without human authorization, can they push to protected branches. Then enumerate every configured bypass actor and the date each was added. Undeclared Stage 3 lives in that list.
The credential exposure result. Compromise a test agent execution context and enumerate every credential obtainable and its remaining validity. One afternoon. It usually reframes the program more than anything else in this chapter.
The oracle map. For every category of work already delegated, name the mechanism that decides correctness and whether the agent can modify it. Most organizations discover at least one work class running with no oracle at all, and that discovery is the finding that justifies the next ten weeks.
32.2 Weeks 3–4: Instrument Before Changing #
Changing anything before this point means you cannot tell whether it helped.
Loop telemetry live, per Section 25: the six events, typed termination, cost as a distribution.
The absorption ratio measured, per Section 29: review hours required against senior engineering hours available. Split change volume by authorship class if the data permits, and if it does not, record that as a provenance finding rather than an inconvenience.
Baselines captured, per Section 27 step 8: human merge rate, revert rate, and review cost per repository, per language. Every later comparison depends on this and it cannot be reconstructed retroactively.
The consumption baseline, per Section 31: actual spend per team, license and consumption separated, against budget.
32.3 Weeks 5–8: Close the Boundary and Build One Oracle #
Two workstreams in parallel, one security-shaped and one verification-shaped.
Close the merge boundary. No self-approval, no self-authorized CI execution, code-owner review on configuration and dependency manifests, concealed-character scanning on ingested content. Verify by demonstration, not documentation: attempt each and confirm failure. Then implement the four hardened-boundary interventions from Section 7.4, which is the highest measured return available in this framework at an 87.7 percent relative reduction in exploit rate with task success essentially unchanged.113
Give every A2-and-above agent a unique identity and a named accountable human, and eliminate agents operating under human credentials.
Build one oracle end to end, per Section 24, for the work class with the strongest published evidence. Do not start with documentation. Take it through all seven steps including the immutability test and the calibration run, because the point of this workstream is that the organization learns the procedure on a case where it works.
Set turn and budget ceilings on every deployment and default-deny unlimited at the platform level.
32.4 Weeks 9–12: Run One Work Class Properly #
One class, one team, full instrumentation, a scheduled decision date.
Register it per Section 27 with an owner, tier, oracle, entry conditions, guards, and stop conditions. Model the review cost and confirm the capacity exists before launch. Run the pilot against the baseline captured in weeks 3 and 4.
In parallel, three things that do not depend on the pilot:
- Instrument review integrity — approval rate, review duration, diff size at approval — and report it.
- Run the first toolchain red team: injection through an issue and a pull request description, a non-existent dependency, a configuration file with concealed characters, an attempted self-approval, and an attempted write to a test file.
- Announce and run the first seeded-defect probe, per Section 29.5.
32.5 What Ninety Days Does Not Buy #
Stage 2 entry requires Level 2 control maturity across all fifteen dimensions in Section 18. Ninety days at this pace typically produces Level 2 on identity, the merge boundary, loop instrumentation, and one work class, with the platform dimensions trailing.
That is the correct outcome, and it should be reported as such. The failure mode is declaring Stage 2 because the visible work is done. The context substrate in Section 28, the evaluation infrastructure in Section 8.5, and provenance verification in Section 15 are quarter-two work, and Stage 3 requires two quarters of operating evidence on top of that.
Anyone promising an AI-native engineering organization in a quarter is selling the tooling, not the operating model.
32.6 The Sequencing Trap #
The temptation is to begin with policy, because policy is fast to write and visible to leadership. Policy written before the ground truth of weeks 1 and 2 governs a population that does not exist and misses the one that does.
The second temptation is to begin with a tool selection. Section 21 exists for that decision, and Section 8 is the reason it is not the first one: the published outcome differences between organizations are platform differences rather than tool differences.
32.7 What This Rests On #
The hardened-boundary result is measured.113 The sequence, the artifact set, and the ninety-day scope estimate are this framework’s construction, derived from the entry criteria in Section 4.2 rather than from any published implementation timeline. No organization has published a staged adoption timeline with outcome data against it.
References cited in this section
1 of 243 · numbering matches the PDF
- 113Thaman, "Reward Hacking Benchmark: Measuring Exploits in LLM Agents with Tool Use," arXiv:2605.02964, May 3, 2026. Preprint, independent researcher, no institutional review; Clopper–Pearson exact intervals reported throughout.arxiv.org/abs/2605.02964 ↗