The AI SDLC / §44
Section 44 of 44 32 min read

References

243 sources. Where a source is vendor-published, correlational, or an unreviewed preprint, the entry says so.

  1. 1International Organization for Standardization and International Electrotechnical Commission, Information Security, Cybersecurity and Privacy Protection — Information Security Controls, ISO/IEC 27002:2022, 3rd ed. (Geneva: ISO, February 2022). Controls 8.25–8.34 govern secure development; 8.30 governs outsourced development.
  2. 2PCI Security Standards Council, Payment Card Industry Data Security Standard: Requirements and Testing Procedures, v4.0.1 (Wakefield, MA: PCI SSC, June 2024). Requirement 6.2.3 governs pre-release review of bespoke and custom code; 6.2.3.1 governs manual review, requiring a reviewer other than the originating code author and management approval.
  3. 3American Institute of Certified Public Accountants, TSP Section 100, 2017 Trust Services Criteria for Security, Availability, Processing Integrity, Confidentiality, and Privacy (With Revised Points of Focus — 2022) (New York: AICPA, 2022). Criterion CC8.1 governs change management.
  4. 4European Commission, Commission Delegated Regulation (EU) 2024/1774 of March 13, 2024 supplementing Regulation (EU) 2022/2554 with regard to regulatory technical standards specifying ICT risk management tools, methods, processes and policies, OJ L, 2024. Articles 15–17 govern ICT project management, systems acquisition and development, and change management.
  5. 5International Organization for Standardization and International Electrotechnical Commission, Information Technology — Artificial Intelligence — Management System, ISO/IEC 42001:2023, 1st ed. (Geneva: ISO, December 2023). Annex A.6 covers the AI system life cycle.
  6. 6International Organization for Standardization and International Electrotechnical Commission, Information Technology — Artificial Intelligence — AI System Life Cycle Processes, ISO/IEC 5338:2023, 1st ed. (Geneva: ISO, December 20, 2023).
  7. 7International Organization for Standardization and International Electrotechnical Commission, Information Technology — Artificial Intelligence — Guidance on Risk Management, ISO/IEC 23894:2023, 1st ed. (Geneva: ISO, February 6, 2023).
  8. 8National Institute of Standards and Technology, Artificial Intelligence Risk Management Framework (AI RMF 1.0), NIST AI 100-1 (Gaithersburg, MD: NIST, January 2023).
  9. 9European Parliament and Council, Regulation (EU) 2024/1689 of June 13, 2024 laying down harmonised rules on artificial intelligence (Artificial Intelligence Act), OJ L, July 12, 2024.
  10. 10Harold Booth et al., Secure Software Development Practices for Generative AI and Dual-Use Foundation Models: An SSDF Community Profile, NIST Special Publication 800-218A (Gaithersburg, MD: NIST, July 2024).
  11. 11National Institute of Standards and Technology, “AI Agent Standards Initiative,” Center for AI Standards and Innovation, announced February 17, 2026.
  12. 12National Cybersecurity Center of Excellence, Accelerating the Adoption of Software and AI Agent Identity and Authorization, concept paper (Rockville, MD: NIST NCCoE, February 2026).
  13. 13National Institute of Standards and Technology, NIST SP 800-53 Control Overlays for Securing AI Systems: Concept Paper (Gaithersburg, MD: NIST, August 14, 2025).
  14. 14Beatrice Nolan, “An AI-Powered Coding Tool Wiped Out a Software Company’s Database, Then Apologized for a ‘Catastrophic Failure on My Part,’” Fortune, July 23, 2025.
  15. 15Veracode, “Spring 2026 GenAI Code Security Update: Despite Claims, AI Models Are Still Failing Security,” March 24, 2026.
  16. 16Microsoft, “Visual Studio Code Release Notes, Version 1.110,” March 2026. The git.addAICoAuthor setting offers off, chatAndAgent, and all, and ships defaulting to off.
  17. 17Addy Osmani, “Comprehension Debt — The Hidden Cost of AI Generated Code,” March 14, 2026. Practitioner essay, not research.
  18. 18Anthropic, “How AI Assistance Impacts the Formation of Coding Skills,” January 29, 2026. Randomized controlled trial, n=52; vendor-affiliated research.
  19. 19CVE-2025-54135 (“CurXecute”), National Vulnerability Database, published August 4, 2025; research disclosure by Aim Security, August 1, 2025. CNA base score 8.5; NVD scores it 9.8. See also Tenable Research, “FAQ: CVE-2025-54135 and CVE-2025-54136, Vulnerabilities in Cursor,” August 2025.
  20. 20CVE-2025-54136 (“MCPoison”), National Vulnerability Database, published August 1, 2025; research disclosure by Check Point Research, August 5, 2025. CNA base score 7.2; NVD scores it 8.8.
  21. 21Kudelski Security, “How We Exploited CodeRabbit: From a Simple PR to RCE and Write Access on 1M Repositories,” August 19, 2025. Disclosed to vendor January 24, 2025; fix deployed January 30, 2025.
  22. 22Pillar Security, “New Vulnerability in GitHub Copilot and Cursor: How Hackers Can Weaponize Code Agents,” March 18, 2025.
  23. 23Google Cloud and DORA, 2025 State of AI-Assisted Software Development Report, September 23, 2025. Survey, approximately 5,000 respondents, fielded 13 June–July 21, 2025.
  24. 24Stack Overflow, 2025 Developer Survey — AI Section, 2025. n=48,962. The 2026 survey opened June 23, 2026; results were not published as of August 2026.
  25. 25Google Cloud and DORA, 2024 Accelerate State of DevOps Report, October 2024. Figures are regression-estimated effects per 25 percent increase in a self-reported adoption index, not measured deltas.
  26. 26GitHub and Microsoft Office of the Chief Economist, “Quantifying GitHub Copilot’s Impact on Developer Productivity and Happiness,” September 7, 2022 (updated May 21, 2024). n=95, single greenfield task, 95 percent CI [21%, 89%]; vendor-conducted.
  27. 27Kevin Demirer, Sida Peng, et al., “The Effects of Generative AI on High-Skilled Work: Evidence from Three Field Experiments with Software Developers,” Management Science. n=4,867.
  28. 28Elise Paradis et al., “How Much Does AI Impact Development Speed? An Enterprise-Based Randomized Controlled Trial,” arXiv:2410.12944, October 16, 2024. n=96.
  29. 29DX, “AI Coding Assistant Pricing and Impact,” 2026. Telemetry across 400+ organizations over 14 months; vendor-published.
  30. 30Faros AI, The AI Engineering Report 2026: The Acceleration Whiplash, 2026. Telemetry, approximately 22,000 developers and 4,000 teams, within-organization design; vendor-published. Figures for review duration vary across the vendor’s own publications and are not presented here as a series.
  31. 31Sien Reeve O. Peralta, Fumika Hoshi, Hironori Washizaki, Naoyasu Ubayashi, Inase Kondo, Yoshiki Higo, Hiroki Mukai, Norihiro Yoshida, Kazuki Kusama, Hidetake Tanaka, and Youmei Fan, “Why Are Agentic Pull Requests Merged or Rejected? An Empirical Study,” 23rd International Conference on Mining Software Repositories (MSR ‘26), arXiv:2605.22534. 11,048 closed agentic pull requests, 9,799 human-reviewed, 717 manually inspected.
  32. 32George Xu, Arjun Subramanian, and Nithilan Karthik, “AI Agent Pull Requests on GitHub: Frequency, Structure, and Merge Conflict Rates,” arXiv:2607.04697, July 6, 2026. 33,596 agent-authored pull requests across 2,807 repositories. Conflict rates rest on 601 intra-agent and 115 cross-agent evaluable pairs; the authors describe the figures as a conservative lower bound measuring textual conflicts only. Preprint.
  33. 33Joel Becker, Nate Rush, Elizabeth Barnes, and David Rein, “Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity,” METR, arXiv:2507.09089, July 10, 2025. n=16, 246 tasks.
  34. 34METR, “We Are Changing Our Developer Productivity Experiment Design,” February 24, 2026. n=57, 143 repositories, 800+ tasks.
  35. 35Veracode, 2025 GenAI Code Security Report: Assessing the Security of Using LLMs for Coding, August 2025.
  36. 36Thomas Claburn, “AI Code Assistants Improve Production of Security Problems,” The Register, September 5, 2025, reporting Apiiro research across tens of thousands of repositories at Fortune 50 enterprises. Vendor research, correlational; velocity measurement questioned in the reporting.
  37. 37WebAIM, The WebAIM Million: The 2026 Report on the Accessibility of the Top 1,000,000 Home Pages, February 2026. Correlational; WebAIM attributes the trend to third-party frameworks and AI-assisted coding as a likely cause.
  38. 38GitClear, The Maintainability Gap: AI Code Quality in 2026, January 2026. 623 million analyzed changes, 2023–2026. Correlational; commits are not labeled by AI authorship; vendor-published.
  39. 39Liu, Widyasari, Zhao, Irsan, and Lo, “Debt Behind the AI Boom: A Large-Scale Empirical Study of AI-Generated Code in the Wild,” arXiv:2603.28592, March 2026. 304,362 verified AI-authored commits. No human-written control group; rates are absolute, not comparative. Preprint.
  40. 40DX, “AI Coding Assistant Pricing,” 2026.
  41. 41Gartner, “Enterprise AI Coding Agent Market,” April 2026.
  42. 42FinOps Foundation, The State of FinOps 2026. 1,192 respondents representing more than $83 billion in annual cloud spend.
  43. 43Linux Kernel Documentation, “Coding Assistants,”.
  44. 44All Things Open, “Assisted-by: How Open Source Projects Are Drawing the Line on AI Contributions,”. Secondary source; individual project policies are independently verifiable.
  45. 45Epoch AI, “LLM Inference Prices Have Fallen Rapidly but Unequally Across Tasks,” March 12, 2025. Decline rates range from 9x to 900x per year depending on task; analysis predates this document by well over a year.
  46. 46European Parliament and Council, Regulation (EU) 2024/2847 of October 23, 2024 on horizontal cybersecurity requirements for products with digital elements (Cyber Resilience Act), OJ L, November 20, 2024. See also European Commission, “Cyber Resilience Act Reporting Obligations,”. The final report is due within fourteen days for an actively exploited vulnerability and within one month for a severe incident.
  47. 47OpenAI, “Deprecations,” OpenAI API documentation.
  48. 48Anthropic, “Model Deprecations,” Claude platform documentation.
  49. 49Gaia Colombo, Leonardo Mariani, Daniela Micucci, and Oliviero Riganelli, “On the Possibility of Breaking Copyleft Licenses When Reusing Code Generated by ChatGPT,” arXiv:2502.05023, 2025. Preprint.
  50. 50Albert Ziegler, “GitHub Copilot Research Recitation,” The GitHub Blog, June 30, 2021 (updated August 16, 2022). 2021 data, Python only, original Copilot model; the only rigorous first-party measurement located.
  51. 51Bloomberg Law, “Copyright Suit Over GitHub AI Coding Tool Vexes Ninth Circuit,” February 11, 2026. Doe v. GitHub, Ninth Circuit no. 24-7700, argued February 11, 2026. No opinion had issued as of August 28, 2026 on the best available docket tracking; verify the docket directly before relying on this position.
  52. 52GitHub, “Introducing Agent HQ: Any Agent, Any Way You Work,” The GitHub Blog, October 28, 2025. Announced and preview capabilities; feature availability is not evidence of adopted practice.
  53. 53Model Context Protocol, “Authorization” and “Security Best Practices,” specification version 2026-07-28.
  54. 54Anthropic, “Donating the Model Context Protocol and Establishing the Agentic AI Foundation,” December 9, 2025; Linux Foundation, “Linux Foundation Announces the Formation of the Agentic AI Foundation,” December 2025.
  55. 55OWASP GenAI Security Project, OWASP Top 10 for Agentic Applications for 2026, December 9, 2025.
  56. 56OWASP GenAI Security Project, OWASP GenAI LLM Top 10 2026, v1.0, August 3, 2026. First edition to weight documented incident data alongside expert consensus.
  57. 57MITRE, “MITRE ATLAS,”. Case study AML.CS0041, “Rules File Backdoor: Supply Chain Attack on AI Coding Assistants.” See also Center for Threat-Informed Defense, “Secure AI v2 Release,” April 2026.
  58. 58Check Point Research, “RCE and API Token Exfiltration Through Claude Code Project Files (CVE-2025-59536),” 2026. Includes CVE-2026-21852 and a project-hooks consent bypass; all fixed in vendor releases.
  59. 59Nx Team, “S1ngularity — What Happened, How We Responded, What We Learned,” Nx Blog, September 2025 (incident August 26, 2025).
  60. 60Socket, “Nx npm Packages Compromised in Supply Chain Attack Weaponizing AI CLI Tools,” August 27, 2025.
  61. 61GitGuardian, “The Nx ‘s1ngularity’ Attack: Inside the Credential Leak,” August 27, 2025.
  62. 62Joseph Spracklen, Raveen Wijewickrama, A. H. M. Nazmus Sakib, Anindya Maiti, Bimal Viswanath, and Murtuza Jadliwala, “We Have a Package for You! A Comprehensive Analysis of Package Hallucinations by Code Generating LLMs,” 34th USENIX Security Symposium, August 2025.
  63. 63Aleksandr Churilov, “The Range Shrinks, the Threat Remains: Re-evaluating LLM Package Hallucinations on the 2026 Frontier-Model Cohort,” arXiv:2605.17062, revised August 9, 2026. Independent preprint, not peer-reviewed; directional only.
  64. 64Chengjie Wang, Jingzheng Wu, Xiang Ling, Tianyue Luo, and Chen Zhao, “Correct Code, Vulnerable Dependencies: A Large Scale Measurement Study of LLM-Specified Library Versions,” arXiv:2605.06279, May 7, 2026. Preprint.
  65. 65Unit 42, Palo Alto Networks, “‘Shai-Hulud’ Worm Compromises npm Ecosystem in Supply Chain Attack,” updated November 26, 2025.
  66. 66Datadog Security Labs, “The Shai-Hulud 2.0 npm Worm: Analysis, and What You Need to Know,” November 2025.
  67. 67GitHub, “Our Plan for a More Secure npm Supply Chain,” The GitHub Blog, September 22, 2025.
  68. 68GitHub, “npm Classic Tokens Revoked, Session-Based Auth and CLI Token Management Now Available,” GitHub Changelog, December 9, 2025.
  69. 69OWASP Foundation, “OWASP Non-Human Identities Top 10,” 2025.
  70. 70GitHub, “Track Copilot Sessions,” GitHub Docs.
  71. 71GitHub, “Risks and Mitigations for GitHub Copilot Cloud Agent,” GitHub Docs. Cited here as a documented control set rather than a product recommendation.
  72. 72SLSA Community, SLSA Specification v1.2, November 24, 2025. The Source Track was added in v1.2; Source Level 4 requires two trusted persons to review all changes to protected branches.
  73. 73Linux Foundation, “Linux Foundation Announces $12.5 Million in Grant Funding from Leading Organizations to Advance Open Source Security,” March 17, 2026. Funds managed through Alpha-Omega and the Open Source Security Foundation; contributors named as Anthropic, AWS, GitHub, Google, Google DeepMind, Microsoft, and OpenAI.
  74. 74Evan Miller, “Adding Error Bars to Evals: A Statistical Approach to Language Model Evaluations,” arXiv:2411.00640, November 1, 2024.
  75. 75Vals AI, “SWE-bench Verified Independent Evaluation,” updated August 19, 2026. Seven of 86 evaluated models at or above 95 percent; the benchmark is saturating.
  76. 76Confident AI, DeepEval Documentation; Ragas, Metrics Documentation. Cited jointly to establish that identically named metrics compute differently across frameworks and that scores are not portable.
  77. 77Naman Jain et al., “LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code,” arXiv:2403.07974, 2024.
  78. 78Justin D. Norman, Michael U. Rivera, and D. Alex Hughes, “Reliability Without Validity: A Systematic, Large-Scale Evaluation of LLM-as-a-Judge Models Across Agreement, Consistency, and Bias,” arXiv:2606.19544, June 17, 2026. Twenty-one judges, approximately 541,000 judgments. Preprint.
  79. 79Blake Bullwinkel, Amanda Minnich, Shiven Chawla, et al., “Lessons From Red Teaming 100 Generative AI Products,” Microsoft AI Red Team, arXiv:2501.07238, January 13, 2025; see also PyRIT, and the attack-success-rate scorecard model described in Microsoft Foundry documentation.
  80. 80CycloneDX, “CycloneDX v1.7 Released,” October 21, 2025; Ecma International, ECMA-424: CycloneDX Bill of Materials Specification, 2nd ed., December 2025.
  81. 81SPDX, “AI Profile” and “Dataset Profile,” SPDX Specification 3.0.1. SPDX 3.1 Release Candidate 1 was published January 26, 2026 and is not final.
  82. 82Cybersecurity and Infrastructure Security Agency et al., 2026 Minimum Elements for a Software Bill of Materials (SBOM), July 29, 2026.
  83. 83Cybersecurity and Infrastructure Security Agency and G7 partners, Software Bill of Materials for AI — Minimum Elements, 2026. Voluntary and nonmandatory. Sources conflict on the exact publication date between May and June 2026; verify before citing a date.
  84. 84Zach Steindler, “Cosign v3 Is Now Available,” Sigstore Blog, October 8, 2025; Hayden Blauzvern, “Rekor v2 GA,” Sigstore Blog, October 10, 2025.
  85. 85OpenSSF, “An Introduction to the OpenSSF Model Signing (OMS) Specification,” June 25, 2025; reference implementation.
  86. 86LaunchDarkly, “AI Configs,” product documentation. Cited as an example of the capability class; vendor documentation, no independent efficacy data.
  87. 87Lingjiao Chen, Matei Zaharia, and James Zou, “How Is ChatGPT’s Behavior Changing Over Time?” arXiv:2307.09009, July 18, 2023. The specific prime-identification task drew methodological criticism; the general finding of behavioral shift under a stable API surface has not been refuted.
  88. 88D. Sculley et al., “Hidden Technical Debt in Machine Learning Systems,” Advances in Neural Information Processing Systems (NIPS 2015).
  89. 89Microsoft, Agent Governance Toolkit, agent-sre package, public preview. Well-specified proposal; no published production efficacy data.
  90. 90OpenTelemetry, GenAI Semantic Conventions. All GenAI spans, events, metrics, and gen_ai.* attributes carry Development stability status; no tagged release as of August 2026.
  91. 91FinOps Foundation, “FinOps for AI,” FinOps Framework technology category. AI usage is currently expressed through existing FOCUS columns rather than AI-native ones.
  92. 92in-toto, “in-toto Attestation Framework Specification v1.2,” March 18, 2024. CNCF graduated project.
  93. 93GUAC, “Graph for Understanding Artifact Composition,” v1.1.0, March 13, 2026. OpenSSF incubating project.
  94. 94European Parliament and Council, Regulation (EU) 2026/1744 of July 8, 2026 amending Regulations (EU) 2024/1689, (EU) 2018/1139 and (EU) 2023/1230 as regards the simplification of the implementation of harmonised rules on artificial intelligence (Digital Omnibus on AI), OJ L, 2026. Entered into force July 27, 2026.
  95. 95Office of Management and Budget, Adopting a Risk-Based Approach to Software and Hardware Security, OMB Memorandum M-26-05 (Washington, DC: Executive Office of the President, January 23, 2026). Rescinds M-22-18 and M-23-16. Note that the CISA attestation form page had not been updated to reflect the rescission as of August 2026.
  96. 96U.S. Food and Drug Administration, Marketing Submission Recommendations for a Predetermined Change Control Plan for Artificial Intelligence-Enabled Device Software Functions, guidance for industry and FDA staff (Silver Spring, MD: FDA, December 2024). Final. A companion lifecycle management guidance issued January 6, 2025 remains in draft.
  97. 97Murugiah Souppaya, Karen Scarfone, and Donna Dodson, Secure Software Development Framework (SSDF) Version 1.1: Recommendations for Mitigating the Risk of Software Vulnerabilities, NIST Special Publication 800-218 (Gaithersburg, MD: NIST, February 2022). Version 1.1 remains the operative final version; a draft revision (SP 800-218r1, SSDF 1.2) was released for comment in December 2025 and had not been finalized as of August 2026.
  98. 98Google LLC, DORA AI Capabilities Model, v2025.1, 2025. Seven capabilities: clear and communicated AI stance; healthy data ecosystems; AI-accessible internal data; strong version control practices; working in small batches; user-centric focus; quality internal platforms.
  99. 99Avishay Balter et al., “Security-Focused Guide for AI Code Assistant Instructions,” OpenSSF Best Practices and AI/ML Working Groups, August 1, 2025.
  100. 100Max Charas and Marc Bruggmann, “1,500+ PRs Later: Spotify’s Journey with Our Background Coding Agent,” Spotify Engineering, November 2025. First-party self-reported adoption figures; quality outcomes not formally quantified.
  101. 101Niklas Gustavsson, “Coding Is No Longer the Constraint,” Spotify Engineering, June 3, 2026. First-party.
  102. 102Gergely Orosz, “How Uber Uses AI for Development,” The Pragmatic Engineer, March 10, 2026. Based on a talk by Uber engineers; figures self-reported.
  103. 103Zohar Einy, “How Uber Built a Software Factory,” Port newsletter, August 24, 2026; and Cameron McClellan, “How Uber Built the Enterprise AI Security Playbook,” Speakeasy, May 28, 2026. Third-party summaries of conference talks, not first-party engineering posts; attribute to the talks.
  104. 104DX, Q2 2026 State of AI Impact in Engineering Report, July 22, 2026. 500+ organizations, telemetry plus survey. Vendor-published.
  105. 105Stephen Toub, “Ten Months with Copilot Coding Agent in dotnet/runtime,” .NET Blog, March 23, 2026. Vendor-adjacent — Microsoft reporting on a Microsoft product — but fully denominated and unusually candid.
  106. 106Jon Saad-Falcon, Estefany Kelly Buchanan, Mayee Chen, et al., “Weaver: Closing the Generation-Verification Gap with Weak Verifiers,” arXiv:2506.18203, June 18, 2025. Preprint.
  107. 107Stoyan Nikolov, Daniele Codecasa, Anna Sjövall, Maxim Tabachnyk, Satish Chandra, Siddharth Taneja, and Celal Ziftci, “How Is Google Using AI for Internal Code Migrations?,” arXiv:2501.06972, January 12, 2025. Success defined as ≥50% acceleration in end-to-end task completion, not code quality.
  108. 108Uber, “uReview: Scalable, Trustworthy GenAI for Code Review at Uber,” Uber Blog. First-party, unaudited, with no independent evaluation. The ~1,500 hours saved weekly is modeled from an assumed ten minutes of second-reviewer time per commit, not measured. The 65% and 51% figures are produced by different methods and are not a matched comparison; see Section 39.2.
  109. 109Alexander Frömmgen and Lera Kharatyan, “Resolving Code Review Comments with ML,” Google Research Blog, May 23, 2023 (and ICSE-SEIP 2024, DOI 10.1145/3639477.3639746); Vijayvergiya et al., “AI-Assisted Assessment of Coding Practices in Modern Code Review,” AIware ‘24, arXiv:2405.13565.
  110. 110Jie Huang, Xinyun Chen, Swaroop Mishra, Huaixiu Steven Zheng, Adams Wei Yu, Xinying Song, and Denny Zhou, “Large Language Models Cannot Self-Correct Reasoning Yet,” ICLR 2024, arXiv:2310.01798.
  111. 111Kristen Pereira, Neelabh Sinha, Rajat Ghosh, and Debojyoti Dutta, “CR-Bench: Evaluating the Real-World Utility of AI Code Review Agents,” arXiv:2603.11078, March 10, 2026. Preprint, vendor-affiliated.
  112. 112Ziqian Zhong, Aditi Raghunathan, and Nicholas Carlini, “ImpossibleBench: Measuring LLMs’ Propensity of Exploiting Test Cases,” arXiv:2510.20270, October 23, 2025. Preprint.
  113. 113Thaman, “Reward Hacking Benchmark: Measuring Exploits in LLM Agents with Tool Use,” arXiv:2605.02964, May 3, 2026. Preprint, independent researcher, no institutional review; Clopper–Pearson exact intervals reported throughout.
  114. 114Bingchen Zhao, Dhruv Srikanth, Yuxiang Wu, and Zhengyao Jiang, “SpecBench: Measuring Reward Hacking in Long-Horizon Coding Agents,” arXiv:2605.21384, May 20, 2026. Preprint, vendor-affiliated.
  115. 115Minh Vu Thai Pham, Tue Le, Dung Nguyen Manh, Huy Nhat Phan, and Nghi D. Q. Bui, “SWE-EVO: Benchmarking Coding Agents in Long-Horizon Software Evolution Scenarios,” arXiv:2512.18470, revised April 4, 2026. Preprint. See also Shaoqiu Zhang et al., “SWE-Explore: Benchmarking How Coding Agents Explore Repositories,” arXiv:2606.07297, June 5, 2026.
  116. 116Cloud Security Alliance AI Safety Initiative, “The Non-Human Identity Governance Vacuum: AI Agents and the Fastest-Growing Unmanaged Attack Surface,” May 20, 2026. Percentages aggregated from third-party industry reports rather than a single CSA-fielded survey.
  117. 117Xiangxin Zhao, Han Li, Shuaiting Li, Tianyi Zhao, Earl T. Barr, Federica Sarro, and He Ye, “Failure as a Process: An Anatomy of CLI Coding Agent Trajectories,” arXiv:2607.09510, July 10, 2026. 1,794 trajectories, >63,000 steps, seven models, three scaffolds. Preprint.
  118. 118Stack Overflow, “Mind the Gap: Closing the AI Trust Gap for Developers,” February 18, 2026.
  119. 119Faros AI, AI Productivity Paradox research report, March 2026, reported in “More Code, More Bugs,” ADTmag, April 22, 2026. 22,000 developers, 4,000+ teams, two years of telemetry. Vendor-published; figures vary across the vendor’s own publications and are not a coherent series.
  120. 120Google Cloud DORA, ROI of AI-Assisted Software Development (2026.01), May 11, 2026. Modeled ROI figures are scenario outputs, not measurements.
  121. 121Hiroki Watanabe, Hao Li, Yutaro Kashiwa, Reid, Iida, and Ahmed E. Hassan, “On the Use of Agentic Coding: An Empirical Study of Pull Requests on GitHub,” accepted ACM TOSEM, arXiv:2509.14745v3. 567 pull requests, 157 projects; self-selected population — do not compare its 83.8% directly to enterprise merge rates.
  122. 122Hamel Husain, Isaac Flath, and Johno Whitaker, “Thoughts On A Month With Devin,” Answer.AI, January 8, 2025. 20 tasks; small-n practitioner evaluation.
  123. 123M. Rastenis, B. Chou, S. Roy Choudhary, and R. Just, “Automated Software Test Generation at Industry Scale Using a Multi-Agent Architecture and Workflow Integration” (AutoCover), ICSE-SEIP ‘26, DOI 10.1145/3786583.3786918.
  124. 124OpenAI, “Why We No Longer Evaluate SWE-bench Verified,” February 23, 2026. Vendor-published.
  125. 125Naman Jain, “Reward Hacking Is Swamping Model Intelligence Gains,” Cursor Blog, June 25, 2026. Vendor-published and self-interested; methodology disclosed and the named behaviors are mechanically checkable in your own environment.
  126. 126Anthropic, 2026 Agentic Coding Trends Report. Predictive trends document with no stated sample size or methodology; the 0–20% full-delegation figure is developer self-report.
  127. 127John Yang, Carlos E. Jimenez, Alexander Wettig, Kilian Lieret, Shunyu Yao, Karthik Narasimhan, and Ofir Press, “SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering,” NeurIPS 2024, arXiv:2405.15793.
  128. 128SWE-agent project, “mini-swe-agent,” GitHub repository, accessed August 28, 2026. Benchmark score is a project self-report, not independently reproduced.
  129. 129Xingyao Wang et al., “OpenHands: An Open Platform for AI Software Developers as Generalist Agents,” ICLR 2025, arXiv:2407.16741.
  130. 130Anthropic, “How the Agent Loop Works,” Claude Agent SDK documentation, accessed August 28, 2026. Vendor documentation; cited as product fact.
  131. 131Thibaud Gloaguen, Niels Mündler, Mark Müller, Veselin Raychev, and Martin Vechev, “Evaluating AGENTS.md: Are Repository-Level Context Files Helpful for Coding Agents?,” arXiv:2602.11988, February 12, 2026, revised June 23, 2026. Preprint.
  132. 132Jai Lal Lulla, Seyedmoein Mohsenimofidi, Matthias Galster, Jie M. Zhang, Sebastian Baltes, and Christoph Treude, “On the Impact of AGENTS.md Files on the Efficiency of AI Coding Agents,” ICSE JAWs 2026. Task success rate explicitly not evaluated.
  133. 133Ali Arabat and Mohammed Sayagh, “Toward Instructions-as-Code: Understanding the Impact of Instruction Files on Agentic Pull Requests,” arXiv:2606.13449. 15,549 agentic pull requests across 148 projects. Preprint.
  134. 134Worawalan Chatlatanagulchai, Hao Li, Yutaro Kashiwa, et al., “Agent READMEs: An Empirical Study of Context Files for Agentic Coding,” arXiv:2511.12884, November 17, 2025. 2,303 files from 1,925 repositories. Preprint.
  135. 135“AGENTS.md,” open format, stewarded by the Agentic AI Foundation under the Linux Foundation.
  136. 136Anthropic, “Effective Context Engineering for AI Agents,” September 29, 2025. Vendor position; no efficacy figures published for the three named long-horizon techniques.
  137. 137Stefan Heule, Emily Jia, and Naman Jain, “Improving Agent with Semantic Search,” Cursor Blog, November 6, 2025. Vendor-published, opposite architectural position to reference 136; each publishes evidence favoring its own approach.
  138. 138Kelly Hong, Anton Troynikov, and Jeff Huber, “Context Rot: How Increasing Input Tokens Impacts LLM Performance,” Chroma Research, July 14, 2025. Publisher sells vector databases and has a commercial interest in the conclusion.
  139. 139Anthropic, “Subagents in the SDK,” Claude Agent SDK documentation, accessed August 28, 2026. Mechanism documented; task-success benefit asserted rather than measured.
  140. 140Chunqiu Steven Xia, Yinlin Deng, Soren Dunn, and Lingming Zhang, “Agentless: Demystifying LLM-based Software Engineering Agents,” Proceedings of the ACM on Software Engineering (FSE 2025), arXiv:2407.01489, DOI 10.1145/3715754.
  141. 141Ryan Ehrlich, Bradley Brown, Jordan Juravsky, Ronald Clark, Christopher Ré, and Azalia Mirhoseini, “CodeMonkeys: Scaling Test-Time Compute for Software Engineering,” arXiv:2501.14723. Preprint.
  142. 142Erik Schluntz and Barry Zhang, “Building Effective AI Agents,” Anthropic, December 19, 2024. Vendor-published.
  143. 143Anthropic, “How We Built Our Multi-Agent Research System,” June 13, 2025. Vendor-published; the same post reports ~15× token usage and states that most coding tasks involve fewer truly parallelizable subtasks than research.
  144. 144Walden Yan, “Don’t Build Multi-Agents,” Cognition, June 12, 2025. Vendor-published; the author has since publicly softened the position.
  145. 145Harrison Chase, “How and When to Build Multi-Agent Systems,” LangChain, June 16, 2025. Vendor-published.
  146. 146Mert Cemri, Melissa Z. Pan, Shuyi Yang, et al., “Why Do Multi-Agent LLM Systems Fail?,” NeurIPS 2025, arXiv:2503.13657. 1,600+ annotated traces, seven frameworks, Cohen’s κ = 0.88. The one large peer-reviewed non-vendor result in this debate.
  147. 147Dat Tran and Douwe Kiela, “Single-Agent LLMs Outperform Multi-Agent Systems on Multi-Hop Reasoning Under Equal Thinking Token Budgets,” arXiv:2604.02460, April 2, 2026. Multi-hop question answering, not software engineering. Preprint.
  148. 148Jiawei Xu, Arief Koesdwiady, Sisong Bei, et al., “Rethinking the Value of Multi-Agent Workflow: A Strong Single Agent Baseline,” arXiv:2601.12307, January 18, 2026. Preprint.
  149. 149Thomas Kwa, Ben West, Joel Becker, et al., “Measuring AI Ability to Complete Long Software Tasks,” METR, arXiv:2503.14499, March 19, 2025; and METR, “Task-Completion Time Horizons of Frontier AI Models,” updated May 8, 2026. Measures task difficulty in human time, not autonomous run duration; measurements above 16 hours are unreliable with the current task suite.
  150. 150GitHub, “About GitHub Copilot Cloud Agent,” GitHub Docs, accessed August 28, 2026. Vendor documentation; cited as product fact.
  151. 151OpenAI, “Run Long Horizon Tasks with Codex,” OpenAI Developers Blog, accessed August 28, 2026. A single documented run, not a measurement.
  152. 152Akshit Sinha, Arvindh Arun, Shashwat Goel, Steffen Staab, and Jonas Geiping, “The Illusion of Diminishing Returns: Measuring Long Horizon Execution in LLMs,” ICLR 2026, arXiv:2509.09677.
  153. 153Philippe Laban, Hiroaki Hayashi, Yingbo Zhou, and Jennifer Neville, “LLMs Get Lost In Multi-Turn Conversation,” arXiv:2505.06120, May 9, 2025. Over 200,000 simulated conversations; tests conversational generation rather than agentic coding trajectories — the transfer is plausible and unproven.
  154. 154Longju Bai, Zhemin Huang, Xingyao Wang, Jiao Sun, Rada Mihalcea, Erik Brynjolfsson, Alex Pentland, and Jiaxin Pei, “How Do AI Agents Spend Your Money? Analyzing and Predicting Token Consumption in Agentic Coding Tasks,” arXiv:2604.22750. Preprint.
  155. 155Sayash Kapoor, Benedikt Stroebl, Zachary S. Siegel, Nitya Nadgir, and Arvind Narayanan, “AI Agents That Matter,” arXiv:2407.01502, July 1, 2024. Preprint.
  156. 156GitHub, “spec-kit: Toolkit to Help You Get Started with Spec-Driven Development,” GitHub repository, MIT licensed. Presents no empirical evidence, benchmark, or user study; its own goals are labeled experimental.
  157. 157Noble Saji Mathews and Meiyappan Nagappan, “Test-Driven Development and LLM-based Code Generation,” ASE 2024, arXiv:2402.13521.
  158. 158Oorja Majgaonkar, Zhiwei Fei, Xiang Li, Federica Sarro, and He Ye, “Understanding Code Agent Behaviour: An Empirical Study of Success and Failure Trajectories,” arXiv:2511.00197, October 31, 2025. Preprint.
  159. 159Vantage, “The Hidden Cost Driver in Agentic Coding Sessions,” April 15, 2026. Vendor modeling, not measurement.
  160. 160HAL leaderboard, Princeton. Late-2024 model-scaffold pairs; cost and accuracy are not correlated across entries.
  161. 161Chengpeng Li, Farnaz Behrang, August Shi, and Peng Liu, “FlakyGuard: Automatically Fixing Flaky Tests at Industry Scale,” ASE 2025, arXiv:2511.14002. Peer-reviewed; deployed autonomously at Uber over six months. The 17.7% end-to-end figure is derived from the paper’s three reported conditional rates, not stated by the authors; plan against it rather than against the 51.8% acceptance rate.
  162. 162John Micco, “Flaky Tests at Google and How We Mitigate Them,” Google Testing Blog, May 27, 2016.
  163. 163Snyk, “Benchmarking Secure-and-Functional Remediation,” August 18, 2026. Vendor-published; ~150 samples, three languages, single-file scope, two runs — the vendor states the sample is too small for significance on narrow gaps.
  164. 164Benjamin Steenhoek, Siva Sivaraman, Renata Saldivar Gonzalez, Yevhen Mohylevskyy, Roshanak Zilouchian Moghaddam, and Wei Le, “Closing the Gap: A User Study on the Real-World Usefulness of AI-powered Vulnerability Detection & Repair in the IDE,” arXiv:2412.14306. 17 professional developers, 24 projects, >1.7M lines; conducted with Microsoft.
  165. 165Charles Covey-Brandt, “Accelerating Large-Scale Test Migration with LLMs,” Airbnb Tech Blog, March 13, 2025. First-party; no dollar cost disclosed, and long-tail files required 50–100 retry attempts.
  166. 166Celal Ziftci, Stoyan Nikolov, Anna Sjövall, Bo Kim, Daniele Codecasa, and Max Kim, “Migrating Code At Scale With LLMs At Google,” arXiv:2504.09691, DOI 10.1145/3696630.3728542. Twelve-month study; three developers, 39 migrations, 595 changes, 93,574 edits.
  167. 167Hyrum Wright, “Large-Scale Changes,” chapter 22 in Software Engineering at Google (Sebastopol, CA: O’Reilly, 2020). Predates AI by a decade; its shard-and-land architecture has not been re-validated for agentic work.
  168. 168Uber, “Uber’s Strategy to Upgrading 2M+ Spark Jobs,” Uber Blog, September 25, 2025. First-party. Deterministic AST transformation; the account contains no language-model use anywhere, though it does not state a rejection of the option. The “85% job migration” figure is a section heading rather than a measured automation rate.
  169. 169Nadia Alshahwan, Jubin Chheda, Anastasia Finogenova, Beliz Gokkaya, Mark Harman, Inna Harper, Alexandru Marginean, Shubho Sengupta, and Eddy Wang, “Automated Unit Test Improvement Using Large Language Models at Meta” (TestGen-LLM), FSE Companion 2024, arXiv:2402.09171.
  170. 170Christopher Foster, Rahul Gulati, Mark Harman, Inna Harper, Ke Mao, Jillian Ritchey, Hervé Robert, and Shubho Sengupta, “Mutation-Guided LLM-based Test Generation at Meta” (ACH), FSE Companion ‘25, arXiv:2501.12862. The 34-point acceptance spread between two teams on the same tool is the finding worth carrying.
  171. 171Mohammed Latif Siddiq, Zhao, Lopes, Casey, and Santos, “Security in the Age of AI Teammates: An Empirical Study of Agentic Pull Requests on GitHub,” Information and Software Technology, arXiv:2601.00477.
  172. 172Rahul Gopu and David Apirian, “60 Million Copilot Code Reviews and Counting,” The GitHub Blog, March 5, 2026. Vendor-published; comment acceptance rate, merge-time effect, and code-quality delta are all absent.
  173. 173Danny Hsu, M. Neu, M. Farrag, and R. Kindi, “Leveraging AI for Efficient Incident Response,” Engineering at Meta, June 24, 2024. 42% is ranking accuracy on a constrained candidate set, not autonomous resolution; no MTTR delta published.
  174. 174Toufique Ahmed, Supriyo Ghosh, Chetan Bansal, Thomas Zimmermann, Xuchao Zhang, and Saravan Rajmohan, “Recommending Root-Cause and Mitigation Steps for Cloud Incidents Using Large Language Models,” ICSE 2023, arXiv:2301.03797.
  175. 175Zhao, Tang, and Qian, “Do Deployment Constraints Make LLMs Hallucinate Citations?,” arXiv:2603.07287, March 7, 2026. 17,443 generated citations; cited here as the closest quantified analogue to documentation risk. Preprint.
  176. 176Danilo Poccia, “Amazon Q Code Transformation,” AWS News Blog, November 28, 2023, updated April 30, 2024. The widely cited “30,000 applications / 4,500 developer-years / $260M” figure originates in executive statements of August 2024 with no published methodology and should be cited as an executive claim, never as a measurement.
  177. 177Goran Petrović and Marko Ivanković, “State of Mutation Testing at Google,” ICSE-SEIP ‘18.
  178. 178Goran Petrović, Marko Ivanković, Gordon Fraser, and René Just, “Practical Mutation Testing at Scale: A View from Google,” IEEE Transactions on Software Engineering, August 2021, DOI 10.1109/TSE.2021.3107634.
  179. 179Zhao, Zhou, and Cohen, “Do Coverage and Mutation Scores of LLM-Generated Test Suites Correlate with Their Effectiveness? (Replicability Study),” PACMSE, July 2026, DOI 10.1145/3832093.
  180. 180Farima Mehrpour and Thomas D. LaToza, “Can Static Analysis Tools Find More Defects? A Qualitative Study of Design Rule Violations Found by Code Review,” Empirical Software Engineering 28, no. 1 (November 2022), DOI 10.1007/s10664-022-10232-4.
  181. 181Alberto Bacchelli and Christian Bird, “Expectations, Outcomes, and Challenges of Modern Code Review,” ICSE 2013. Still the best evidence on what review actually does versus what practitioners believe it does.
  182. 182Wang, Pradel, and Liu, “Are ‘Solved Issues’ in SWE-bench Really Solved Correctly? An Empirical Study,” arXiv:2503.15223, March 19, 2025. Preprint.
  183. 183Jiawei Liu, Chunqiu Steven Xia, Yuyao Wang, and Lingming Zhang, “Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code Generation” (EvalPlus), NeurIPS 2023, arXiv:2305.01210.
  184. 184Jesse Toth, “Scientist,” GitHub Engineering Blog, February 3, 2016, updated December 3, 2020. The canonical documented shadow-verification implementation; no defect-catch figures published.
  185. 185Gabor, Lynch, and Rosenfeld, “EvilGenie: A Reward Hacking Benchmark,” arXiv:2511.21654v2, May 17, 2026. Preprint; Cambridge Boston Alignment Initiative and MIT FutureTech.
  186. 186Salim, Latendresse, Khatoonabadi, and Shihab, “Tokenomics: Quantifying Where Tokens Are Used in Agentic Software Engineering,” arXiv:2601.14470, January 20, 2026. n = 30 tasks, one framework, one model; its “review” stage is a model reviewing model output, not CI compute. Preprint.
  187. 187Atif Memon, Zebao Gao, Bao Nguyen, Sanjeev Dhanda, Eric Nickell, Rob Siemborski, and John Micco, “Taming Google-Scale Continuous Testing,” ICSE-SEIP 2017.
  188. 188Haider and Zimmermann, “Understanding Dominant Themes in Reviewing Agentic AI-authored Code,” MSR ‘26, arXiv:2601.19287. 19,450 inline review comments across 3,177 agent-authored pull requests.
  189. 189He, Shao, Chen, Gao, Zhang, and Sheng, “Use Property-Based Testing to Bridge LLM Code Generation and Validation,” arXiv:2506.18315. The paper assumes rather than demonstrates that model-generated properties are more reliable than model-generated implementations. Preprint.
  190. 190Ma, Zhang, Cao, Liu, Zhang, Luo, Zhang, and Chen, “Rethinking Verification for LLM Code Generation: From Generation to Testing,” arXiv:2507.06920v2, July 2025. Names the “homogenization trap.” Preprint.
  191. 191Kubernetes SIG Apps, “Agent Sandbox,” kubernetes-sigs/agent-sandbox, v0.4.6, May 14, 2026. Pre-1.0.
  192. 192Anthropic, “Claude Code Sandboxing,” Engineering blog, October 20, 2025, and “Choose a Sandbox Environment,” Claude Code documentation. Vendor-published; the sandbox runtime is open-sourced and the architecture claims are verifiable in the released code.
  193. 193Matt Mathew, Prasad Borole, Meng Huang, Sergey Burykin, Gaurav Goel, and Bayard Walsh, “Solving the Identity Crisis for AI Agents,” Uber Engineering Blog, May 21, 2026. First-party; the strongest published enterprise agent-identity implementation.
  194. 194Amit Arora and Omri Shiv, “Governing AI Assets at Scale with MCP Gateway and Registry,” AWS Open Source Blog, June 17, 2026. Vendor-published, Apache-2.0, genuinely implementable.
  195. 195Spotify, “Portal MCP / Actions Registry,” Backstage documentation; and Tyson Singer, “Introducing Xirp,” Spotify Portal blog, August 10, 2026. Vendor-published; internal adoption figures self-reported.
  196. 196Natalie Lung, “Uber Caps Usage of AI Tools Like Claude Code to Manage Costs,” Bloomberg, June 2, 2026 (paywalled), corroborated by Simon Willison, June 3, 2026.
  197. 197CNCF and SlashData, “Platform Engineering Tools Maturing as Organizations Prepare for AI-Driven Infrastructure,” March 24, 2026. Survey of 400+ professional developers, fielded Q4 2025.
  198. 198Gartner, “2026 Hype Cycle for Agentic AI,” April 15, 2026. 17% of organizations have deployed AI agents; over 60% expect to within two years. Analyst-published; underlying sample not disclosed publicly.
  199. 199Yasmin Moslem and John D. Kelleher, “Dynamic Model Routing and Cascading for Efficient LLM Inference: A Survey,” arXiv:2603.04445v2, April 21, 2026. Reported figures are benchmark results from individual papers, not enterprise production results. Preprint.
  200. 200GitHub, “About Premium Requests,” GitHub Docs. Vendor documentation; cited as product fact for the quota mechanism.
  201. 201Northflank, “AI Sandbox Pricing Comparison (2026),” May 5, 2026. Vendor comparing itself against competitors; cite only the third-party list prices, which are independently checkable.
  202. 202Sergio De Simone, “AI Code Review at Scale: LinkedIn’s Multi-Agent Approach,” InfoQ, August 22, 2026. 5,230 sampled review comments across 1,727 pull requests.
  203. 203Bowen Qin and Yi Xie, “Agent Retrieval Bench: Evaluating Repository Context Retrieval for Coding Agents,” arXiv:2607.24882, July 2026. 427 samples across 25 repositories. Preprint.
  204. 204Taj Shorter, “Inside Shopify’s AI-First Engineering Playbook,” Bessemer Venture Partners, April 1, 2026. Third-party interview; figures self-reported by Shopify with no methodology.
  205. 205Microsoft, “What Is Microsoft Entra Agent ID,” Microsoft Learn, documentation dated April 14, 2026, updated June 24, 2026. No formal general-availability date is stated in the documentation; do not assert GA.
  206. 206modelcontextprotocol/registry, community MCP registry, in preview with the API frozen at v0.1.
  207. 207Erik Brynjolfsson, Bharat Chandar, and Ruyu Chen, “Canaries in the Coal Mine? Six Facts about the Recent Employment Effects of Artificial Intelligence,” Stanford Digital Economy Lab, August 2026; and “Canaries, Interest Rates, and Timing,” February 9, 2026. ADP administrative payroll microdata, balanced panel of 3.5–5 million employees per month — the highest-quality evidence in this domain.
  208. 208SignalFire, State of Tech Talent Report 2026. Assigns primary causation to the end of the zero-interest-rate era rather than to AI; genuinely disagrees with reference 207 on cause while agreeing on fact.
  209. 209Hao-Ping Lee, Advait Sarkar, Lev Tankelevitch, Ian Drosos, Sean Rintel, Richard Banks, and Nicholas Wilson, “The Impact of Generative AI on Critical Thinking,” CHI 2025, DOI 10.1145/3706598.3713778. n = 319 knowledge workers, 936 usage instances.
  210. 210Raja Parasuraman and Dietrich H. Manzey, “Complacency and Bias in Human Use of Automation: An Attentional Integration,” Human Factors 52, no. 3 (June 2010): 381–410, DOI 10.1177/0018720810376055.
  211. 211Atlassian, State of Developer Experience 2025, July 9, 2025. n = 3,500 developers and managers across six countries. Vendor-published.
  212. 212Matthew Skelton, “Team Topologies as the Infrastructure for Agency with AI,” QCon London, March 2026, reported InfoQ, March 2026; and Olivier Wulveryck, “Who Does What? Team Topologies for the Agentic Platform,” June 22, 2026. Conference-stage and practitioner reasoning respectively; neither is measured.
  213. 213Duolingo, AI-first memo (April 2025) and subsequent clarifications, with the substantive reversal reported in Fortune, April 13, 2026. Primary memo text not retrievable; all memo language is secondhand through contemporaneous reporting.
  214. 214Klarna, “Klarna AI Assistant Handles Two-Thirds of Customer Service Chats in Its First Month,” press release, February 27, 2024; reversal reported Entrepreneur, May 2025; steady state reported Fortune, October 10, 2025. The agent-equivalence figure is 700 in the primary source and 800 in later coverage; use 700.
  215. 215Salesforce, statements by Marc Benioff on the Q4 FY25 earnings call, February 26, 2025; hiring announcement reported Fortune, April 27, 2026; and earnings call, May 28, 2026. The 30% engineering productivity claim has never been substantiated or repeated with a methodology.
  216. 216Nicholas Carlini, “Building a C Compiler with a Team of Parallel Claudes,” Anthropic Engineering, February 5, 2026. n = 1 human; vendor-published.
  217. 217Microsoft, Work Trend Index 2025: The Year the Frontier Firm Is Born, April 24, 2025. Coined “human-agent ratio” and explicitly supplied no formula, benchmark, or tested figure.
  218. 218Ipek Ozkaya, Anita Carleton, Sebastián Echeverría, et al., The AI Adoption Maturity Model v1.0, Carnegie Mellon Software Engineering Institute, June 30, 2026. General AI adoption rather than agentic engineering specifically.
  219. 219Sabry E. Farrag, “The Productivity-Reliability Paradox: Specification-Driven Governance for AI-Augmented Software Development,” University of East London, arXiv:2605.01160, May 2026. Multivocal review of 67 sources; a proposal, not validated in the field. Preprint.
  220. 220Indeed Hiring Lab software development postings index, FRED series IHLIDXUSTPSOFTDEVE, 74.57 as of August 21, 2026; and Lightcast forward-deployed-engineer posting data reported in IT Brew, December 19, 2025.
  221. 221IBM entry-level hiring statements, reported Axios, February 13, 2026. Executive claim; no published outcome data.
  222. 222Coinbase, statements by Brian Armstrong on the Cheeky Pint podcast, August 22, 2025, reported TechCrunch, August 22, 2025. The number of engineers dismissed was never disclosed; the reported AI-authored code percentage is unverified.
  223. 223Tobi Lütke, internal memo posted publicly to X, April 7, 2025, reported Inc., April 2025. No reversal reported as of August 2026, and no published outcome data.
  224. 224Nx Team, security advisory GHSA-cxm3-wv7p-598c, nrwl/nx, August 27, 2025. First-party; contains the incident timeline and the enumerated remediations.
  225. 225Snyk, “Weaponizing AI Coding Agents for Malware in the Nx Malicious Package,” August 2025. Vendor-published; reproduces the agent-invocation code verbatim.
  226. 226StepSecurity, “Supply Chain Security Alert: Popular Nx Build System Package Compromised with Data-Stealing Malware,” August 27, 2025. Vendor-published; per-version publish timeline in UTC.
  227. 227Wiz Research, “s1ngularity Supply Chain Attack,” August 27, 2025, and “s1ngularity’s Aftermath,” September 2025. Vendor-published telemetry analysis; the refusal-rate findings are uncorroborated, and the firm revised its own second-phase repository count upward mid-investigation.
  228. 228Ashish Kurmi (StepSecurity) and Charlie Eriksen (Aikido), quoted in Connor Jones, “Supply Chain Attack Hits Nx Build System,” The Register, August 27, 2025.
  229. 229Cybersecurity and Infrastructure Security Agency, “Supply Chain Compromises Impact Nx Console and GitHub Repositories,” alert, May 28, 2026. A separate, later incident affecting the same ecosystem; CVE-2026-48027.
  230. 230Jason Lemkin, public posts on X and SaaStr, July 2025, including the retrospective. Single-participant self-published account; figures are as reported and were not independently verified.
  231. 231Amjad Masad, public thread on X, July 21, 2025, with remediation detail corroborated in Connor Jones, “Replit Responds,” The Register, July 22, 2025, and Fast Company, July 21, 2025. No written postmortem was subsequently published.
  232. 232Andrew Cornwall (Forrester Research), Matthew Flug (IDC), and Torsten Volk (Enterprise Strategy Group), quoted in Beth Pariseau, “Replit AI Agent Snafu ‘Shot Across the Bow’ for Vibe Coding,” TechTarget, July 2025.
  233. 233AI Incident Database, incident 1152. Characterizes the events as reported and alleged; the claims have not been independently verified.
  234. 234Uber, “How Uber Executed a JUnit Migration at Massive Scale,” Uber Blog, April 7, 2026. First-party; contains the statement that generative AI was attempted for the migration and abandoned in favor of deterministic transformation.
  235. 235Spotify, “Fleet Management at Spotify (Part 3): Fleet-wide Refactoring,” Spotify Engineering, May 2023; and “Part 1: Spotify’s Shift to a Fleet-First Mindset,” April 2023. First-party; no language models are described in either, and no failure or rollback rate is published.
  236. 236Spotify, “Background Coding Agents: Context Engineering,” Spotify Engineering, November 2025. First-party.
  237. 237Spotify, “Background Coding Agents: Predictable Results Through Strong Feedback Loops,” Spotify Engineering, December 2025. First-party; source of the judge-veto rate.
  238. 238Devon Edwards Joseph, “Background Coding Agents: Supercharging Downstream Consumer Dataset Migrations,” Spotify Engineering, April 22, 2026. First-party; the manual-effort comparison is an estimate.
  239. 239Nadia Alshahwan, Mark Harman, Inna Harper, Alexandru Marginean, Shubho Sengupta, and Eddy Wang, “Assured LLM-Based Software Engineering,” arXiv:2402.04380, February 6, 2024. The methodological parent of the two Meta deployment papers.
  240. 240Mark Harman, Peter O’Hearn, and Shubho Sengupta, “Harden and Catch for Just-in-Time Assured LLM-Based Software Testing: Open Research Challenges,” FSE ‘25 Companion, arXiv:2504.16472. Source of the statement that only four of the six claimed assurances are verifiable guarantees.
  241. 241Itamar Friedman, “We Created the First Open-Source Implementation of Meta’s TestGen-LLM,” Qodo (formerly CodiumAI), May 20, 2024. Vendor-published by a competitor; the only substantive independent engagement with the TestGen-LLM paper located.
  242. 242Uber, “Running a Software Factory Efficiently at Uber Scale,” Uber Blog, August 27, 2026. First-party; adoption and cost-trend figures self-reported.
  243. 243Reddit thread by u/NegativeWeb1, May 2025, and coverage in Habr, “On Reddit, They Discovered That AI Copilot on GitHub Is Slowly Driving Microsoft Employees Crazy,” May 21, 2025; primary artifact (Copilot-authored, opened May 20, 2025, closed unmerged). Practitioner reaction, not an evaluation.

Version 1.0, September 2026. This framework is offered as practitioner guidance and does not constitute legal, regulatory, or financial advice. Standards versions, regulatory dates, and published figures were verified against primary sources in September 2026 and move quickly; verify all regulatory mappings against current obligations in your jurisdiction before relying on them. Where a source is vendor-published, correlational, self-reported, or an unreviewed preprint, the reference entry says so. Appendix C lists the pending events that would occasion a revision.

PDF