← All writing

Tenants of the Same Building: What September 3 Revealed About AI Failover

Three AI providers degraded inside one overlapping window. The explanation the industry reached for was the CDN layer. The only shared dependency anybody has documented sits lower and is far harder to contract around: a single leased building in Memphis, occupied by a lab and its own competitor.

Three Hours in Memphis

At 6:30 a.m. Pacific on Thursday, September 3, Grok stopped answering. It stayed down for roughly three and a half hours. Later that day SpaceXAI posted an explanation and an apology, and the second sentence is the one worth reading twice.1

“We are sorry for the issues you may have experienced with Grok following an outage at our Memphis compute center this morning. We’d also like to apologize to our impacted compute partners.”

— SpaceXAI, September 3, 2026

Compute partners. Plural, unnamed.

Anthropic’s status page logged an incident that opened at 13:26 UTC and closed at 16:23 UTC, two hours and fifty-seven minutes of elevated errors across seven models in three families: Mythos 5.1 and Mythos 5, Fable 5.1 and Fable 5, and Opus 5, 4.8, and 4.6.2 OpenAI posted its own incident the same morning, “Elevated errors across ChatGPT and Codex,” covering fifteen ChatGPT components and four in Codex. OpenAI’s status page carried no root cause; The Register reported that a routing error had made both products unavailable for some users, with impact running roughly 7:43 to 8:17 a.m. Pacific.3

Inside a single overlapping window, three providers and four widely used products had degraded, and the industry did what it always does when several things break at once. It went looking for a shared upstream.

Cloudflare, the usual suspect, said flatly that it was “not experiencing any significant service disruptions.”4 The major cloud status dashboards showed nothing relevant. Neither Anthropic nor OpenAI confirmed any link to the Memphis event, and SpaceXAI did not name the partners it had apologized to.5 So the definitive causal chain does not exist in public, and this article will not pretend otherwise.

What does exist in public is a lease. On May 6, 2026, Anthropic announced an agreement to use all of the compute capacity at Colossus 1, the SpaceXAI facility on an industrial tract of South Memphis adjacent to the Boxtown neighborhood: more than 300 MW and over 220,000 GPUs by Anthropic’s own description.6 Independent tracking puts the site at 340 MW of IT power and roughly 276,000 H100-equivalents, assembled from about 150,000 H100s, 50,000 H200s, and 30,000 B200s, at a capital cost near $12.9 billion.7

One building. One landlord who is also a competitor. Two frontier labs whose products degraded inside the same window.

The signal The interesting fact about September 3 is not that several AI services broke at once, but that the only shared dependency anybody has actually documented sits at a layer almost no enterprise procurement process inspects, and that the layer is physical.

The Explanation Everyone Reached For

In the first hours after the outages, the analyst commentary converged on a familiar hypothesis. Brian Jackson, principal research director at Info-Tech Research Group, observed that simultaneous failures of this kind could stem from shared infrastructure such as a CDN, DNS, or a cloud platform.8 That was a reasonable read, and parts of it remain unresolved. What can be said with confidence is narrower: the CDN hypothesis was ruled out within hours, and the one shared dependency anybody has since documented sits a layer beneath all three candidates.

The instinct is worth examining because it reveals how the industry has been trained to think about correlated failure. For fifteen years the shared-upstream story has almost always been true at the network edge. It was true on February 20 of this year, when a bug in an automated cleanup task at Cloudflare misread an API query parameter, treated a request for prefixes pending deletion as a request for all of them, and withdrew 1,100 BGP prefixes covering roughly a quarter of the company’s BYOIP footprint. That incident ran six hours and seven minutes and took down core CDN, Spectrum, Magic Transit, and dedicated egress along the way.9 That is what a genuine upstream-dependency event looks like, and enterprises have built reasonable muscle memory around it.

September 3 was a different animal, at least for two of the three providers. The edge was fine. The models were fine. What appears to have wobbled beneath Grok and Claude was the floor they were standing on. OpenAI’s incident carries its own reported explanation and no publicly documented tie to Memphis, and that distinction matters: the shared-substrate story accounts for a pair, not a quartet.

There is a second thing the aftermath revealed, and it is arguably more actionable than the outage itself. In the absence of any authoritative attribution, speculative causes circulated freely, including a widely shared claim that a particular hyperscaler’s region had suffered ingress failures, resting on a single user-submitted status report and no statement from the provider named. Nobody could adjudicate it. That is not a media failure. It is a disclosure vacuum, and it exists because no party in the chain is obligated to say whose hardware served your request.

Diversification That Stops at the Contract

Here is where the story sharpens: Anthropic is not a cautionary tale about lazy infrastructure strategy. It is actually close to the opposite.

Over the eighteen months preceding the outage, the company assembled what is probably the most deliberately diversified compute posture of any frontier lab. On April 20, 2026, it expanded its Amazon relationship to up to five gigawatts of capacity, committing more than $100 billion to AWS technologies over a decade, with Amazon investing $5 billion immediately and up to $20 billion more on top of $8 billion already deployed.10 Two weeks earlier, it had announced a partnership with Google and Broadcom covering multiple gigawatts of next-generation TPU capacity beginning in 2027, disclosing alongside it that run-rate revenue had passed $30 billion, up from roughly $9 billion at the end of 2025.11 By April 2026, it was running over a million Trainium2 chips, roughly 500,000 of them activated at Project Rainier in Indiana. Its earlier October 2025 Google Cloud agreement contemplates up to a million TPUs, with well over a gigawatt coming online during 2026 alone.12

Three accelerator architectures. Two competing hyperscalers. A separate leased GPU campus. If contractual diversification were sufficient, this would be the reference implementation.

2h57m
Claude elevated-error incident, Sept 3, 2026
7
Claude models affected simultaneously
340 MW
IT power, Colossus 1, Memphis
~276k
H100-equivalents at the single leased site

Fig. 1 — The September 3 incident windows, plotted in UTC. Three providers degraded inside a span of roughly three and a half hours; only the Grok and Claude windows have a documented common dependency.

And yet seven models across three families degraded at once for the better part of three hours. Whatever the precise mechanism, the lesson survives the uncertainty: diversification measured in contracts, chip architectures, and vendor logos did not produce isolation measured in served requests. Capacity is fungible on a spreadsheet and stubbornly physical everywhere else. A given model version, at a given moment, is being served from specific racks in a specific hall drawing power from a specific substation, and no amount of multi-cloud posture at the corporate level changes that at the request level.

Enterprises have made precisely the same category error one layer up. The standard resilience pattern of 2026 routes traffic through a gateway with two or three model providers configured behind it, on the reasonable theory that competitors do not fail together. September 3 is the counterexample. Your primary and your failover can be tenants of the same landlord, and nothing in either contract would tell you.

What a Gigawatt Actually Rents

To understand why this arrangement exists, follow the money rather than the technology.

The compute layer beneath the model providers has organized itself into a fairly legible three-tier structure. At the bottom sit the landlords, frequently converted bitcoin miners with energized land and powered shells. In the middle sit the neoclouds, firms that lease those buildings and fill them with silicon. At the top sit anchor customers, the labs and hyperscalers whose demand makes the whole capital stack financeable. One market analysis puts the neocloud tier at roughly $48 billion in annualized revenue as of mid-2026, with six major operators carrying combined valuations above $150 billion, and projects the category toward $300 billion by 2030.13

The structural oddity of this tier is how little of it is owned. CoreWeave reported a revenue backlog of approximately $104 billion as of June 30, 2026, spread across forty-nine leased data centers. Fluidstack has contracted for roughly 1.4 gigawatts across five leased U.S. sites held in bankruptcy-remote vehicles and owns no buildings at all. And the leases run ten to fifteen years against silicon that turns over every two to five, a duration mismatch that obliges tenants to refill the same building with newly financed silicon three or more times inside a single lease term.13

Fig. 2 — Neocloud leases run ten to fifteen years against accelerator generations that turn over every two to five, producing a structural duration mismatch across the compute layer.

That mismatch is why capacity gets subleased, swapped, and reassigned at speed, and why a lab’s physical footprint on any given Thursday is a moving target rather than a stable fact. Colossus 1 illustrates the point neatly. It was built for one company’s training runs, and once that company moved training to a newer site, the older campus was leased in its entirety to a competitor. From the outside, nothing about Claude’s API surface changed. Underneath, the address did.

Power makes the picture more concrete still. Colossus began life with only 8 MW of grid supply, ran on fourteen mobile generators at 2.5 MW apiece while it waited, and required a $24 million substation investment plus TVA and utility approval to reach 150 MW of additional grid capacity.14 This is a facility whose electrical story has been improvised at every stage, in a region where the utility had already proposed new gas generation on the grounds that it could not meet growing demand. Uptime Institute’s 2026 analysis is blunt about where impactful failures actually originate: power systems remain the leading cause, dominated by UPS, transfer switch, and generator faults, with external fiber and connectivity issues rising fast and tending to produce the longest disruptions.15

Where the Dependency Map Goes Blank

Put the pieces together and the enterprise exposure comes into focus.

Availability at this layer is already worse than most buyers assume. The Cloud Security Alliance, in a May 2026 research note on AI compute concentration, found that neither OpenAI nor Anthropic had sustained 99% annual availability, which works out to more than three and a half days of downtime a year.16 That is roughly two orders of magnitude weaker than the availability posture enterprises demand of a payments processor or an identity provider, and yet AI inference is now embedded in workflows with comparable blast radius. Nearly three-quarters of organizations report using AI to automate processes across multiple business functions, and most have done little to account for the business interruption that creates.17

The gap is not that enterprises lack a failover plan. Many have one. The gap is that the plan is validated against the wrong dependency graph.

Resilience questionCan you answer it today?Where the answer actually lives
Which model providers serve this workload?Yes, in the gateway configYour own architecture
Which cloud regions are those providers using?Sometimes, for enterprise tiersProvider documentation, partially
Which physical facility served the request?NoProvider’s leasing arrangements, undisclosed
Do your primary and failover share a facility?NoTwo separate contracts, neither of which says
Who owns the building and the power feed?NoLandlord tier, three contracts removed from you
What are you owed when it fails?Usually service creditsNot proportionate to business interruption

Notice that the questions get less answerable as they get more determinative. The one an enterprise can answer with confidence, which providers are configured, is the one with the least bearing on whether a correlated failure takes both of them down. The technology analyst Carmi Levy warned that as AI agents take on wider-scale workflows, enterprises “could find themselves ‘uncomfortably exposed’ when AI hits the brakes,” and called the day a wakeup call to IT leaders who have “largely ignored what it’ll cost them if these increasingly critical platforms suddenly go dark.”8 The harder problem sits upstream of the costing exercise: you cannot price an exposure whose shape you are not permitted to see.

The implication A failover architecture is only as independent as the physical infrastructure underneath it, and that infrastructure is currently undisclosed by every major provider. Enterprises are therefore buying redundancy they cannot verify and, in at least one documented arrangement, may not have.

What the Builders Are Signaling

Vendors understand this better than their customers do, and the evidence is in their capital structure rather than their messaging.

Consider the lengths to which the guarantor layer has gone. Google has backstopped neocloud lease obligations at multiple sites so that landlords could issue senior secured debt, with guarantees at Lake Mariner, Barber Lake, and Abernathy reported at $3.2 billion, $1.73 billion, and $1.3 billion respectively, structured so that Google assumes the leases outright if the tenant defaults inside six years. Nvidia has committed roughly $6.3 billion in rent-back arrangements with CoreWeave through 2032 and $1.5 billion with Lambda, alongside direct equity across several operators.18

Firms do not collateralize other companies’ bonds for capacity they consider substitutable.

The second signal is disintermediation. Anthropic has since gone around the middle tier entirely, signing a direct arrangement with a landlord in Kentucky reported at roughly $19 billion over twenty years.18 Labs are pushing down the stack toward the power and the concrete, which is a rational response to exactly the exposure September 3 illustrated. The uncomfortable implication for buyers is that the labs are solving this problem for themselves, on a multi-year construction timeline, and no equivalent remedy is being extended to their customers in the meantime.

The third signal is what the labs say when describing why they run three accelerator architectures. The stated rationale includes resilience explicitly, the ability to shift workloads across clouds if one supplier is constrained.12 That capability is real at the planning horizon of weeks. It is not a hot standby, and on September 3 it did not behave like one.

Designing for the Failure You Cannot See

None of this argues for slowing AI adoption, and it does not argue for self-hosting frontier models, which almost no enterprise should attempt. It argues for a resilience posture that stops assuming vendor plurality delivers physical independence. Five moves are worth putting on the near-term agenda.

  1. Start by writing substrate disclosure into contracts. At enterprise commitment levels you have leverage, and the asks are modest: the cloud regions and facility operators serving your traffic, notification when serving capacity for your workloads migrates between facilities, and a commitment that designated primary and secondary capacity does not share a site or a power feed. Most providers will resist the third. Some will agree to the first two, and the answers alone will change how you architect.

  2. Rebuild the dependency inventory below the API. Map each AI-dependent workflow to its model provider, then to that provider’s disclosed infrastructure, then to whatever you can establish about facilities and operators. The map will have holes. Document the holes explicitly and treat each one as an accepted risk with a named owner, because an undocumented dependency is not an absent one.

  3. Separate genuine failover from nominal failover in your own architecture. A gateway with three providers configured is not resilient if you have never run a live cutover under production load, if the fallback path has different context limits or tool-calling semantics, or if prompts have been tuned to one provider’s behavior. Test the cutover on a schedule, and measure what degrades rather than assuming the traffic simply lands.

  4. Add a non-AI floor to any process where interruption is expensive. This is the least fashionable recommendation and probably the most valuable. For each critical workflow, define what happens during a three-hour degradation, whether that means queueing for later processing, routing to a smaller self-hosted open-weights model as a stopgap, or falling back to a manual path that staff have actually practiced. Levy’s point about maintaining human skill is not nostalgia; it is capacity planning.

  5. Price the exposure honestly in the business case. If the observed availability of frontier APIs is materially below 99%, then any workflow whose downtime costs more than the service credits it would receive is carrying uncompensated risk. Uptime Institute found that 57% of operators put the cost of their most recent major outage above $100,000, and one in five above $1 million.15 Those numbers belong in the AI business case alongside token spend, and today they almost never appear.

The Floor Beneath the Abstraction

The promise of the API era was that infrastructure stopped being your problem. For most purposes that promise held, and the abstraction was worth every dollar. What September 3 exposed is that the abstraction has a floor, and the floor is a converted appliance factory in a Memphis neighborhood, drawing power from a grid whose utility had proposed new generation to keep up with regional demand, leased in full to a company whose landlord competes with it. Nobody chose that arrangement as a risk posture. It emerged from a capacity shortage, a duration mismatch, and a financing structure that rewards moving fast. Enterprises inherited it without being told, and they will keep inheriting the next one on the same terms unless they start asking, in writing, what is underneath.

References

  1. SpaceXAI, post on X, September 3, 2026.
  2. “Elevated errors for multiple models,” Claude Status, incident of September 3, 2026 (13:26–16:23 UTC).
  3. “Elevated errors across ChatGPT and Codex,” OpenAI Status, incident of September 3, 2026; Thomas Claburn, “True AI-pocalypse as ChatGPT, Claude, and Grok all go down at once,” The Register, September 3, 2026. OpenAI published no root-cause analysis; the routing-error attribution is The Register’s reporting.
  4. Cloudflare statement as reported in The Register, September 3, 2026.
  5. “SpaceXAI apologizes for outage that affected Grok and other ‘compute partners,’” Engadget, September 3, 2026. Neither Anthropic nor OpenAI confirmed a connection to the Memphis facility.
  6. “Anthropic to use all of SpaceX-xAI’s Colossus 1 data center compute,” DataCenterDynamics, May 2026, reporting Anthropic’s announcement of May 6, 2026.
  7. “Colossus 1,” Epoch AI AI Data Centers directory, accessed September 2026. Note the discrepancy between Anthropic’s stated “more than 300 MW” and “over 220,000 GPUs” and Epoch’s independent estimate of 340 MW IT power and ~276,000 H100-equivalents; Epoch’s independent estimates are used where the two conflict, with Anthropic’s own lower figures noted alongside.
  8. Brian Jackson, principal research director, Info-Tech Research Group, and Carmi Levy, technology analyst, quoted in “ChatGPT, Claude, and Grok all went down at once; enterprises need a backup plan,” Computerworld, September 2026.
  9. “Cloudflare outage on February 20, 2026,” Cloudflare Blog, February 2026.
  10. “Anthropic and Amazon expand collaboration for up to 5 gigawatts of new compute,” Anthropic, April 20, 2026.
  11. “Anthropic expands partnership with Google and Broadcom for multiple gigawatts of next-generation compute,” Anthropic, April 6, 2026. TPU capacity is projected to come online beginning in 2027; the revenue figure is run-rate, not annual recognized revenue.
  12. “Inside Anthropic’s Multi-Cloud AI Factory: How AWS Trainium and Google TPUs Shape Its Next Phase,” Data Center Frontier, 2026.
  13. “Neoclouds: AI’s $150 Billion Middle Layer That Owns Almost Nothing,” Measured AI, 2026. Market sizing and the $300 billion 2030 figure are analyst projections, not reported results.
  14. “Fury from campaigners as Elon Musk’s xAI gets 150MW for Colossus supercomputer in Memphis,” DataCenterDynamics, 2025.
  15. “Annual Outage Analysis 2026,” Uptime Intelligence / Uptime Institute, 2026. Cost figures reflect 2025 survey respondents describing their most recent major outage.
  16. “AI Compute Concentration and Systemic Risk,” Cloud Security Alliance AI Safety Initiative, May 9, 2026.
  17. Arti Deshpande, Robert Stines, and Mike Vaughan, “AI is becoming a single point of failure — and most companies don’t see it,” CIO, June 10, 2026. The article states the three-quarters figure without citing an underlying survey; it is reproduced here as the authors’ assertion rather than as an independently sourced statistic.
  18. Measured AI, 2026, reporting guarantee structures at Lake Mariner, Barber Lake, and Abernathy, Nvidia rent-back arrangements, and Anthropic’s direct landlord agreement in Kentucky. Figures are as reported by the analyst and have not been confirmed in company filings.