The Week the Copilots Became One
On August 13, Microsoft began merging its consumer and commercial Copilot applications, rolling the change out to a small subset of users with a color-coded cue indicating whether someone was signed in to a business or personal account. The Microsoft 365 app becomes the Microsoft Copilot app. Podcasts, Group Chat, and consumer Deep Research began retiring on August 18.1
Those are small changes, but the thing they are structural groundwork for is not. Two weeks earlier, on the fiscal fourth-quarter earnings call, Satya Nadella put the plan on the record: “Copilot is evolving rapidly, from chat to Cowork to Autopilots. This quarter, we are bringing these Copilot experiences together, including code, in one super app.” Pressed later in the same call on how Microsoft 365 Copilot converts from pilots to deployments, he returned to product shape rather than pricing: chat, Cowork, Autopilot, and Code all converging into “this flagship super app that various roles can use it.”2
Read that last clause carefully, because it is the entire premise. Not one app for one role but one app for every role, from the developer running an agent against a repository to the analyst reconciling a spreadsheet. Fortune first reported the project in May, describing an internal effort running under the slogan “Delivering one Copilot,” led by Jacob Andreou, promoted to head of Copilot in March after Nadella consolidated the previously separate consumer and commercial teams. The reporting was blunt about the motivation: Microsoft found that customers dislike shifting between Copilot tools, and fewer than 4.5 percent of the 450 million customers of its Microsoft 365 suite were paying for Copilot features at the time.3
Microsoft is not early here, and it is not alone. OpenAI absorbed its standalone Codex desktop application into a unified ChatGPT desktop app on July 9, launching ChatGPT Work alongside it and beginning to sunset the Atlas browser.4 Anthropic released Claude Cowork in January as a knowledge-work companion to Claude Code, took it generally available, and extended it to web and mobile in July.5 Cursor pushed outward from the editor into cloud, CLI, mobile, Slack, and GitHub surfaces.6 Four companies, four starting positions, one convergent bet.
The signal beneath the signal The super app is a procurement event before it is a workflow event. Enterprises are being offered consolidation at the contract layer while the number of places their people actually work continues to climb.
What the Failed Super Apps Were Missing
Western technology has attempted this before and mostly lost. X has pursued an “everything app” for years. Uber wanted to be the operating system for city life. PayPal, Meta, and most recently Disney have all bundled toward the same horizon. The post-mortems converge on a structural explanation rather than a cultural one: super apps flourished in markets where the smartphone arrived before entrenched desktop habits and where payment and identity infrastructure was immature enough that an integrated app solved a genuine friction problem. American users, by contrast, learned modular tools first, one app for maps and another for music, and never developed the switching pain that a super app relieves.7
M.G. Siegler put the design objection more sharply than most, observing that once a service accumulates enough menus and drop-downs, you have reinvented Microsoft Office, which is tolerable for business and miserable for consumers.8 It is a fair warning, and it lands differently in an enterprise context than a consumer one, because enterprises have been buying suites happily for thirty years.
What is different this time is that AI supplies something the previous attempts never had: a reason for the parts to know about each other. A ride-hailing app and a payments app share a user account and nothing else. An agent that drafted your quarterly narrative and an agent that opened the pull request implementing the pricing change can share context, memory, permissions, and audit trail. The consolidation has a technical rationale rather than a merchandising one, which is why it deserves a different reading than the last five attempts.
That rationale is real. Whether it translates into one application is a separate question, and the evidence on that point is not going the way the marketing does.
The Buy Side and the Use Side Have Parted Ways
Enterprise procurement is genuinely consolidating. Gartner projects worldwide end-user spending on AI models and platforms will reach $64.3 billion in 2026, up 63.4 percent from $39.3 billion, and its framing of who wins is explicitly about control rather than capability: the biggest winners will be vendors that help enterprises manage where and how AI is used across the business. As spending shifts toward usage-based pricing, senior principal analyst Arunasree Cheparthi expects buyers to gravitate toward platforms that help them choose tools, monitor performance, enforce policy, and hold costs down.9 That is a consolidation thesis, and it is well founded.
The practitioner data points the other way. JetBrains published fresh results from its Developer Ecosystem Survey on August 18, drawing on more than 15,000 professional developers worldwide and statistically reweighted for global representation. Between May and July this year, 90 percent of professional developers were using AI coding agents at work at least weekly, with 68 percent using them daily. Underneath that headline, the tool distribution is startling. Claude Code reached roughly 39 percent adoption at work worldwide, up from 18 percent in January, and 47 percent in the United States. Codex climbed from 3 percent to 16 percent. GitHub Copilot declined from 29 percent a year earlier to 21 percent. Cursor slipped from 18 percent to 12 percent. OpenCode reached 7 percent, Google Antigravity 6 percent, and JetBrains AI around 9 percent.10
Add those figures together and they clear 100 percent by a comfortable margin. That is not a methodological error; instead, it is the finding. Developers are not choosing a tool; they are assembling a toolkit, and the toolkit is getting larger in exactly the period when their employers are trying to reduce vendor count.

One more line from that survey deserves an enterprise reader’s attention: 39 percent of GitHub Copilot users access it, among other surfaces, inside JetBrains IDEs. The vendor ships a first-party editor, a first-party desktop application, and a first-party CLI, and two in five of its users still reach it through somebody else’s interface. Multiply that behavior across a portfolio and the shape of the problem becomes clear. The seat is consolidated. The surface is not.
The most direct evidence comes from a vendor publishing data against its own marketing. Anthropic sampled 1.2 million anonymized Claude Cowork sessions from more than 600,000 organizations between May 11 and May 31, classifying them into a twenty-category taxonomy. Business process and operations work accounted for 33.4 percent of sampled sessions and content creation for 16.4 percent. Software development came to 8.7 percent.11 Cowork descends directly from Claude Code, is marketed on that lineage, and more than nine in ten of its sessions have nothing to do with code.

That is not a failure. It is a clean natural experiment. Given one vendor, one account, and a genuine choice of surface, developers went to the terminal and everyone else went to the desktop app. The populations the super app promises to unite selected themselves apart.
Seats, Surfaces, and the Inference Paradox
For a CIO, the divergence creates three exposures that arrive on different budget lines.
The first is utilization. Microsoft reported more than 30 million paid Microsoft 365 Copilot seats in the fourth quarter, with net seat adds more than doubling sequentially, alongside $90.0 billion in quarterly revenue and 43 percent Azure growth.12 Against a commercial base above 450 million, that is meaningful traction and thin penetration at the same time, and both readings are defensible. Nadella’s own counter is worth quoting fairly: he told analysts that the time from license purchase to high usage has collapsed from months to days and that usage intensity now sits on par with Outlook and Teams. The honest position is that the top-line number tells you what was sold and nothing about what your organization will use. That gap is the thing to instrument, and it does not close because the icon changed.
The second is inference economics, and it is the most underpriced risk in this entire category. Gartner published a prediction on August 17 that inference costs per agentic workflow will rise more than fivefold through 2028, describing what it calls the Inference Paradox: better unit economics escalating the total cost of AI without a clear path to commensurate value. Routing a task to an agentic reasoning model rather than a basic chatbot raises provider inference costs by at least five times, and considerably more as complexity grows. Senior director analyst Will Sommer framed the consequence for anyone building or buying these systems: “Defaulting to generic autonomous intelligence will result in unbounded costs orders of magnitude higher than those of optimized product ecosystems.”13
Sit that alongside the super app premise. A single general-purpose front door, used by every role, that routes a request for a meeting summary through the same reasoning-heavy path as a multi-file refactor is a textbook instance of defaulting to generic autonomous intelligence. Specialized surfaces are not merely a habit developers refuse to give up. They are a tiering mechanism, and tiering is where the margin lives. Gartner’s own forecast reinforces the point from the spending side: domain-specific and specialized models are projected to grow 210 percent this year, faster than any other segment in the market.14
The third exposure is governance, and it arrives last and hurts longest. GitHub Copilot revenue accelerated more than 60 percent quarter over quarter following a June shift to usage-based pricing, and Microsoft’s Agent 365 registered nearly 40 million agents across tens of thousands of companies within two months of launch.15 Consumption pricing plus agent proliferation plus a consolidated front door is a combination that makes spend easy to incur and hard to attribute. If the super app becomes the place where every role does everything, the audit question shifts from “who has a license” to “which role, running which agent, on which model tier, against which data, at what cost.” Very few organizations can answer that today.
“Product leaders cannot rely on more efficient token economics to rationalize AI costs. Each successive generation of AI capability will necessitate more, and often more expensive, tokens.”
— Will Sommer, Sr. Director Analyst, Gartner, August 17, 2026
Four Companies, One Bet, Four Different Hedges
The most useful thing about the vendor landscape right now is that all four leading builders have publicly hedged the very consolidation they are marketing. Read the architecture rather than the announcement and a consistent pattern emerges: what is being unified is the shell, not the work.
OpenAI’s July 9 release is the clearest case. The Codex desktop app became the ChatGPT desktop app, and what sits inside is a three-mode environment: Chat, Work, and Codex, on every plan including Free. Codex was not folded into a general assistant; it was preserved as a distinct mode with its own icon option and default view, and Sam Altman stated plainly that Codex is not going anywhere. The Codex CLI persists untouched.16 One binary, three doors, and the developer door still opens onto a terminal.
GitHub made the same architectural concession more explicitly, and in some respects more thoughtfully. The Copilot app launched at Build in June as an agent-native desktop control center, with a single My Work view, isolated git worktrees per session, canvases as bidirectional human-agent work surfaces, and cloud and local sandboxes. In the same release, GitHub shipped a redesigned Copilot CLI with voice input and scheduled tasks explicitly to keep terminal-first developers in the terminal, and positioned the app as complementary to VS Code rather than a replacement. GitHub’s own announcement calls it “one runtime, many surfaces,” and describes Memory and the /chronicle capability as carrying context across sessions started in the app, the CLI, VS Code, or on GitHub.17 That is a company consolidating the runtime and the context layer while deliberately declining to consolidate the interface. It is a more defensible architecture than the super app framing suggests, and it is worth noting that the same organization’s marketing sits beneath a corporate narrative pushing the opposite.
Anthropic runs two products rather than one, and its own telemetry explains why. Claude Code owns the terminal; Cowork owns the desktop, web, and mobile surfaces for the other ninety-one percent of sessions. The company has not attempted to merge them, and the usage data suggests that restraint is empirical rather than accidental.
Cursor presents the most interesting measurement puzzle. Its enterprise footprint and revenue trajectory have grown rapidly while JetBrains’ globally representative survey shows work adoption falling from 18 percent to 12 percent over the same window. Both can hold simultaneously if revenue is concentrating among fewer, larger, higher-spending accounts while breadth thins, and that reading has direct procurement implications for anyone standardizing on it. The point is not that any vendor is failing, but rather, seat count, revenue, and actual practitioner adoption have become three different numbers, and vendors will quote whichever one flatters.

| Layer | Consolidating? | What a CIO should require |
|---|---|---|
| Commercial contract and identity | Yes—genuinely | Single tenant, SSO, unified entitlements |
| Context, memory, and audit trail | Yes—and this is the real prize | Portable context across every surface |
| Governance and policy enforcement | Partly—varies sharply by vendor | Agent registry, per-role policy, spend caps |
| Model routing and inference tier | No—and it should not | Task-appropriate tiering, not one default |
| Interface and daily surface | No—practitioners are diverging | Fund CLI, IDE, and app as one entitlement |
Buying the Shell Without Losing the Rooms
The strategic error available here is not adopting a super app. It is buying interface consolidation and booking the savings as if workflow consolidation came with it. A few moves separate the two.
Negotiate at the entitlement layer, not the application layer. Ask any vendor whether a single seat covers every surface the seat holder will actually use, including the terminal, the IDE plugin, the desktop app, and the mobile client. If the answer requires a second SKU for the CLI, the consolidation is cosmetic, and your engineers will route around it within a quarter.
Instrument utilization by role before you standardize. Aggregate seat activation is nearly meaningless in a product where a developer and a financial analyst have almost nothing in common except a login. Break adoption down by function, measure at daily rather than the vendor-default 28-day window, and expect the distribution to be bimodal. If it is not, you are probably measuring license assignment rather than work.
Treat model tiering as a procurement requirement, not an implementation detail. Ask explicitly how the platform routes a low-complexity request, whether administrators can set tier policy per repository or per workload, and what the cost differential looks like across tiers. GitHub’s configurable low and medium review effort levels are a concrete example of what this control looks like when a vendor builds it; the useful move is to demand the equivalent from everyone on the shortlist and to be skeptical of any platform that cannot show it.
Fund the specialized surface deliberately. Practitioner evidence is unambiguous: terminal-native and IDE-native tools are where a large share of senior engineering work happens, and the same is now true in reverse for knowledge workers who will never open a shell. Budget for both, resist the temptation to declare a single blessed interface, and let usage rather than architecture diagrams decide where people work. A short set of practices helps here: publish an approved-surface list rather than an approved-app list, standardize agent configuration files across tools so context is portable, require every agent action to carry role and cost attribution, keep one sanctioned sandbox pattern for agent execution, and review the surface list quarterly rather than annually.
Finally, plan for the audit you cannot currently pass. Agent registration, per-role policy, and spend attribution are the controls that make consolidated agentic work defensible, and they are uneven across the field. Ask for them by name in the evaluation, not after the rollout.
The Question Behind the Front Door
Nadella described a flagship app “that various roles can use,” and that is a reasonable thing to build. The evidence simply suggests the roles will keep using it differently, and that the value of consolidation sits behind the interface rather than in it. Portable context, one identity, one policy engine, one bill you can actually read: those are worth restructuring a contract for. A shared icon is not. The organizations that get this right will negotiate for the runtime, instrument the rooms separately, and stay indifferent about which door their people walk through.
References
- Sebastian Herrera, “Microsoft Begins to Merge Consumer and Enterprise Copilot Apps in Push for Super App,” Fortune, August 13, 2026.
- Microsoft Corporation, FY26 Q4 earnings call transcript, July 29, 2026 (prepared remarks and analyst Q&A).
- Sebastian Herrera, “Exclusive: Microsoft Is Building a Super App That Combines Coding, Chat, and Other Copilot AI Tools,” Fortune, May 29, 2026.
- “OpenAI Launches ChatGPT Work and Unveils Unified Desktop App with Codex Built In,” Neowin, July 9, 2026; OpenAI, “ChatGPT Is Now a Partner for Your Most Ambitious Work,” July 9, 2026.
- Anthropic, Claude Cowork release notes and product announcements, January–July 2026; “Anthropic Says Claude Code Transformed Programming. Now Claude Cowork Is Coming for the Rest of the Enterprise,” VentureBeat, February 25, 2026.
- Cursor (Anysphere) product documentation and changelog, 2026 (cloud, CLI, mobile, Slack, and GitHub surfaces).
- Dev Patnaik, “Disney’s Super App Dilemma,” Forbes, May 21, 2026.
- M.G. Siegler, “The Age of the ‘Super App’ — Again and Again,” Spyglass, May 3, 2026.
- “Gartner Forecasts Worldwide AI Platforms and Models Market to Grow 63% in 2026,” Gartner, July 20, 2026.
- Mikhail Bogdanov, “AI Coding Agents: Adoption Trends,” JetBrains Research, August 18, 2026, based on the Developer Ecosystem Survey 2026 (more than 15,000 professional developers, statistically reweighted).
- Anthropic, “How People Are Using Claude Cowork,” July 2026 (1.2 million sampled sessions, more than 600,000 organizations, May 11–31, 2026).
- Microsoft Corporation, FY26 Q4 results and earnings call, July 29, 2026.
- “Gartner Predicts AI Inference Costs Per Agentic Workflow Will Increase More Than Fivefold Through 2028,” Gartner, August 17, 2026.
- “Gartner Forecasts Worldwide AI Platforms and Models Market to Grow 63% in 2026,” Gartner, July 20, 2026 (Table 1, DSLMs and specialized GenAI models).
- Microsoft Corporation, FY26 Q4 earnings call, July 29, 2026 (GitHub Copilot revenue growth and Agent 365 registrations).
- “ChatGPT Work and Codex Now Share One Desktop App: What Actually Changed,” Developers Digest, July 15, 2026; Neowin, July 9, 2026.
- Mario Rodriguez, “GitHub Copilot App: The Agent-Native Desktop Experience,” The GitHub Blog, June 2, 2026.