← All writing

The Compute Price Nobody Could See

The first futures market for AI compute missed its October 5 launch after regulators questioned whether the underlying price can be trusted. Their answer describes the enterprise AI bill as well: a volatile input, priced privately, passed through by vendors whose token rates are strategic decisions.

Forty-Five Days and a Letter From Washington

On August 11, CME Group and Silicon Data, a GPU benchmarking firm backed by the trading house DRW, announced two futures contracts tied to the hourly rental price of Nvidia’s H100 and B200 accelerators. Each contract would represent a month of rent, with trading scheduled to open on October 5 under NYMEX rules.1 The contracts will now miss that date by at least a month. On September 21, the Commodity Futures Trading Commission informed the exchange that it was extending its review by forty-five days, to November 9. The agency cited what it called “novel or complex” questions, among them the risk that a fragmented and lightly observed rental market could be manipulated.2

The delay is procedural, and it may prove short. However, the document standing behind it describes the enterprise AI market as accurately as it describes GPU rentals. A month earlier, the Commission had published a formal request for comment on compute derivatives, and in it the agency described the market for AI compute more candidly than any vendor ever has. Price formation, the Commission wrote, happens primarily in opaque bilateral transactions; pricing varies dramatically across providers, regions, and contract structures; and dominant participants may wield enough pricing power to make a reference index manipulable.3 The Commission then stated its preliminary conclusion on the core question directly.

“The Commission preliminarily believes that compute may not yet exhibit certain of the characteristics of commodities that typically underlie a commodity derivatives market, including fungibility, standardization, and sufficient liquidity.”

— Commodity Futures Trading Commission, Request for Comment on the Listing of Compute Derivatives Contracts, August 21, 2026

That sentence was written about GPU-hours, yet it describes with uncomfortable precision how most large organizations buy artificial intelligence. Enterprises buy seats, credits, and tokens under privately negotiated agreements whose prices bear an undisclosed relationship to the compute underneath them, and only a minority rent GPU capacity by the hour. Whenever the futures market opens, it will produce the first public forward price for that underlying input. For buyers, therefore, the more immediate lesson is that the enterprise AI bill already carries compute-price risk that almost no one has measured.

The signal Regulators paused the first compute futures market because the price it would settle against is fragmented, privately negotiated, and difficult to verify. Enterprise AI contracts sit on top of that same price, two or three layers removed, with even less visibility into how it moves.

Most of the Market Trades in Private

Beyond the delay itself, the Commission’s request stated its preliminary understanding that non-public bilateral agreements “carry the majority of economic value” in compute markets and tend to be negotiated privately.3 Therefore, any public index measures the visible remainder of the market: on-demand rates, marketplace transactions, and posted prices. Silicon Data’s H100 index draws observations from neocloud providers, hyperscalers, colocation markets, and private rental platforms, then standardizes each one for machine specification, rental term, and geography before publication.4 CME’s own product page specifies that both contracts will track the neocloud series, meaning the non-hyperscaler segment, of on-demand rental pricing.5

That segment is useful, but it covers a narrow slice of the market. After all, most of the capacity that frontier labs and hyperscale tenants consume moves through multi-year reservations whose terms never appear in any index. Those terms have also behaved differently from the spot market. SemiAnalysis, which launched its own one-year H100 rental index in April, found that one-year contract pricing rose almost 40%, from a low of $1.70 per GPU-hour in October 2025 to $2.35 by March 2026.6 Over the same months, on-demand capacity sold out across every GPU type, and holders of locked-up instances refused to release them. Silicon Data’s neocloud H100 series, meanwhile, climbed from $2.57 on May 4 to $2.75 on July 27.7 In August, Bank of America’s semiconductor analyst Vivek Arya described GPU spot rental prices as near all-time highs, citing roughly $2.80 per hour for the H100 and $5.66 for the B200.8

Those readings share a direction, and the direction matters because the H100 is now three years old and older accelerators normally cheapen as newer silicon arrives. Instead, the price of an aging chip rose through most of 2026. The indexes agree on the trend while disagreeing on the level, and that gap between readings is precisely the problem the Commission is trying to understand.

Fig. 1 — H100 rental pricing, October 2025 through August 2026. One-year contract rates and neocloud on-demand rates both rose while the chip aged, and Bank of America's August spot estimate sits at the top of the observed range.

The manipulation concern follows from how such an index is assembled. If a settlement index is built partly from rates that capacity providers post themselves, then the parties best positioned to move the settlement price are the same parties selling the capacity. The analyst Dave Friedman illustrated the problem this month with a hypothetical in which a provider holding a short futures position temporarily lowers an eligible posted quote. If that observation carries enough weight, the provider profits on the contract. Such a provider, he noted, need not be an unusual bad actor, since providers are the natural short hedgers in compute and their lenders would want them to sell forward.9 The Commission asks nearly the same question in its own text.

A compute futures market will very likely form regardless of the delay. The policy direction is unambiguous, and CFTC Chairman Michael Selig framed the inquiry in strategic terms, stating that “America cannot win the AI race without a robust derivatives market for compute.”10 Consequently, the more plausible outcome is a market that launches later with tighter benchmark governance. Such a market would measure one slice of compute very well. For an enterprise buyer, the practical question is how far that slice sits from the price actually being paid.

Input Prices Rise While Token Prices Fall

Through the summer of 2026, the price of the input and the price of the output moved in opposite directions. While GPU rental rates held near their highs, three leading model providers each made a significant pricing decision in quick succession.

On July 30, OpenAI cut the price of GPT-5.6 Luna by 80% and GPT-5.6 Terra by 20%, leaving the flagship Sol unchanged, and attributed the reductions to efficiency work across its stack.11 On August 10, Anthropic made the introductory $2 and $10 per million token rate for Claude Sonnet 5 permanent, canceling a scheduled September 1 increase to $3 and $15 that would have raised prices by 50%.12 Google, for its part, lists its current Gemini Flash models at an introductory $0.75 and $3.75 per million input and output tokens through December 31, 2026.13 Standard pricing of $1.50 and $7.50 per million tokens takes effect on January 1, 2027.

~40%
Rise in one-year H100 rental contract pricing, Oct. 2025 to Mar. 2026
−80%
OpenAI price cut on GPT-5.6 Luna, July 30, 2026
2×
Scheduled Gemini Flash price step on January 1, 2027
Nov. 9
End of extended CFTC review of CME compute futures

The strongest objection to treating compute prices as an enterprise concern lives inside these announcements. Token prices can fall while GPU-hours get more expensive because each GPU-hour produces more tokens over time through better batching, caching, routing, and kernel optimization. OpenAI said, for example, that improvements including kernels rewritten by its own model cut the end-to-end cost of serving GPT-5.6 by 20%.11 If efficiency reliably outruns input inflation, a buyer can safely ignore the rental market. That argument is partially correct, since efficiency gains compound with every hardware generation and every serving optimization.

However, the providers’ own financial results suggest that efficiency has not insulated them from the input. According to investor documents reported by The Information, OpenAI’s inference costs quadrupled in 2025. Its adjusted gross margin fell to 33% from 40% the prior year, well short of its 46% target, in part because the company had to buy more expensive computing capacity on short notice when demand exceeded forecasts.14 Anthropic lowered its own 2025 gross margin projection by ten points, to 40%, after inference costs on Google and Amazon servers came in 23% above plan.15 These two companies rank among the largest and most sophisticated buyers of compute in the world, yet both were caught out by the price of capacity they had not secured in advance.

Fig. 2 — Direction of selected input and output price moves in 2026. GPU rental benchmarks rose while list token prices were cut, held, or discounted on a dated schedule.

Therefore, the token price an enterprise pays is best understood as a managed price: a commercial decision layered over a volatile input and shaped by competition, promotional strategy, and access to capital. When the capital absorbing the difference becomes scarcer or more expensive, the managed price tends to move, a dynamic explored in an earlier analysis of the end of the AI subsidy economy. Interestingly, a dated introductory rate is the most transparent version of that arrangement because it tells the buyer when the managed price ends.

Where the Volatility Lands Once the Seat Is Gone

For two decades, enterprise software pricing insulated buyers from input costs. A per-seat license put the vendor on the hook for whatever it cost to serve the customer, and the buyer’s exposure arrived once a year at renewal. AI pricing has been moving steadily away from that model. GitHub transitioned every Copilot plan to usage-based billing on June 1, 2026, replacing premium request units with AI Credits calculated from input, output, and cached tokens at each model’s listed API rate.16 Business and Enterprise seat prices stayed unchanged, with an allotment of credits attached to each seat. Cursor made a comparable shift a year earlier, explaining that the hardest requests on newer models cost an order of magnitude more than simple ones and that API-based pricing was the best way to reflect that difference.17

Gartner has warned that this shift among coding-agent vendors, from seat-based licensing to consumption-based pricing, introduces highly variable cost structures, and it predicts that AI coding costs will surpass the average developer’s salary by 2028.18 The broader market is moving in the same direction. Additionally, Gartner forecasts worldwide spending on AI models and platforms of $64 billion in 2026, up 63.4% from 2025, with spending on generative AI models alone growing 117%. The firm expects that as usage-based pricing becomes harder to predict, buyers will favor platforms that help them monitor and control cost.19

Usage pricing changes which party holds which risk, and the distinction is easy to lose inside a single invoice line. The table below separates the two exposures that matter: unit risk, meaning the price per token, credit, or GPU-hour, and volume risk, meaning how many units the organization ends up consuming.

Billing modelUnit-price riskVolume riskWhen compute costs reach the buyer
Flat per-seat licenseVendor absorbsVendor absorbsAt renewal, as a list-price change
Seat with pooled usage creditsVendor within allotment; buyer beyond itBuyer, above the allotmentWhenever allotments, rates, or promotions change
Per-token API at list priceVendor sets and may reviseBuyerOn any rate change, including dated introductory expirations
Committed-spend agreementDiscount fixed for the termBuyer, including shortfall on unused commitmentAt term end
Reserved GPU capacityRate fixed for the termBuyer bears utilizationAt renewal, directly at prevailing rental rates

Most enterprises now hold several of these arrangements at once, and each one passes compute costs through on a different schedule. The common thread is that consumption-based models transfer volume risk to the buyer immediately and defer unit risk to the next repricing event, which the vendor usually controls. A CFO building a 2027 AI budget today can see list prices, a handful of dated promotional schedules, and last year’s consumption trend. However, what the CFO cannot see is any forward view of the input that ultimately sets the vendor’s floor.

A futures curve would supply that third view while leaving the first two conditions untouched. A public forward price for H100 and B200 rental, extending out several years, would give finance teams a directional signal about whether the cost floor beneath their vendors is rising or falling. However, the same curve offers little as a hedge. An enterprise that pays per token for a hosted model cannot offset that exposure with an H100 rental future. Hardware generation, serving efficiency, and vendor margin all sit between the two prices, and any one of them can move independently. The exception is the minority of organizations that rent GPU capacity directly for training or for self-hosted open-weight inference. For them, a regulated rental contract may eventually become a genuine treasury instrument, with all the hedge-accounting and governance questions that implies.

Every Layer of the Supply Side Is Already Hedging

While enterprises wait for a price signal, nearly every other participant in the compute economy is building instruments to manage its own exposure.

Chip suppliers are turning hardware into financeable collateral. On August 10, Nvidia announced partnerships with Apollo, BlackRock, Blackstone, Brookfield, Goldman Sachs, and KKR to establish financing platforms intended to mobilize more than $500 billion of third-party capital for AI infrastructure over time, describing its compute as “an investable asset.”20 However, Nvidia’s subsequent quarterly filing is more measured and notes that these memorandums of understanding and other preliminary arrangements may not lead to definitive agreements.21 Additionally, critics raised a separate concern the same day, observing that a supplier arranging financing for its own customers could revive fears about circular AI financing, in which trouble at one major company ripples through the ecosystem.22

Exchanges are competing to own the benchmark, and CME was not the first to announce a product. Intercontinental Exchange announced plans in May for cash-settled GPU compute futures based on Ornn’s Compute Price Index, which Ornn describes as built only from printed transactions. ICE’s head of futures markets called the compute market “in desperate need” of a globally accepted pricing mechanism.23 Architect Financial Markets is building a venue with contracts across multiple GPU vendors, and Kalshi has published forward curves for Nvidia GPU rental prices as reference data.24 OneChronos, a New York firm, is awaiting approval for a marketplace of its own. Furthermore, BlackRock chief executive Larry Fink told the Milken conference in May that he believes buying futures of compute will become a new asset class.25

Capital providers, lenders, and infrastructure operators need these prices because they underwrite long-lived assets against uncertain future rental income. The four largest hyperscalers alone have guided to more than $700 billion in combined capital expenditure for 2026, financed through a widening mix of debt, tenant prepayments, and partnership capital.26 Model providers, meanwhile, manage their own exposure through the one lever fully under their control: the price they charge. The OpenAI cut, the Anthropic freeze, and the Google introductory schedule are hedges of a kind, each calibrated to a view of future serving cost that the buyer never sees.

Consequently, every constituency upstream of the enterprise, including suppliers, financiers, exchanges, operators, and model providers, is building tools to price and transfer compute risk. The enterprise, meanwhile, remains the one participant whose exposure arrives indirectly through vendor pricing decisions, without a benchmark or an instrument of its own.

Fig. 3 — The compute derivatives calendar, May through November 2026: exchange announcements, the Nvidia financing platforms, the CFTC request for comment, the review extension, and the dates that remain open.

Put the Compute Assumption in Writing

Although most enterprises should never trade a compute future, every enterprise should understand how its AI spending depends on the price those futures will eventually discover. Managing that dependency calls for budgeting and contract discipline, and five practices carry most of the load:

Map the exposure chain. For each material AI commitment, identify the billing unit, the party holding unit-price risk, the party holding volume risk, and the event that triggers repricing. The table above can serve as a working template. Most organizations will find that their exposure is larger than the licensing line suggests because agentic workloads convert seat-based products into consumption-based ones through allotments and overages.

Budget introductory pricing at its standard rate. A dated promotional price is a scheduled liability, and the finance model should record it as one. Google’s published January 1 step is the clearest current example, but the same discipline applies to any promotional credit pool or launch discount. Where a vendor cancels a scheduled increase, as Anthropic did, the budget can recover the difference, which is a far better outcome than discovering the increase in the first invoice of a new fiscal year.

Negotiate the mechanics of price changes. Contract language on notice periods for rate changes, caps on unit-price increases at renewal, and the right to migrate to a successor model at an equal or lower effective cost per task will outlast any single rate card. Effective cost can also diverge from the posted rate, since a newer tokenizer, a more verbose reasoning mode, or a longer default context can raise the cost of completing a task even when the posted rate falls.

Decline index-linked pricing until the benchmark earns trust. Once a public compute price exists, some vendors will propose pricing that floats with it. That arrangement may eventually make sense, but the regulator’s own questions about fragmentation, manipulability, and the gap between neocloud spot rates and the capacity that serves enterprise workloads should settle before any buyer accepts an index as the basis for its bill.

Treat the regulatory calendar as market intelligence. The CFTC comment file closes October 20, and submissions will become public. Whether hyperscalers, neoclouds, and model providers file comments, and what they disclose about how much of their business runs through bilateral contracts, will reveal more about the true shape of compute pricing than any vendor briefing. Finally, when contracts do begin trading, the slope of the forward curve belongs in the scenario section of the AI budget alongside consumption growth.

The Meter Behind the Meter

Every enterprise AI invoice is a meter reading taken from another meter that the buyer has never been shown. The first attempt to publish that second reading has stalled for now because the people charged with protecting market integrity are not yet convinced that the number would be trustworthy. Their caution also confirms the underlying point that the price of compute is volatile, largely private, and set by parties whose incentives differ from the buyer’s, and that it already reaches the enterprise through every token, credit, and renewal. Organizations that write their compute assumptions into budgets and contracts now will be ready to use the forward curve when it arrives, while the others will learn its lessons from their vendors’ next price change.

References

  1. CME Group, “CME Group and Silicon Data to Launch Compute Futures on October 5 to Unlock New Way to Hedge AI Risks,” press release, August 11, 2026.
  2. “CME’s Nvidia GPU Futures Hit a Regulatory Speed Bump,” Finimize, September 22, 2026, reporting a CFTC letter of September 21, 2026, first reported by The Information. Launch timing remains subject to the extended review.
  3. Commodity Futures Trading Commission, “Request for Comment on the Listing of Compute Derivatives Contracts,” RIN 3038-AF77, Federal Register 91, no. 161 (August 21, 2026): 54259–54264.
  4. Silicon Data, “H100 Rental Price Index,” product and methodology page, accessed September 2026.
  5. CME Group, “Compute Futures,” product page, accessed September 2026.
  6. SemiAnalysis, “The Great GPU Shortage – Rental Capacity – Launching our H100 1 Year Rental Price Index,” SemiAnalysis, April 6, 2026.
  7. Silicon Data, “H200 vs H100 Rental Prices: The Premium That Doubled,” Silicon Data Blog, August 4, 2026. Values are for the Neo-Cloud tier; hyperscaler rates are published separately.
  8. “Bank of America Sends Blunt Message to Nvidia Stock Investors,” TheStreet, August 12, 2026, summarizing a Bank of America research note dated August 7, 2026. Figures are analyst estimates of spot rental rates.
  9. Dave Friedman, “Compute Futures Need More Than a Useful Price Index,” Substack, September 22, 2026. The manipulation scenario is presented by the author as a hypothetical and does not assess Silicon Data’s specific methodology.
  10. “Great Question. Let’s Break It Down: CFTC Considering Compute Futures,” Willkie Compliance Concourse, August 2026, quoting CFTC Chairman Michael S. Selig.
  11. OpenAI, “Advancing the Price-Performance Frontier with GPT-5.6,” July 30, 2026; “OpenAI Cuts Prices for Two of Its GPT-5.6 AI Models as Companies Grow Sensitive to Costs,” CNBC, July 30, 2026.
  12. Anthropic, “Introducing Claude Sonnet 5,” June 30, 2026, updated August 10, 2026, and Claude API pricing documentation as reported by Neomanex, September 14, 2026.
  13. Google Cloud, “Generative AI Pricing,” Gemini Enterprise Agent Platform documentation, accessed September 2026.
  14. Matthias Bastian, “OpenAI Adds $111 Billion to Its Cash Burn Forecast as AI Costs Spiral beyond Projections,” The Decoder, February 21, 2026, reporting internal financial documents first reported by The Information. Margin figures are company-reported to investors and are not audited.
  15. Sri Muppidi, “Anthropic Lowers Gross Margin Projection as Revenue Skyrockets,” The Information, January 2026. Figures are internal projections reported by The Information and are not audited.
  16. GitHub, “GitHub Copilot Is Moving to Usage-Based Billing,” The GitHub Blog, April 27, 2026.
  17. Michael Truell, “Clarifying Our Pricing,” Cursor Blog, July 4, 2025.
  18. Gartner, “Gartner Predicts AI Coding Costs Will Surpass Average Developer’s Salary by 2028 as Token Consumption Surges,” press release, June 24, 2026.
  19. Gartner, “Gartner Forecasts Worldwide AI Platforms and Models Market to Grow 63% in 2026,” press release, July 20, 2026.
  20. NVIDIA, “NVIDIA Partners With Apollo, BlackRock, Blackstone, Brookfield, Goldman Sachs and KKR to Establish AI Compute Infrastructure Financing Platforms to Mobilize Over $500 Billion of Third-Party Capital,” press release, August 10, 2026.
  21. NVIDIA Corporation, Form 10-Q for the quarter ended July 26, 2026, filed with the U.S. Securities and Exchange Commission.
  22. “Nvidia, Wall Street Partner on $500B AI Financing,” Axios, August 10, 2026.
  23. Intercontinental Exchange, “ICE and Ornn to Launch GPU Compute Futures Contracts,” press release, May 19, 2026.
  24. “CFTC Plans AI Compute Derivatives Consultation as CME Targets October 5,” Finance Magnates via TradingView, August 20, 2026.
  25. “The Push to Create a Futures Market for AI Compute,” Axios, August 12, 2026.
  26. “2026 Hyperscaler Capex Tops US$700bn – Analysis,” TMT Finance, August 18, 2026, drawing on second-quarter 2026 earnings call transcripts for Alphabet, Amazon, Meta, and Microsoft. Figures reflect company guidance, not reported full-year results.