This website uses cookies

Read our Privacy policy and Terms of use for more information.

// AI Deep Dive

The $100B Question: Whose AI Stack Wins?

The model is the cheapest thing you'll ever switch. The contract, the compliance letter, and the SDK are the commitment.

// Executive Summary

Enterprises are about to lock in Q4 platform decisions during the loudest launch week of the year. In seven days, OpenAI shipped a new flagship, Google shipped its third Flash model in six weeks, Nvidia agreed to buy Hugging Face, and Anthropic's IPO came into view. The evidence says the model is the cheapest thing you'll ever switch. The real commitment is the cloud contract it bills through, the compliance letter it ships with, and the developer tooling your teams wire it into. Enterprise generative AI spend more than tripled to $37 billion last year. Hold that flat for a three-year contract and this is a $100-billion-plus decision.

  1. Spend follows the catalog, not the leaderboard. All three hyperscalers now let third-party frontier models burn down committed spend: Claude on AWS retires existing commitments, OpenAI models on Bedrock count toward AWS commitments, and Claude in Microsoft Foundry draws down MACC. With roughly $470 billion in upfront cloud commitments sitting on enterprise books, the model already inside your contract starts with a lead no benchmark closes.

  2. The compliance letter is not the same letter. Anthropic reaches FedRAMP High through Bedrock GovCloud, Google's Agent Platform, and Claude for Government. OpenAI's ChatGPT Enterprise and API are authorized at FedRAMP 20x Moderate. xAI's Grok for Government is still "in process". Three vendors, three different answers to "can we use it."

  3. Everyone is multi-model already, so the lock-in moved down a layer. a16z found 81% of enterprises run three or more model families, up from 68% a year earlier. The switching cost isn't the model. It's the agent framework and the protocol: MCP, now under the Linux Foundation, reports 97 million monthly SDK downloads.

Buy the stack the way you'd have bought a database in 2005: on the contract, the compliance letter, and the developer ecosystem. Treat this week's model launches as pricing news, not platform news.

// The Hook

The Loop I Paid For

I run a fleet of AI agents every day — coding agents, research agents, a couple of writers — and if you sat behind me for an afternoon you'd notice something the demos leave out. A lot of what gets sold as autonomy is brute force. The agent doesn't reason its way to done. It gets run again. And again. Until something passes.

Builders have a nickname for this. They call it the Ralph Wiggum loop, after the Simpsons kid who keeps cheerfully failing upward. You put the agent in a shell loop, hand it the same prompt every pass, and let it bang on the problem until the tests go green. Anthropic liked the idea enough to ship an official plugin for it. Thoughtworks put the Ralph loop on its Technology Radar this spring. And here's the honest part: it works more often than it has any right to.

Here's the other honest part. I pay for every lap. When a loop finally converges, I paid for every pass before it, including the ones that produced nothing. When it doesn't converge — and not every loop finishes — I paid for all of them and got a log file. Nobody sends you an invoice line that says "attempts." It shows up as tokens, and tokens show up as a number that's bigger than the demo implied.

So I've stopped asking which model is smartest. That's the leaderboard question, and I wrote about that last week. This week's question is the one that actually matters when you're staring at a Q4 budget: whose stack finishes the job, and who is charging you for the attempts? The answer has very little to do with the model and almost everything to do with what's wrapped around it.

// The Deep Dive

Deep Dive: Whose AI Stack Wins?

Five moves that decide the platform war, and why none of them is a benchmark.

The conventional wisdom in enterprise AI procurement has a comforting shape: pick the best model, wrap it in your own tooling, and switch when something better comes along. Prices fall, versions roll every few weeks, and every vendor swears its API is a drop-in replacement for the other guy's.

Then you look at what actually happened this week, and almost none of it is about the models. It's about who owns the catalog you buy from, who signs the letter your auditor reads, and whose SDK your developers have already stopped noticing they use.

The Field, as of This Week

Start with the labs, because every one of them made a platform move this week dressed up as a model move.

OpenAI shipped its new GPT flagship, which it calls Astra, on Thursday. It is priced at $10 per million input tokens and $50 per million output, carries a context window past a million tokens, and is rolling out across ChatGPT tiers, the API, Azure, and Amazon Bedrock, with enterprise workspaces switched off by default until an admin opts in. It's also the first OpenAI model to hit the company's top "Critical" tier for cybersecurity risk under its own framework. The enabling fact is five months old: Microsoft's exclusive license ended in April. Azure stays first in line, but ChatGPT's maker now sells through Amazon's catalog too.

Anthropic's news was financial and political rather than technical. The company filed confidentially to go public in June at a $965 billion valuation, and its annualized revenue passed $65 billion by the end of July, with a Wall Street debut expected as soon as this fall. In Massachusetts it broke with OpenAI to support a bill that would require large frontier developers to submit to independent catastrophic-risk evaluations every 120 days, with its state-policy lead arguing the industry shouldn't "grade its own homework". And on Tuesday it published an open-source blueprint for shopping and merchant agents with Visa, Mastercard, and Shopify attached. Read the fine print: the blueprint "leaves payment to you." Claude is selling the reasoning layer of commerce, not the wallet.

Google shipped its newest Flash model on Wednesday at $0.75 per million input tokens and $3.75 per million output tokens, an introductory price that doubles on January 1. That's the third Flash release in about six weeks, and it lands on a platform Google renamed this spring from Vertex AI to the Gemini Enterprise Agent Platform. The scale is the real story: Google Cloud posted $24.8 billion in quarterly revenue, up 82%, and Sundar Pichai told investors the cloud backlog stands at $514 billion, with nearly 90% of the Fortune 100 using Gemini Enterprise. Gemini is also the only frontier model family you can't buy on a rival cloud.

xAI finished the opposite move. Its newest Grok release went generally available on Bedrock on August 19, on Google's Agent Platform on August 21, and on Microsoft Foundry on August 26, putting it in all three hyperscaler catalogs inside a single week. It's priced at $2 per million input and $6 per million output.

Open weights had a strange week too. Meta, which went proprietary with its Muse models in the spring, released Muse Glimmer under Apache 2.0 in August. Mistral announced a European compute build-out targeting a gigawatt by 2030, plus regional inference endpoints. And Nvidia signed a definitive agreement to acquire Hugging Face for roughly $11.9 billion plus about $1 billion in retention equity, closing in the first half of 2027, with a public pledge that the hub "will remain an open platform for the entire AI ecosystem". The distribution point for open weights now belongs to the company that sells the chips they run on. Two datapoints frame the stakes: Menlo's enterprise survey put open source at 11% of enterprise LLM workloads, down from 19%, while Vercel's production traffic shows open-weight models running 29% of tokens on less than 4% of spend. Different measures, same lesson: open weights are the commodity layer, and somebody just bought the commodity exchange.

Integration Gravity: Why the Catalog Beats the Model

Here's the pattern underneath all of that. Every lab except Google spent the summer getting into the other clouds' catalogs, and Google spent it turning its own catalog into a moat.

The reason is a boring line in a procurement contract. Enterprises pre-buy cloud in bulk, and the analysts at Omdia estimate about $470 billion of those upfront commitments are sitting on books right now. When a model is listed in your cloud's marketplace, spending on it retires against the commitment you already signed. Google's own documentation says 100% of Marketplace spend can count toward minimum-commitment obligations. Microsoft's says Claude usage shows up as one Azure line item with MACC drawdown. From the CFO's chair, the model that's in the catalog is free money and the model that isn't is a new vendor.

I watched this exact movie in the 2000s. Oracle didn't win most of its database seats in bake-offs. It won them because it was already on the enterprise license agreement, and adding a schema cost nothing while adding a vendor cost procurement a quarter. Linux didn't win the data center when it won a benchmark. It won when it showed up on every hardware vendor's price list. Distribution beat quality both times, and quality caught up later.

The spend data says it's happening again. Menlo's tracking has Anthropic at 40% of enterprise LLM API spend, OpenAI at 27%, and Google at 21%, with OpenAI down from 50% two years earlier. Anthropic is the vendor with the broadest cloud footprint. That isn't the whole explanation, but it isn't a coincidence either.

There's a fair objection here: if Claude, Grok, and OpenAI's models are all on Bedrock, Foundry, and Google's platform, the catalog stops favoring any one lab. True. But look at who that leaves standing. When every frontier model is in every catalog, gravity stops pulling toward a lab and starts pulling toward your cloud. Over a three-year contract, that's a bigger commitment than any model choice, and nobody puts it on a slide.

The Compliance Moat: Who Signs the Letter

Regulated buyers don't ask which model is smartest. They ask whose name is on the authorization letter, and the answers diverge more than the marketing does.

Anthropic gets to FedRAMP High three ways: through Bedrock in AWS GovCloud, through Google's platform with Assured Workloads, and through its own Claude for Government. OpenAI's ChatGPT Enterprise and API platform are authorized at FedRAMP 20x Moderate, which is a real authorization but a different tier, with High available only by going through Azure Government. xAI's Grok for Government sits at "Agency Authorization In Process". Healthcare is the same story in miniature: Anthropic's BAA requires zero data retention, xAI's requires a signed BAA plus a zero-retention API, and Google's Cloud BAA covers the Agent Platform as a service.

That last phrase is the tell. Google's own FedRAMP guidance says individual models "aren't independently authorized". The authorization belongs to the platform. So the compliance moat, like the catalog, mostly belongs to the cloud, and the lab inherits it by being inside the fence. That's why a frontier model showing up in GovCloud is bigger news for a regulated buyer than the same model scoring two points higher on a coding benchmark.

The Massachusetts split fits the pattern. Anthropic saying yes to independent third-party evaluations every 120 days while OpenAI's policy lead argues that "inconsistency doesn't mean safer" is a policy disagreement on the surface. Underneath, it's positioning. A lab that volunteers for outside audits is telling banks, hospitals, and agencies that it plans to be the easy one to approve. Whether that's conviction or strategy, the buyer gets the same benefit.

Developer Lock-In: The SDK Is the Contract

The a16z survey of Global 2000 executives found 81% now run three or more model families in production. Multi-model is the default, which means the model was never the lock-in. The lock-in is whatever your developers built the agents in, because that's the code that doesn't get rewritten when the leaderboard changes.

LangGraph pulled about 63 million PyPI downloads in the last 30 days. OpenAI's Agents SDK pulled about 33 million. Meanwhile the Model Context Protocol moved to the Linux Foundation in December, co-founded by Anthropic, OpenAI, and Block, with 97 million monthly SDK downloads and roughly 10,000 servers. That donation is the most important lock-in story of the year, and it barely made the trades. It means the tool-connection layer is now vendor-neutral by charter, in the same way the Apache Software Foundation neutralized the web server. Anything you build on a vendor-specific SDK above that layer is a three-year decision. Anything you build on the protocol is portable.

This is also where my Ralph loop comes back. A retry loop is trivial to write in any harness. What differs is who bills for the laps. Per-token pricing means the vendor gets paid whether the loop converges or not, and the price of a wasted lap varies wildly: OpenAI's new flagship charges $50 per million output tokens, Google's new Flash charges $3.75, xAI's $6. A loop that doesn't finish costs about thirteen times more on one stack than another, and none of them refund the attempts. The stack that "wins" for a given workload is the one that finishes the job in the fewest paid laps, and you only learn that by metering your own runs.

Can a New Lab Still Win? Thinking Machines as the Test Case

Which brings us to the expensive question: with the catalogs, the compliance letters, and the SDKs mostly spoken for, can a new lab still become a platform?

Thinking Machines Lab is the cleanest test. Mira Murati's company raised a $2 billion seed round at a $12 billion valuation in July 2025, shipped a fine-tuning service called Tinker, and in July released Inkling, a 975-billion-parameter mixture-of-experts model, under Apache 2.0. This week TechCrunch reported it is in talks to raise about $1 billion at a $40 billion valuation, with Accel considering the lead. Two co-founders have already left for OpenAI. None of that is closed.

Look at what the company actually sells, though. It isn't selling a catalog listing. It isn't selling a FedRAMP letter. It's selling open weights and a training API, which is to say it's selling the thing enterprises fine-tune inside somebody else's stack. That's a layer, not a platform. The other 2026 test cases rhyme: Safe Superintelligence took a reported $5 billion from Nvidia for compute and still has no product. Reflection AI closed $2.5 billion at a $25 billion pre-money as an "open frontier lab." And the smaller shops mostly ended up as licensing-and-hiring deals with Google, or shut down and sent the staff to Chrome.

So yes, a new lab can still win a layer: open weights, fine-tuning, a specialty vertical. What it can't win in 2026 is the platform, because the platform is defined by three assets that take a decade and a hyperscaler to accumulate. The Nvidia purchase of Hugging Face tells you where the new labs' distribution will live: in the chip company's catalog, one rung below the clouds.

Four Ways Buyers Get This Wrong

  1. Buying the leaderboard. Ranking is a tiebreaker. If your 2027 platform choice is being made from a benchmark screenshot, you're choosing the most expensive variable and ignoring the three that compound.

  2. Reading the cloud's compliance letter as the model's. Google says it plainly: the model isn't independently authorized, the platform is. If your auditor asks who signed, the answer had better be a name you can call.

  3. Standardizing on a vendor SDK instead of a protocol. The SDK is a three-year commitment disguised as a convenience. Build to MCP and the open interfaces above it, and let the model underneath be the thing you swap.

  4. Pricing the demo instead of the loop. The demo converges on the first pass. Production converges later, or never. If you don't meter attempts per completed task, you're buying tokens and calling it outcomes.

// Key Takeaways
  1. Treat this week's launches as pricing news, not platform news. A new flagship at $50 per million output tokens and a new Flash at $3.75 change your unit economics. Neither changes which cloud your commitment sits in.

  2. Buy through the contract you already hold, and know what it costs to leave. With $470 billion in cloud commitments outstanding and every major model now retiring against those commitments, the cloud is the platform decision. The lab is a line item inside it.

  3. Ask who signs the compliance letter. Anthropic's three routes to FedRAMP High, OpenAI's Moderate authorization, and xAI's in-process status are three different procurement timelines. Know which one you're on before the pilot, not after.

  4. Standardize on protocols, meter the attempts. With 81% of enterprises already multi-model, the model is swappable and the framework isn't. Build on MCP, and track cost per finished task, not cost per token.

// What This Means for Your Planning

The planning question for Q4 isn't "which model should we standardize on." It's "which stack finishes our work at a price we can predict, inside a contract we can leave." Most boardrooms are still asking the first question because it's the one the vendors keep answering. The honest answer to the second one usually points at the cloud, not the lab.

Phase 1: Inventory the gravity (weeks 1–2). List every cloud commitment you hold, its remaining balance, and which frontier models are already inside it. Then pull a month of agent logs and count attempts per completed task, by model. That second number is the one nobody has, and it decides whether a cheaper model is actually cheaper.

Phase 2: Run the vendor-commitment test (weeks 3–6). Put every AI vendor on your shortlist through the same three questions, in writing. Score the answers, not the pitch.

Phase 3: Write the exit before you sign the entry (by quarter close). Wherever you commit, build to the protocol layer, keep your prompts, evals, and skills in your own repo, and put the cost of leaving in the contract review.

The three questions for every vendor in your Q4 budget:

  • What does it cost to leave? Not the contract penalty. The rewrite. If the answer is "swap the API key," you built well. If the answer involves the phrase "our agent framework," you're already locked in and just haven't been billed for it yet.

  • Which cloud does it live in? If the model isn't in your catalog, every dollar you spend on it is a dollar that doesn't retire your commitment.

  • Who signs the compliance letter? The lab, or the cloud it rides on? If your auditor can't tell the difference, your risk committee can't either.

Here's the question I'd bring to the budget meeting: for every AI dollar we plan to spend next year, how many of those dollars buy finished work, and how many buy attempts? Whoever can answer that is the one who should own the stack decision.

Your AI Sherpa,

Mark R. Hinkle
Founding Publisher, The AIE Network
Follow me on LinkedIn

If you want to get in contact or give me feedback, reply to this email. I read every single one of them.