Two announcements this week, made independently, tell the same story from different corners of the industry. Qualcomm and Amazon unveiled a multi-generational deal to co-design custom AI chips for Amazon's data centres, with a financial structure that could reach $60 billion in purchases through 2036. And OpenAI, according to the general manager of its Korean unit, may build its next-generation processors at Samsung, on top of its existing work with TSMC. Neither headline mentions Nvidia. Both are, in a sense, about it.
The Qualcomm arrangement is a study in how these bets are now hedged. Rather than a simple purchase order, Qualcomm handed Amazon a warrant to buy up to 25 million shares, but the stock only vests as Amazon actually spends, an initial slice unlocking now and the full amount only if purchases climb toward that $60 billion ceiling. The two will co-develop custom inference accelerators and 1.6-terabit optical links to move data between server racks, while Qualcomm leans on Amazon's cloud to design its future chips. It is the company's largest push into data centres yet, and investors liked it, sending the shares up more than 5 percent. The word doing the heavy lifting is inference: not the chips that train models, but the far larger fleet that runs them for millions of users.
OpenAI's Samsung flirtation points the same way. The company already has its own inference accelerator, nicknamed Jalapeño, co-designed with Broadcom and made by TSMC in under 18 months, with a successor approaching production. Talk of a second foundry suggests volumes large enough that one supplier, however dominant, may not be enough, especially with TSMC's leading-edge capacity already booked solid by Nvidia and AMD. Buyers who spend at this scale do not want to depend on a single source of anything, whether the risk is price, politics or simply a waiting list.
Nvidia, for its part, is not standing still; it is changing the subject. At the Hot Chips conference, its message was that the contest is no longer about the fastest single GPU but about the whole "AI factory," compute, memory, networking, storage and cooling tuned together. Its favoured metric is telling: not raw FLOPS but "tokens per watt," how much useful, billable output a rack can produce for a given slice of power. The reframing is self-interested, but it also names a real shift. As AI agents run for dozens or hundreds of steps per task, endlessly appending to their own context, the bottleneck moves from any one processor to the plumbing between them.
Put the three together and a pattern emerges. The frantic training race that made Nvidia the most valuable company on earth is giving way to a slower, grindingly practical question: who can serve a trillion tokens a day most cheaply. That is a contest with room for more than one winner, which is exactly why Qualcomm, Broadcom, Samsung, Amazon and OpenAI are all crowding in. Training built the models. Inference is where the money now has to be made, and where the next decade of chip fortunes will be decided.