On paper it reads like a company giving ground. Nvidia, which sells the most sought-after AI chips on earth, is now helping a competitor's accelerator slot directly into its own racks. The competitor is d-Matrix, a Microsoft-backed inference-chip startup valued at around $2 billion, and its Raptor processors will plug into Nvidia-powered systems using a technology called NVLink Fusion, with the first rack-scale products due in 2027. Read quickly, it looks like Nvidia inviting a rival to eat its lunch. Read carefully, it is closer to a toll operator waving more cars onto its road.
The key is understanding what NVLink actually is. It is the high-speed fabric that lets a rack full of chips behave like a single computer, sharing memory and passing data at enormous bandwidth. For an inference chip like Raptor, that connection matters more than the silicon itself, because inference accelerators live or die on how fast they can reach a model's weights and hand tokens back to the host. By opening the fabric to outsiders, Nvidia turns a proprietary bus into something closer to an industry standard. d-Matrix is not the first through the door: Marvell joined earlier this year, and Groq is being integrated at the rack level too.
The economics explain why Nvidia would want this. Even when a partner supplies the main accelerator, Nvidia keeps charging for the networking, the CPUs, the NVLink switches, the storage and the CUDA software wrapped around all of it. Data-center networking alone brought in more than $40 billion last quarter, up 138 percent from a year earlier. Nvidia now measures revenue per gigawatt of data center at roughly $18 billion for its older Hopper generation, $25 billion for Blackwell and $40 billion for the coming Vera Rubin, precisely because the "AI factory" it sells is now far more than a pile of GPUs. As Jensen Huang put it, "compute is revenue."
The same logic is pulling in the industry's most self-reliant company. The Information reports that Apple, building an enterprise AI-inference server around its in-house M8 Ultra chip for a possible 2029 launch, is weighing NVLink Fusion to connect those chips. Apple designs its own silicon and rarely leans on anyone; that its engineers reportedly see Nvidia's interconnect as the best option available says a lot about how durable the fabric has become. Accelerators change hands quickly. Interconnect standards, once hyperscalers build their racks around them, tend to stay put.
There is a real risk buried in the strategy, and it is worth naming. Inference is exactly where custom chips compete best on cost and power, and it is where Nvidia's fattest margins sit. If Raptor and its kind prove several times more efficient per token, hyperscalers could shift their highest-volume work off Nvidia GPUs while keeping the fabric, leaving Nvidia collecting a smaller toll. Management already expects gross margins to bottom around 71 to 72 percent this quarter. But that is the bet Nvidia is making: better to own the road every rival must drive on than to try to win every car. The concession is the moat.