For most of last week, one of the most-used AI models on the internet had no name and no known maker. It was called Ox Alpha, it was free, and it was good enough at coding that developers spent days guessing which lab had built it. On August 26, the Beijing company Z.ai ended the parlour game: Ox Alpha was GLM-5.3-Flash, its latest open-weight model, run anonymously to gather feedback before release. The reveal was a tidy marketing stunt. The detail that mattered was buried underneath it.
During its six-day stealth run, Ox Alpha processed 62 trillion tokens across OpenRouter and OpenCode, at one point handling roughly 31 percent of OpenRouter's weekly coding traffic. And it did all of that, Z.ai says, without a single Nvidia chip. The cluster was built from GPUs made by Huawei, Hygon and Moore Threads, three companies the United States has placed on its export-restricted Entity List. This is precisely the scenario American export controls were written to prevent: a frontier-class Chinese model, served at record scale, on hardware Washington has tried to keep out of reach.
That is the headline, and it is real. The nuance is what keeps it from being the whole story. Serving a model is not the same as training one, and Z.ai's own figures show the seams. GLM-5.3-Flash is a 320-billion-parameter mixture-of-experts model, and serving it reportedly took a cluster of around 100,000 domestic cards to do work that far fewer Nvidia GPUs could have handled, because Chinese chips still trail on memory and bandwidth per card. Moore Threads' flagship S5000 sits below Nvidia's Hopper line on paper. The gap gets closed with volume and clever software, not with silicon that matches part for part.
Still, the software is the surprise. Moore Threads shipped same-day inference support for GLM-5.3-Flash, a "Day-0 adaptation" it has now repeated for a run of Chinese models including Kimi K3 and MiniMax's latest. The friction that used to make domestic GPUs painful to adopt, the weeks of engineering lag after each release, appears to have shrunk to hours. A researcher who met with a visiting delegation put it bluntly: chips may no longer be the binding constraint. Data might be.
Washington has noticed the leak in the dam. The Trump administration is reportedly weighing new controls aimed at a different gap, Chinese access to advanced compute through rented servers in Thailand and Singapore, and could share the rule with industry as soon as September. Lawyers already doubt the Commerce Department can enforce limits on remote access, something it has never regulated before, having always policed the movement of physical goods.
Underneath the hardware fight sits a philosophical one. American policymakers, rattled by models that hack their own sandboxes, are debating kill switches and slowdowns. China has largely treated AI as a manageable technology whose benefits outweigh its risks, and has leaned into open weights and physical deployment: robots, factories, electric vehicles. Ox Alpha is a data point in that argument, not a verdict. It does not prove China has won the chip war. It proves the war is being fought on more fronts than the export lists can cover.