Something unusual is happening in American software. Raffi Krikorian, the chief technology officer at Mozilla, switched much of his daily work to Kimi K3, a model from Chinese startup Moonshot, within days of its launch. Coinbase says it is moving to Chinese models to trim costs. Over the past month, the five most popular models on OpenRouter, a platform that routes traffic across providers, were all Chinese. The story of the year in AI is no longer just about who builds the smartest model. It is about who can make advanced intelligence cheap enough to use everywhere.
The numbers behind the shift are striking. On GPQA Diamond, a demanding science-reasoning benchmark, the leaders remain American: Claude Mythos 5, Gemini 3.1 Pro and Claude Fable 5 all cluster around 94 percent. But Alibaba's Qwen and the open-weight models GLM-5.2 and DeepSeek V4 Pro sit only a few points behind, at 90 to 92 percent. The gap in capability has narrowed to a rounding error. The gap in price has not. One analysis of 33 models found the strongest Chinese systems costing up to 100 times less to run, depending on the workload.
That combination has split the market in two. At the top, complex agentic work such as coding and long multi-step tasks still flows to a handful of frontier labs, because a small error rate compounds badly over 50 steps; Anthropic alone is said to account for around 40 percent of enterprise API spending. But the vast commodity layer beneath it, the classification, support and routine generation, has already tipped. On OpenRouter, the American share of traffic fell from roughly 74 percent to 20 percent over the past year while Chinese models climbed toward half. Tokens flow to the cheap models; dollars still flow to the best ones.
The twist is who benefits. For a year, software companies feared AI would make them obsolete. Instead, cheap open-weight models may be the best thing to happen to them. Because many can be downloaded and run in-house, firms can reserve premium American models for the hardest jobs and route everything else to cheaper alternatives. Vercel says open-weight models rose from 4 percent of its gateway traffic in January to 55 percent in July. Analysts at William Blair call the trend "unambiguously good" for software margins, and enterprise software stocks battered in the first half of 2026 have rebounded sharply.
None of this sits comfortably in Washington. A White House official last week accused Moonshot of covertly building K3 off Anthropic's Fable through "distillation," the practice of training a model to mimic a stronger one. Treasury Secretary Scott Bessent threatened sanctions. Then Beijing inverted the charge, accusing American firms of distilling Chinese models during their own training. The mutual finger-pointing is telling: distillation is standard across the industry, and the sudden concern for intellectual property from companies that built their systems on scraped public data has not gone unnoticed.
There is a catch buyers should weigh. Chinese open-weight models arrive with lighter guardrails, shifting responsibility for misuse onto whoever deploys them, and they carry political constraints trained deep into their weights: researchers found DeepSeek's answers on China-related controversies stayed pro-Beijing even after jailbreak attempts. Self-hosting buys data control, not neutrality. But for most business tasks, "good enough" has already won. The real risk to American labs is not that China seizes the top benchmark spot. It is that the market decides the top spot is no longer worth the premium.