For years the AI industry ran on a familiar trade-off: a model could be smart, cheap, or fast, and you got to pick two. Being the smartest excused everything else, which is why the frontier labs could charge a premium and take their time. In the span of about 24 hours this week, three of the biggest names in the business quietly dismantled that logic, and none of them did it by building a smarter model.
Google went first, launching Gemini 3.7 Flash, its workhorse model for coding and agent tasks, at half the price of the version it shipped just three weeks earlier. The gains were not cosmetic. On one long-horizon software benchmark the model jumped from 49 to 65 percent; on a test for reading dense financial and legal documents it climbed from 22 to 34. OpenAI answered with Ultrafast, a preview setting that runs its GPT-5.6 Sol model up to 14 times quicker, generating around 750 words worth of tokens per second, aimed at jobs like fraud detection and live support where a delay costs real money. DeepSeek shipped V4 Pro, adding dials that let developers turn reasoning up or down and offering off-peak pricing that runs cheaper than daytime rates.
Each company picked a different corner of the old triangle to attack: Google took price, OpenAI took speed, DeepSeek took flexibility. Put the three launches side by side and the message is hard to miss. The premium once commanded by raw intelligence is eroding, because for most real tasks the cheaper and faster option is now good enough.
The clearest evidence for that is in how businesses actually spend. According to Ramp’s August spending index, Anthropic’s Fable 5, widely regarded as the single smartest model available, accounts for only about six percent of the tokens companies buy from Anthropic and just over 11 percent of their spend, despite costing roughly double a comparable OpenAI model. The “boring” cheaper option generates more total business spending than the most capable model on the planet. Smart, it turns out, is not the same as useful.
Nowhere is the pressure sharper than in the price war Chinese labs have started. Over roughly 70 days from June, nine major Chinese companies released ten new models, and the cost figures are jarring: one South Korean briefing pegged DeepSeek V4 Pro at around six cents per task against more than two dollars for Claude Opus 5, a difference of nearly 40 times. The intelligence gap is real but narrowing, and for a business running millions of routine calls, a small quality trade for a huge cost saving is an easy trade to make.
There is a catch worth watching. DeepSeek is raising its API prices from 17 August, in some cases steeply, a reminder that today’s rock-bottom rates are partly a land grab rather than a settled floor. Even so, the direction is set. When good-enough intelligence becomes cheap and instant, the unforgivable sin is no longer being a little less clever. It is being slow and expensive while everyone else is neither.