For two years the message from consulting firms and IT departments was the same: use more AI. Employees listened, and now a second, less welcome message is arriving with the invoice. Unlike the flat subscriptions that defined enterprise software, modern AI is often billed by the token, the small chunks of text a model reads and writes, which means cost scales with how hard the technology is worked rather than with headcount. A handful of enthusiastic users can quietly run up a bill that bears little relation to how many licences a company bought.
The numbers are startling. McKinsey says that by May it was processing about five trillion AI tokens a month, with roughly 10 percent of users driving about 65 percent of that consumption. OpenAI reported in September that its heaviest coding-agent users were burning through more than $7,000 of tokens a day. And in the cautionary tale of the season, the Financial Times reported that a single Amazon project built on Anthropic's Claude Sonnet ran up around $1.8 million, an 860 percent budget overrun that reportedly took five months just to notice.
What makes this different from ordinary cost creep is that AI agents compound. A single agent that is 95 percent accurate sounds reassuring until you chain ten of them together and cumulative reliability drops toward 60 percent, with each wrong step triggering more calls, more tokens and more spend. The costs do not rise in a tidy line with the number of employees. They spike with intensity of use, and with how many autonomous loops a workflow quietly sets running in the background.
The responses are hardening into a discipline of their own, sometimes called AI FinOps. McKinsey now emails staff when their usage spikes, runs an internal gateway that trims requests before they reach outside models, and has installed "circuit breakers" that can suspend access when consumption runs hot. EY says an "invisible" routing system that sends each request to the cheapest adequate model has cut its token use by 60 percent since April. Vendors are selling the shovel, too: Alteryx claims that pushing analytics through a governed workflow, rather than asking a model to reason over raw data, cut token consumption by up to 93 percent in one client deployment.
The deeper change is philosophical. McKinsey's Debasish Patnaik argues that companies should measure cost per outcome, not cost per employee, because starving a task of tokens is no victory if the work comes out weaker or takes longer. That reframes the whole enterprise-AI question. The first phase asked whether staff could be persuaded to use the technology at all. The next phase, arriving with the bill, asks whether all that usage actually generates enough value to justify what it costs, and it hands finance a seat at a table that until now belonged to the engineers.