Three very different reports this week, from a hiring-data firm, two large incumbents, and the builder of a new benchmark, all arrive at the same unglamorous conclusion. The breakthrough model everyone obsessed over for three years has quietly become a commodity. The money, and the difficulty, now sit in everything wrapped around it: the data, the plumbing, the governance, and the judgment to know when not to use AI at all.
Start with the job market, which tends to tell the truth before the press releases do. Analysis by the hiring-data firm Draup found AI-related postings at banks like JPMorgan, Citigroup and Capital One up 49 percent this year, to nearly 140,000 listings. The single fastest-growing skill was "agent orchestration," the art of making multiple AI agents work together on a task, up a startling 1,721 percent. Draup's chief executive called it "the hottest skill on Wall Street." Notably, the demand is not mainly for people who build models. It is for "forward-deployed engineers" who can embed AI into a trading desk or a back office, and for a fast-growing cohort in governance and risk: mentions of "responsible AI" rose 657 percent, and governance-related skills now outnumber references to actually training and running models by nearly two to one.
The incumbents tell the same story in a more cautious register. At a Fortune summit, Bank of America's technology chief Hari Gopalkrishnan said one of the biggest mistakes he sees is rushing to AI "when deterministic models do a plenty good job." Every project at the bank passes through a review covering 16 "pillars" of risk, and plenty of ideas are rejected in favour of a simple app or a decision rule. The bank is spending heavily anyway, around $400 million on roughly 140 uses for about $800 million in benefit, and doubling that budget next year, but it routes work through an orchestration layer that sends easy tasks to cheap open-weight models and only the hard reasoning to expensive ones. S&P Global's Sally Moore put the underlying lesson in four words: "data is the currency within AI." She described helping one bank lift the accuracy of an AI system from about 60 percent to 98 percent, and the gain came from better data and context, not a cleverer model.
That claim gets its sharpest test in the third report. Writing in CIO, the builders of a new "Enterprise-Bench" held the model constant and changed only the system around it. Their structured-memory setup answered 94.3 percent of tasks correctly; a general coding agent using the same underlying model managed 63.6 percent, while burning far more tokens per correct answer. The authors' point is that enterprises keep misdiagnosing an integration failure as a model failure. A question like "which customers are affected by this bug, and what is the revenue exposure?" is not hard because the model cannot reason. It is hard because answering it means stitching together a support ticket, a product component, an engineering issue and a CRM record, each named differently and governed by different permissions. Buying a smarter model does not repair a missing relationship between two databases.
There is a tidy through-line here, and a warning inside it. The benchmark authors argue that autonomy should be "earned" rather than switched on, with agents proving they can reliably retrieve and reconcile facts before they are trusted to act. Bank of America is making exactly that bet in practice, keeping humans in the loop and expanding only as its controls mature. For all the noise about ever-larger models, the enterprises actually putting AI into production are converging on a less exciting but more durable idea: the frontier model is now the cheap, interchangeable part. Owning your data, your context and your guardrails is the thing that lasts.