Earlier this week OpenAI published a trove of 722 papers, written largely by an unreleased AI model, claiming advances on 372 open problems spanning algebra, geometry and theoretical computer science. Some papers assert progress on famous targets such as the Riemann hypothesis and the Birch and Swinnerton-Dyer conjecture, each worth a million dollars as a Millennium Prize Problem. The company said its aim was to "enable further progress in mathematics." The mathematical community, by and large, has reacted as though something had been done to it rather than for it.
The sharpest line came from the Association for Human Mathematics, whose statement was reshared and effectively endorsed by Fields Medallist Terence Tao. "Releasing over 700 files at once is not a demonstration of scholarship," it read, "but a demonstration of power." That framing matters, because the complaint is not really that the machine is wrong. It is about who now has to do the unglamorous work of finding out.
Mathematics has always settled truth slowly, by expert referees checking each step before a result enters the literature. A long paper can take years to vet. OpenAI's release inverts that economy: it produces claimed proofs faster than anyone can read them, then leaves the verification bill with the humans. One senior researcher, writing in The Conversation, said a paper solving one of his favourite problems was so unintelligible that as a journal editor "it would have gone straight into the bin." OpenAI's own ChatGPT, asked to assess a headline result, called it "a serious hallucination" that should "never be cited, submitted, or circulated as a proof without a complete expert audit."
There is reason for caution beyond wounded pride. OpenAI has already retracted three papers over an elementary error and amended several others whose mistakes invalidated their results, which suggests thin vetting before publication. The company says it formalised 300 of the main results in Lean, a system that mechanically checks each logical step. But a Cambridge and King's College London paper found at least two discrepancies between the plain-language proof and the Lean code behind OpenAI's earlier Navier-Stokes claim, concluding that autoformalised proofs should not be trusted without the same peer review as any other. Machines checking machines is not yet the same as being sure.
It is hard to separate the mathematics from the moment. OpenAI is reportedly chasing a valuation near 1.4 trillion dollars, and research-level mathematics is exactly the kind of feat that makes talk of artificial general intelligence sound plausible to investors. A capability demonstration and a scientific contribution can look identical from the outside; the difference is whether the people who have to live with the result were consulted. Here they were not. The field's own Advisory Group on Mathematics and Artificial Intelligence had explicitly asked labs to stop testing advanced problems on proprietary models, and OpenAI cited the group for legitimacy while ignoring its central request.
Twenty-five Fields Medallists have now warned that mass-produced solutions could "destroy fertile ground instead of breathing life into new ideas." The worry is concrete: open problems are not just puzzles but the training ground where young mathematicians learn the craft. Carpet-bomb them with unreadable machine proofs and you may answer the questions while hollowing out the practice of asking them. Whether this week marks a genuine leap or an expensive act of theatre, the real test is the old one. The community will decide what these papers are worth, slowly, once the dust settles.