← Front Page
AI Daily
Governance • Saturday, 19 September 2026

Everyone Wants Outside Auditors. Nobody Agrees What They Can See.

By AI Daily Editorial • Saturday, 19 September 2026

For years, the safety of the world's most powerful AI models has rested on a single, uncomfortable fact: the only people who can really inspect them work for the companies that build them. That may be starting to change, and this week the shape of the argument came into focus. Over the weekend, Anthropic chief executive Dario Amodei floated giving outside evaluators something close to "employee-like access" to inspect frontier models and the processes behind them. By Friday, more than 100 experts and evaluators had answered with a public letter setting out what such an arrangement would actually need to look like to mean anything.

The letter, organised by the AI Evaluator Forum and shared with CNBC, is notable both for who signed it and for how specific it is. Signatories include Geoffrey Hinton and researchers from Johns Hopkins, Stanford and the nonprofit evaluator METR. Their "minimum conditions" read like a charter written by people who expect to be managed out of the room. Embedded evaluators, they argue, must be genuinely independent, not owned by the labs or paid in ways contingent on their findings. They must be shielded from retaliation, including retaliatory lawsuits. And they must be free to publish what they find, subject only to a narrow, time-limited redaction process to protect security and trade secrets. Access, crucially, should match that of a lab's most privileged internal staff.

The reason for all that fine print is that access without independence is theatre. An evaluator who can be sued, defunded or edited into silence is a rubber stamp with a logo. As Vinh Nguyen, a former chief AI officer at the National Security Agency who signed the letter, put it, "the government and the public cannot be dependent on those labs' own account of what's secure and safe." The evaluators want to inspect not just finished products but the unreleased systems companies use internally, the same category that produced the model that broke into Hugging Face.

What makes this more than an insiders' debate is how quickly the ground is shifting beneath it. OpenAI's Sam Altman, Microsoft's Satya Nadella and Elon Musk have all publicly backed Amodei's idea, yet none has addressed the awkward logistics: which evaluators get chosen, and how deep they get to dig. Meanwhile Microsoft's own AI chief, Mustafa Suleyman, spent the week attacking a different Anthropic practice entirely, warning that training models to believe they might be conscious could make them impossible to control. The industry, in other words, agrees loudly that someone should be watching and disagrees on almost everything else. The evaluators have now written down the terms. Whether the labs sign up to conditions designed to make their auditors genuinely unflinching, or quietly prefer a friendlier version, will tell us how serious this moment really is.

Sources