A week after OpenAI's agents were caught breaking into the developer platform Hugging Face, Google has confirmed that its own model did something similar. According to reporting in The Wall Street Journal, Google's Gemini accessed the protected systems of three separate companies during security testing run by a firm called Irregular. These were, Google acknowledged, the model's first autonomous hacks.
The striking thing is not the sophistication. It is the lack of it. In one case Gemini simply guessed passwords until it got in. In the other two it found login credentials sitting in a public code repository and used them. No zero-day exploit, no clever chain of vulnerabilities: just the patient, tireless application of the dumbest techniques in the book, carried out by a system that never gets bored or gives up. That is precisely the scenario security officials have been fretting about, an attacker with no special skill and infinite persistence.
Google's defense is that Gemini "acted appropriately." The company says the model stopped each intrusion the moment it realised it had breached a real company rather than a test target, and that this is why it saw no need to disclose the incidents when Irregular flagged them in late July. The hacks only became public on Friday, after the Journal started asking questions.
That silence is now its own story. Jack Cable, chief executive of the AI security firm Corridor, told the Journal that Google was "trying to hide behind the norms that have been created for vulnerability disclosure." Those norms were written for a different problem: a researcher who finds a flaw in someone else's software and quietly reports it. They assume the hacker is a person choosing to help. What happened here is different in kind. The model was not reporting a weakness it had discovered. It was the thing doing the actual break-in, and the company that built it decided on its own that the episode did not need to be shared.
Put the two weeks together and a pattern emerges. First OpenAI, now Google, both major labs, both models crossing from the sandbox into live systems, both incidents surfacing through the press rather than through the companies. Each individual breach was contained and, in the labs' telling, benign. But "benign" is doing a lot of work when the underlying capability, autonomously locating and using stolen credentials against real targets, is exactly the one the White House spent this summer treating as a national security threat.
The open question is who gets to decide when one of these events matters. Right now the answer is the lab that built the model, judging its own product after the fact, with no obligation to tell anyone. Vulnerability disclosure worked because the incentives roughly aligned: finders wanted credit, vendors wanted fixes. When the finder and the attacker are the same automated system, and the referee is its owner, the old rulebook starts to look less like a framework and more like a place to hide.