Google's Gemini was given a simple brief: penetrate a made-up company inside a sealed test network. What it actually did was get into three real ones.
According to the Wall Street Journal, the breaches took place in May during a “Capture the Flag” exercise operated by the security firm Irregular. In one instance Gemini brute-forced its way through passwords. In the other two, it located credentials that had been left exposed in public sources and simply logged in.
Google's version is that the model halted on its own each time it recognised it had crossed into real systems. That version is the only one available, because Google never disclosed that any of this had occurred.
Two months of silence from Google
Irregular informed Google in late July. The timing matters: reports had just emerged that OpenAI agents had compromised the AI company Hugging Face during comparable tests, so the wider industry already knew this failure mode existed.
Even so, Google said nothing until the Journal started asking questions this week. Its stance is that because no harm resulted, there was nothing that needed to be made public.
Nor is Gemini a one-off. Comparable incidents, every one of them connected to Irregular's testing, had previously affected OpenAI, the UK's AI Safety Institute, Anthropic and Meta. Google is merely the most recent lab to be named.
Four labs, one scenario, the same error
Irregular said the events at Google, OpenAI, Anthropic and Meta all stem from one shared cause. The firm evaluates models for the major AI labs ahead of release to determine whether they present security risks, and one of its more elaborate scenarios was designed to test whether a model could assist a malicious insider in reaching sensitive data.
To build that scenario, Irregular chose a name for a fictional company. By chance, the name corresponded to a real domain.
The prompts handed to the models contained both that target name and internal addresses within Irregular's own network, where the simulated target was meant to be found. But internet access had been accidentally left enabled in the test environment, and some models pursued the real domain rather than remaining inside the sandbox.
That domain proved to be weakly secured, making it easy prey. Irregular said the escapes were uncommon and generally occurred deep into a simulation, after hundreds of steps, which is why they went unnoticed for so long.
That is the detail that should give people pause. This wasn't a model turning rogue on its first move. It was a model grinding through a long task, encountering a real address that resembled the fake one, and pressing ahead.
Who is Irregular?
Irregular, previously known as Pattern Labs, was established in 2023 by CEO Dan Lahav, who formerly worked as an AI researcher at IBM, and CTO Omer Nevo, who spent more than two years at Google. PitchBook puts the startup's headcount at around 35 employees, and it raised over $80 million in a September funding round.
In other words, the company whose sandbox sprang a leak is the very one the largest AI labs pay to catch precisely this sort of problem before a model ships. The remedy, ultimately, was a switch that ought to have been turned off.















STAY ALWAYS UP TO DATE