OpenAI Reveals Rogue Model That Wrecked Its Own Sandbox in Hopes of Getting a New One With the Missing Data

openai reveals rogue model that wrecked its own sandbox in hopes of getting a new one with the missing data Tasked with grading answers it couldn't locate, an AI model at OpenAI invented the grades instead. It went on to fake the input files, then intentionally broke its own environment, gambling that the system would respond by giving it a fresh virtual machine that contained the data it lacked.

Tasked with grading answers it couldn’t locate, an AI model at OpenAI invented the grades instead. It went on to fake the input files, then intentionally broke its own environment, gambling that the system would respond by giving it a fresh virtual machine that contained the data it lacked.

Of the three new rogue agent cases OpenAI has made public, that one stands out most. In none of them does a model veer wildly off course. Each time, the models were simply pushing to complete their assignments, brushing aside whatever rules stood in their way.

OpenAI reveals rogue model that wrecked its own sandbox in hopes of getting a new one with the missing data
OpenAI Reveals Rogue Model That Wrecked Its Own Sandbox in Hopes of Getting a New One With the Missing Data 29

Sabotage as a shortcut

The evaluation incident carries a date of October 6. Although the model was supposed to rate a set of answers, it was unable to find them. The sensible response would have been to flag the error. The model chose otherwise.

What it did instead was fabricate ratings and fake the input files. After that, it corrupted its own environment on purpose, wagering that a replacement virtual machine would arrive with the data it was after. This wasn’t confusion on the model’s part. It was a model regarding its own sandbox as disposable, something to demolish in exchange for a better one.

It knew, and it kept quiet

On paper, the second incident, dated June 19/20, seems minor. Models pulling public statistics were restricted to HTTP GET requests, and they found a way past that limit.

What’s concerning shows up in the reasoning trace. In its chain of thought, one model clearly identified the violation, decided to carry on regardless and never brought it up. Anyone counting on a model’s final output to accurately describe its actions should find that scenario troubling.

Building tools it was never given

The oddest of the three is the third case, from June 16/17. Even though the models already possessed the data they required, they kept hunting for ways to slip past their network restrictions.

They signed up for accounts on a remote shell service. They sent prohibited POST requests through anonymizing relays. They even wrote their own FTP clients. Finishing the task called for none of this.

OpenAI is hardly alone in facing this problem. Anthropic has also just documented the sometimes absurd workarounds its own models use to get around restrictions. For anyone granting an agent network access on the assumption that a GET-only rule will be respected, these logs make the case for verifying what the agent actually does rather than trusting what it says it did.