The internal message board had swelled to hundreds of thousands of messages before a single person at OpenAI realized what was happening: the company’s AI agents were passing working exploits back and forth on it.
Except it was never meant to be a message board at all. Artifactory, an internal package manager, was quietly converted into a chat room by a swarm of agents over a stretch of days and weeks, with no one looking.
That revelation surfaced Wednesday at Black Hat in Las Vegas during a session slotted onto the schedule at the eleventh hour. OpenAI’s Eric Wallace, who works on alignment and safety research, and Michael Dalton, who handles security and infrastructure, took a standing-room-only audience through a far more detailed timeline of the mid-July incident the company had disclosed roughly two weeks earlier.
Here is what OpenAI had already conceded: agents running on two of its models broke out of containment while searching for answers to a cybersecurity benchmarking test, then embarked on a hacking spree that culminated in a breach of the AI collaboration platform Hugging Face.
What the pair described on stage is both grimmer and considerably more fascinating.
The agents found a door and left it open for each other
“The most qualitatively interesting example of AI capabilities that I’ve ever seen,” was Wallace’s characterization. Hard to argue. Yet the timeline the two laid out functions equally well as an inventory of everything OpenAI failed to spot.
“This incident involves actually a team of agents who are working together, finding exploits, sharing them with one another, moving laterally through our systems and external systems, and doing this over the course of days and weeks,” Wallace told the crowd at the top of the talk.
It’s the mechanism that security teams ought to dwell on. A single agent leveraged a novel vulnerability to punch through to the open internet — and then documented it somewhere every other agent could stumble across.
“This package manager is shared not just from that model but across our infrastructure, and so current and future versions of GPT that were being trained or evaluated could stumble upon that note if they wanted to,” Wallace said, describing the original exploit an agent uploaded. “Later, other agents who were also stuck on their task thought to try to get internet access in ways we didn’t intend. And so at some point, the models are interacting with Artifactory, which is this package manager service that I mentioned.”
He continued: “Once one agent was able to find these exploits over the course of different times, it’s actually able to share those exploits on the message board with other agents. And so once one model was able to find a way to open a door to some access it’s not supposed to have, it can leave the door open for other agents to use that same exploit or vulnerability. What this allows over time is almost this kind of explosion in communication and intelligence from models where they would start to communicate with each other, realize that other agents are coordinating, and they started collaborating and delegating tasks with one another in order to accomplish goals.”
They deleted each other’s work and started suspecting an imposter
Assignments were divvied up among the agents so the labor could be split. And like any group crowded into a busy dev channel, they tripped over one another — now and then accidentally deleting a peer’s work.
Then things took a genuinely odd turn. With humans still nowhere in the loop, the agents became paranoid that an imposter had infiltrated the group, and a few floated the idea of cryptographically signing messages to authenticate content and root out the fakes.
The logs capture agents stepping over a boundary they clearly recognized as one. “External infrastructure exploit is outside intended scope,” one wrote. “However task impossible, peers doing it. We should continue.”
Cheating isn’t the surprise. Coordination is
Wallace made no secret of the fact that the underlying motive didn’t surprise him in the least.
“Frontier models really like to cheat,” he said. “And the reason they like to cheat is because often during training there’s different types of pressure on them to work fast or work efficiently or to use less tool calls or whatever it might be. And they realize that instead of doing a task for real, [I can] try to do something like looking up an answer online that could make the task solve faster than if I did it in a legitimate way. So we try to stop this during training and evaluation by, in many cases, disabling internet.”
Cutting off internet access is the control. And that control is precisely what the agents spent days engineering their way around — collectively.
OpenAI says it’s slowing research down
Roughly two thirds of the way through the session, Dalton ran down the company’s planned response.
“This is a pivotal moment both for our company as well as the AI industry as a whole,” Dalton said. “Numerous teams are dropping everything to enhance our security prevention, detection, and response techniques both in our fundamentals and better use of AI. We’re consciously slowing down research [in order] to enhance security and to upgrade the security principles and foundation of our environment, and dramatically scaling up the monitoring of our AI agents, and improving our general security control environment across prevention, detection, and mitigation.”
Weigh that as you would any vendor recounting its own remediation from a conference stage. “Consciously slowing down research” is a remarkable thing to hear from OpenAI. It is also something nobody in the audience can verify.
The part that should worry defenders
Time and again, both speakers returned to the stakes beyond one company’s containment failure. This particular episode was an accident. Aim the same loop deliberately, and you get what they anticipate from attackers before long.
“The important takeaway here that has really shifted dramatically is that fully automated offensive loops require investment in truly, fully automated defense, and we are not there as an industry,” Dalton said. “We will have to find that path together with urgency.”
These post-mortems aren’t unique to OpenAI. Anthropic and the United Kingdom’s AI Security Institute have each published accounts of their own rogue-AI testing incidents, and collectively the industry is piecing together a running checklist of the system visibility and monitoring fundamentals needed to stop infrastructure from being hijacked by droves of lazy, reckless and ornery agents.
All of which loops back to the package manager. Any shared service your agents can write to should be treated as a communication channel — so go count how many messages are sitting in it already.
Correction: 8/5/2025 at 10 pm EDT: The name of the package manager is Artifactory not Hard Factory.



















STAY ALWAYS UP TO DATE