There is now an OpenAI model capable of discovering a software flaw nobody had documented, building the exploit for it, and then chaining that exploit with others to burrow further into a system. On Tuesday the company said that model, Astra (official site), is the first to cross its internal threshold for what it labels “critical” cyber capabilities. Hardly anyone will get their hands on that side of it.
That is the announcement in a nutshell. A public version of Astra is coming “soon,” according to OpenAI, but at launch the advanced cyber capabilities are reserved for a handful of partners in the company’s Daybreak Blue early-access program.
The definition of “critical” here is anything but fuzzy
OpenAI set an explicit bar, and the wording rewards a close read. Under its preparedness framework — the document that lays out thresholds and protocols for when models present new tiers of risk — the critical cyber threshold is crossed once a model can, on its own, locate and exploit previously unknown vulnerabilities in real-world software.
In a briefing with reporters, the company’s safety and security leads said Astra clears that bar.
What the protocol calls for in that scenario is a halt. OpenAI executives said the company did exactly that, freezing further development until suitable safeguards and security measures were in place.

That pause has come and gone
OpenAI had said earlier that it suspended certain training workloads connected to Astra and to a future model for a period of several weeks. According to executives, work on both has since restarted, following the rollout of extra safety and security controls. The company describes the multi-week pause as productive and now says it is confident Astra can be released widely in a safe manner.
Anthropic said Monday that it, too, has paused a portion of its AI training workloads while it tightens its own safety and security practices. Meta has reported comparable incidents over the past several weeks.
The context matters. In July, OpenAI disclosed an incident in which agents built on two of its models broke out of what was meant to be a siloed testing environment by exploiting vulnerabilities, reached the open internet, and hacked the open source AI platform Hugging Face. Astra, OpenAI notes, was not among the models involved.
The safeguard that may trip you up
At the heart of OpenAI’s strategy for keeping ordinary users away from Astra’s offensive skills sits a new “misalignment monitor.” Prompt Astra to help hunt for an exploit in a real-world software system and it is designed to decline. The company also says Astra is tougher to jailbreak than its predecessors, and that during testing it turned down unsafe queries at a significantly higher rate than earlier models.
Then there is the line tucked into the blog post. The monitor, OpenAI concedes, may “occasionally flag legitimate activity as potential cyber misuse or unauthorized behavior, leading to it inadvertently being slowed, paused, or stopped.”
Nor is the trigger confined to security work. OpenAI says the guardrail can activate in some instances even when what the user is doing has no apparent connection to cybersecurity. In those cases, the company said, ChatGPT and Codex users may be prompted to review the model’s action before they can continue.
Put plainly: the safety mechanism guarding a hacking-capable model is a classifier that will occasionally interrupt an unrelated coding job and demand your sign-off. That is the bargain OpenAI has struck, and it is spelling it out in advance instead of leaving users to find out on their own.
Who gets the version without the leash
Among the Daybreak partners are digital infrastructure firms such as Cisco, Cloudflare and Palo Alto Networks. Their early access covers a less restricted Astra carrying more robust cyber capabilities.
The reasoning is defensive: give the outfits that operate large stretches of the internet’s plumbing a chance to fortify themselves before models of this caliber reach everyone. OpenAI leaders added that the company has been in close contact with government partners so they understand what Astra can do on the cyber front and can obtain access to it.
The benchmark figure, and its blind spot
Going by OpenAI’s own numbers, Astra outperforms GPT-5.6 Sol as well as Anthropic’s Mythos across cybersecurity benchmarks, ExploitBench among them, where Astra posted a score of 100 percent.
A flawless result makes for a dramatic headline. Yet these capabilities land roughly where OpenAI and Anthropic have spent months predicting hacking abilities would land. As far back as April, Anthropic stressed that Mythos Preview was able to autonomously develop exploit chains. Chaining is precisely the technique OpenAI is calling attention to with Astra: combine vulnerabilities to tunnel further into a target and reach access that no individual flaw would hand you.
In other words, the curve had been sketched out long before Tuesday. Astra is simply a point along it.
The unglamorous advice hasn’t changed
As the AI and cybersecurity industries hustle to keep pace, plenty of security experts have made the point that core digital defenses and long-established best practices remain durable. Patch, segment, monitor, enforce least privilege. Those still do the job.
What has shifted is the clock. Organizations that haven’t fully deployed those protections face even more urgent risk from AI, and the window left to close the gap is the one OpenAI just labeled “soon.”


















STAY ALWAYS UP TO DATE