Roughly seven weeks. That’s the gap between OpenAI putting GPT-5.6 Sol, Terra and Luna in users’ hands and the debut of a completely new model family, GPT-6 Astra, which the company describes as “the most intelligent and aligned model in the world.”
It may also be the final major release for a while. Back in August, OpenAI resolved to slow frontier model development after one of its own models hacked the AI platform Hugging Face, and Astra arrives under the cloud of that decision.
What Astra is actually designed to do
The selling point isn’t nicer writing. It’s your computer. According to OpenAI, Astra is strongest at computer-use and browsing, along with software engineering, cybersecurity, science and general professional work.
In the launch demo, the model handles 3D modeling and puts together slideshows, often juggling multiple tasks across separate domains simultaneously — ordering food while coding a game, for instance.
Everything rests on one claim: that Astra can keep a multi-step workflow coherent on your machine, hopping in and out of the browser without losing the thread. OpenAI says it manages this with “strong visual judgment” and without straying from the original prompt or directions. Allegedly.
The benchmark that comes with an asterisk
Numbers came with the launch, as they always do. The standout is 98.6% on ARC-AGI-3, an industry benchmark for tackling unfamiliar problems.
Handle that figure with care. Not every model run through the benchmark is set up the same way, and architectural differences — whether a model has persistent memory, for example — can shift the outcome.
The remaining results are less murky. Astra posted 57.7% on the coding benchmark Terminal Bench 4.0 and 59.3% on the Agent’s Last Exam, a measure of agentic ability, with both figures well above GPT-5.6 Sol. Additional benchmarks appeared in OpenAI’s announcement, and the gist is that Astra now sits at or near the top of most leaderboards.
The cybersecurity results are the real story
This is where the launch turns genuinely uncomfortable. On ExploitBench, which gauges how well a model can exploit software vulnerabilities, Astra earned a perfect score. GPT-5.6 Sol managed 78.5% on the same test.
OpenAI also reports results on SRE-Bench, which tests reverse engineering of software binaries without access to raw source code: “Astra solved 88.0% of tasks in a single attempt and 99.2% within four attempts, compared with 55.9% and 68.7% for GPT-5.6 Sol, respectively.”
Put those two figures side by side and you have a model that moved from competent to almost flawless at locating and exploiting software flaws inside a two-month release cycle. The responsible-deployment question is right there in the open.
OpenAI’s response is alignment work — spanning everything from adhering to templates to communicating more transparently — plus a model engineered to flatly refuse advanced cybersecurity tasks. Beefed-up protections should also let Astra withstand jailbreak attempts while making misuse easier to spot, the company says. Refusal training and a perfect ExploitBench score now coexist inside the same model, and users will discover how well that arrangement holds before OpenAI publishes anything on it.
Who gets it, and what it costs
Per OpenAI, Astra is “rolling out today to a limited set of organizations and over the coming days will become available to all ChatGPT Plus, Pro, Business, and Enterprise users, as well as through the OpenAI API and AWS.”
Via the API, the rate is $10 per million input tokens and $50 per million output tokens. That’s the pricey end of the market, and OpenAI is aware of it — which explains why the company would prefer you stop tallying tokens at all.
“What you actually want, and I think the market is starting to really wake up to, is the price per task,” OpenAI President Greg Brockman said. “It’s just about: can you get the thing done for an appropriate cost at appropriate speed?”
That argument holds only if Astra nails the job on the first try. At $50 per million output tokens, a model that takes three swings at your workflow is one you can’t justify, and the SRE-Bench gap between one attempt and four is precisely what determines your invoice. Push your own agentic task through it before migrating a team, and tally retries rather than tokens.

















STAY ALWAYS UP TO DATE