GPT-6 Astra is the first model making OpenAI willing to declare the "AGI era"
Key Points
OpenAI is launching GPT-6 Astra, its most capable model to date. Paying ChatGPT customers and cloud platforms should get access in the coming days.
In benchmarks, Astra significantly outperforms its predecessor GPT-5.6 Sol and Anthropic's Fable 5 models across key disciplines including logic, math, and software engineering.
Token prices are 2.5x higher and on par with Anthropic's Fable 5.1, but OpenAI argues that the cost per completed task is actually lower depending on the use case.
Update – Sep 3, 2026
Added Codex and GPT-6 Pro details and the launch video
OpenAI has shipped GPT-6 Astra, its most capable model to date. President Greg Brockman says it might already qualify as "AGI" or is at least within reach, meaning an AI system that outperforms humans at most economically valuable work by OpenAI's own definition.
GPT-6 Astra is rolling out first to select organizations through OpenAI's Daybreak program , with broader availability for ChatGPT Plus, Pro, Business, and Enterprise customers expected in the coming days. It will also be accessible through the API and cloud platforms like AWS Bedrock and Microsoft Azure. Pro, Business, and Enterprise subscribers get access to GPT-6 Astra Pro, a higher-performance variant, though enterprise workspace admins need to activate the model manually.
Astra was pretrained on more than 100,000 GPUs at the Stargate facility in Texas. OpenAI researcher Aidan Clark called it the company's largest training run ever. The jump from Sol to Astra represents a bigger capability gain than the jump to Sol from earlier models, Clark said, in part because previous AI models played a role in monitoring training. Ad
In benchmarks OpenAI published, Astra scores well above its predecessor GPT-5.6 Sol and Anthropic's Fable models. Astra hits top marks across a range of disciplines: logical reasoning (99.9 percent on ARC-AGI-3, though under its own test conditions ), math (97.6 percent on FrontierMath Tier 4 v2), software engineering (74.1 percent on DeepSWE v1.1), expert knowledge (96 percent on GPQA Diamond), engineering (95.9 percent on BenchCAD), and cybersecurity (100 percent on ExploitBench). Ad
Computer Use
Professional
BenchCAD cost: ~43% below Sol, ~86% below Fable 5.1.
Coding
Terminal-Bench 4.0 cost: ~9% below Sol, ~63% below Fable 5.1. Ad
Academic
Lower-cost settings: Terminal-Bench Science 61.1% at ~27% lower cost; GPQA Diamond 94.9% at ~37% lower cost. Prime gaps improved from 240 to 186, and a large-gap bound term improved for the first time in over 80 years.
Science and Health
Fable 5 and 5.1 are not included in LifeSciBench, GeneBench Pro, and MedChemBench because they reject most questions. Ad
Cybersecurity
SRE-Bench within four attempts: 99.2% versus 68.7% for Sol. Astra found two previously unknown zero-days during evaluation. Ad
Alignment (lower is better except for Impossible ExploitGym)
Impossible-task scope test: Sol exceeded its authorized target 48% of the time, Astra 0%. Astra is 3x less likely to misstate its own capabilities.
Long Context
Abstract Reasoning
In scientific work, the model reportedly improved a mathematical result on prime gaps and set new records in biology, chemistry, medicine, and physics evaluations. OpenAI also positions Astra as a model that can reliably operate a computer the way a human would, and on OSWorld 2.0, which measures that ability, Astra scored 72.6 percent at about 40 minutes per task compared to Sol's 65.7 percent at roughly 75 minutes. "Anything you can do on a computer, Astra can do for you. Fast," the company claims .
Alongside Astra, OpenAI is also updating its Codex coding environment. A new experimental feature lets the model take notes across multiple context windows during long sessions, rather than compressing everything into a single summary each time. Earlier context windows remain searchable, so Astra can look up requirements or test results from previous messages even if they weren't captured in the notes. OpenAI plans to make this the default in the coming weeks.
More expensive than Sol, but OpenAI says cheaper per task
GPT-6 Astra costs $10 per million input tokens and $50 per million output tokens in standard mode through the API. Fast mode, which promises 2.5x speed, doubles the price, making Astra 2.5x more expensive than GPT-5.6 Sol and putting it in the same price range as Anthropic's Fable 5.1 .
Brockman argued that token prices are becoming a poor way to compare models, since OpenAI's tokens aren't the same as a competitor's and aren't even comparable across its own model families. What matters, he said, is the price per completed task , and OpenAI is already experimenting with that pricing model . On DeepSWE v1.1, Astra's top configuration cuts estimated API costs per task by about 57 percent compared to Sol, according to the company.
During the press briefing, Brockman acknowledged there's no clearly defined AGI moment, saying the team originally thought there would be an obvious threshold everyone would recognize when OpenAI was founded. That's not how it played out, and the transition has been more gradual than expected, a position OpenAI laid out last spring . Brockman closed the briefing by saying , "Welcome to the AGI era." OpenAI CEO Sam Altman had previously said he expects a model he'd call AGI by the end of the year .
First model to hit OpenAI's critical cybersecurity threshold
Astra is the first model that OpenAI classifies as "critical" under its Preparedness Framework . That means the model can find previously unknown vulnerabilities and build exploit chains across well-defended systems when given the right tools and access, all without a human guiding it step by step. OpenAI also admits that Astra's written reasoning is harder to monitor than GPT-5.6 Sol's.
On ExploitBench , a benchmark that measures a model's ability to discover and exploit real software vulnerabilities, Astra scores a perfect 100 percent. OpenAI says the model discovered two previously unknown vulnerabilities during evaluation, which the company then reported to the affected vendors. Human experts confirmed that Astra can identify novel zero-day vulnerabilities across several software categories, including browsers and operating systems.
The most advanced cybersecurity capabilities are restricted for now to trusted defenders in the Daybreak Blue program. OpenAI had previously delayed Astra's release to run more safety testing. These dual-use capabilities cut both ways: an agent that can autonomously find a vulnerability helps a defender patch it just as easily as it helps an attacker exploit it.
AI News Without the Hype – Curated by Humans
Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section.