OpenAI has introduced GPT-6 Astra, its newest frontier model and what the company describes as the “world’s most intelligent and aligned model.”
That is a big claim, even by AI launch standards.
Astra arrives with eye-catching results across mathematics, abstract reasoning, software engineering, cybersecurity and computer use. OpenAI is also positioning the model as something more than another intelligence upgrade. The emphasis this time is on completing real work: navigating software, working across long tasks, handling professional documents and making judgment calls without constantly handing control back to the user.
And then there is the AGI question.
OpenAI President Greg Brockman has suggested the company may have reached that territory, adding another layer to what was already going to be one of the most closely watched AI releases of the year.
GPT-6 Astra Posts Huge Gains on ARC-AGI-3
The number grabbing most of the attention is 99.9% on ARC-AGI-3, a benchmark built around unfamiliar environments that test whether an AI system can work out new rules rather than simply reproduce patterns it has already seen.
There is an important detail behind that headline. ARC Prize reported that Astra reached 99.9% using its Provider Adapter harness, which preserves reasoning state between requests and uses context compaction. Under the organization’s standard harness, Astra scored 62.7%. Both results were state-of-the-art, but the difference shows why benchmark methodology matters almost as much as the final number.
Perhaps the more interesting result came from efficiency. ARC Prize found that Astra used fewer actions than the median human participant on 96% of tested levels. It was not simply solving the environments. It was learning how to navigate them with surprisingly little wasted movement.
Mathematics and Science Get Another Big Jump
Astra also scored 98% on FrontierMath Tier 4, according to OpenAI, putting the model near the ceiling of one of the more demanding mathematical evaluation sets currently being used for frontier systems.
OpenAI says Astra has already contributed to work involving long-standing mathematical problems, while the wider launch includes improvements across scientific reasoning and specialist research tasks. The company is increasingly pitching its models as research collaborators rather than just systems that answer scientific questions.
That distinction matters. A model that can reason about a scientific problem is useful. One that can inspect data, operate specialist software, generate plots and continue working across several steps starts to look much more like a research agent.
Cybersecurity Is Where Astra Gets More Complicated
One benchmark result stands out for a different reason.
GPT-6 Astra scored 100% on ExploitBench, compared with 78.5% for GPT-5.6 Sol. OpenAI also reported a 42.4% success rate on ExploitGym, up from 30.3% for its previous frontier cyber model.
Those capabilities pushed Astra into the Critical cybersecurity capability level under OpenAI’s Preparedness Framework — the first broadly deployed OpenAI model to cross that threshold. The company says the model can, with the right tools and permissions, discover previously unknown vulnerabilities and develop methods for exploiting sophisticated systems.
During testing, Astra reportedly discovered two previously unknown zero-day vulnerabilities. OpenAI said it is disclosing those issues to the relevant maintainers.
That makes the release unusually useful and unusually uncomfortable at the same time. The same model that could help security teams find vulnerabilities before attackers do could also make advanced offensive capabilities easier to scale. OpenAI says stronger monitoring and restrictions have been added around those higher-risk tasks.
Astra Is Built to Actually Use Computers
Benchmark numbers are only part of this release.
OpenAI has put a lot of Astra’s training into computer use, letting the model interact with software, websites and professional workflows rather than stopping after generating instructions.
The company says Astra can fill out online forms, update CRM systems, organize calendars, conduct research, work inside document editors, analyze datasets, build websites and perform frontend quality checks.
On OSWorld 2.0 simulations, Astra reportedly completed computer-use tasks in roughly 47% less time than GPT-5.6 Sol while also improving its success rate. OpenAI says updates to the Codex harness combined with Astra can deliver around 1.9x faster task completion on Mind2Web.
This may prove more important than another few percentage points on a reasoning leaderboard. The AI race is increasingly about whether models can finish work instead of merely explaining how someone else should do it.
Coding Gets Smarter and Long Tasks Lose Less Context
OpenAI is also calling Astra its strongest software engineering model so far.
The model has been designed to keep track of complicated projects over much longer periods. In Codex, Astra can maintain notes across context windows and search previous context when it needs to recover requirements, earlier test results or decisions made during a long coding session.
That addresses one of the irritating problems with agentic coding systems: they can be brilliant for 20 minutes and then quietly forget why they made an architectural decision an hour earlier.
Astra’s API supports a context window of roughly 1.05 million tokens with as many as 128,000 output tokens, giving developers considerably more room for large repositories, extended research sessions and complex document workflows.
Independent Benchmarks Tell a Less Convenient Story
OpenAI’s internal numbers make Astra look dominant almost everywhere. Independent testing is messier.
Artificial Analysis initially placed GPT-6 Astra at 61 on its Intelligence Index, putting it among the leading frontier models but behind systems including Claude Fable 5.1 and Meta’s Muse Spark 1.3 on the composite evaluation. Artificial Analysis also found Astra roughly level with GPT-5.6 Sol on that particular intelligence measure.
That does not necessarily contradict OpenAI’s benchmark results. It shows how difficult it has become to declare a single “best AI model.”
Astra can dominate ARC-AGI-3, cybersecurity and several agentic evaluations while another model leads a broader composite intelligence index. Different tests increasingly reward very different kinds of intelligence.
This is probably what frontier-model competition is going to look like from here. Fewer clean victories. More arguments over which benchmark actually resembles the work people care about.
GPT-6 Astra API Pricing Starts at $10 and $50 Per Million Tokens
That performance is not cheap.
OpenAI has priced GPT-6 Astra at $10 per million input tokens and $50 per million output tokens through its API. Cached input costs $1 per million tokens. Prompts exceeding 272,000 input tokens also receive higher long-context pricing.
The sticker price is considerably higher than GPT-5.6 Sol, although raw token pricing does not tell the whole story.
Artificial Analysis found Astra can use substantially fewer tokens for some coding workloads, while OpenAI’s early partners have also reported lower token consumption on complex workflows. In other words, developers may pay more for each token while needing fewer tokens to finish the same job.
That is going to matter far more to businesses than the headline API rate.
The awkward part for OpenAI is that Claude Fable 5.1 is competing at the same $10/$50 base API price, creating an unusually straightforward comparison between two frontier systems.
Most ChatGPT Users Still Have to Wait
Despite the scale of the announcement, Astra did not arrive for everyone at once.
OpenAI started the rollout with a limited group of organizations. The company says access will expand over the coming days to ChatGPT Plus, Pro, Business and Enterprise users, along with the OpenAI API, Microsoft Azure and AWS Bedrock.
Pro, Business and Enterprise customers are also expected to receive access to GPT-6 Astra Pro, a higher-end version aimed at particularly demanding workloads.
For a launch this big, that staggered availability is slightly frustrating. Benchmarks can dominate AI Twitter for a few days, but independent users still need to put Astra through messy codebases, unreliable websites, giant spreadsheets and badly written prompts before anyone knows how much of the performance survives outside controlled evaluations.
Has OpenAI Reached AGI?
This is where the Astra launch gets philosophical very quickly.
OpenAI executives have spoken about the model as something that may mark the beginning of an AGI era, with Brockman publicly suggesting Astra could eventually be remembered as that turning point.
Whether that makes Astra “AGI” is another question entirely.
There is still no universally accepted benchmark that flips from zero to one and announces that artificial general intelligence has arrived. Even inside the AI industry, the definition shifts depending on whether people are talking about reasoning, autonomous work, economic usefulness, human-level performance or the ability to learn unfamiliar tasks.
Astra gives the argument considerably more fuel. It does not settle it.
Why GPT-6 Astra Matters
The flashy answer is the 99.9% ARC-AGI-3 score.
The more consequential answer may be everything happening around it.
Astra can operate computers faster, sustain longer workflows, write and debug software, work with professional documents and perform cybersecurity tasks that would have looked wildly ambitious for a commercial AI system only a few model generations ago.
It also comes with warning signs. Cyber capabilities have reached a level that OpenAI itself classifies as Critical. Independent leaderboards do not show Astra crushing every competitor. And most customers did not get immediate access when the announcement landed.
That makes this release more interesting, not less.
The frontier is no longer moving in a straight line where one new model simply beats the old model on everything. GPT-6 Astra looks exceptionally strong at agency, computer use, coding and certain forms of reasoning. Claude Fable 5.1 remains a serious rival. Other models are winning their own evaluations.
The next comparison will not happen on a benchmark chart.
It will happen when millions of people finally hand these systems real work.
Sources
OpenAI — GPT-6 Astra: A New Generation of Intelligence
Read the official OpenAI announcement

