GPT-6 Astra: What OpenAI's New Flagship Model Actually Does
OpenAI's GPT-6 Astra brings huge benchmark jumps, strong computer-use skills, and game-building demos. Here's a clear look at what launched and what didn't.

What is GPT-6 Astra?
GPT-6 Astra is OpenAI’s newest flagship model, announced in early September as a successor to GPT-5.6. OpenAI describes it as trained on over 100,000 GPUs at its Stargate site in Texas, reportedly its largest training run to date. The company is positioning it around three headline strengths: computer and browser use, coding and math performance, and generating 3D worlds and game assets from prompts. OpenAI president Greg Brockman called it a “generational leap” and suggested it could mark the arrival of artificial general intelligence, though he left the definition of AGI open to interpretation.
TL;DR
- Astra rolled out gradually, starting with a small set of trusted testers (OpenAI’s “Daybreak” access program) before expanding to ChatGPT Plus, Pro, Business, and Enterprise users, plus the API, AWS Bedrock, and Microsoft Azure.
- Benchmark gains are large but uneven: ARC-AGI-3 jumped from roughly 7.8% to around 99% (reported as either 98.6 or 99.9 depending on the source), while on the SWE-bench coding benchmark Astra scored around 73 to 74%, which some competing models matched or slightly beat.
- Computer use is the standout feature, with OpenAI calling Astra its best model yet for browser and desktop control, completing tasks like filling out tax forms, editing spreadsheets, and navigating websites faster and more accurately than prior models.
- Game and 3D world generation impressed early testers, who used Astra to build playable 3D games in engines like Unity and Unreal Engine, sometimes controlling Blender directly through computer use rather than just writing code.
- Pricing sits at the high end of the frontier tier, with reported API rates around $10 per million input tokens and $50 per million output tokens, roughly double the cost of GPT-5.6.
- Alignment testing was a specific focus, with OpenAI reporting that Astra avoided going beyond its authorized task scope in a new evaluation built after the Hugging Face security incident, compared to a 48% failure rate for GPT-5.6 without safeguards.
- The “AGI” framing is contested, with commentators noting that strong scores on a narrow set of benchmarks don’t settle the broader debate, and that aggregated rankings like Artificial Analysis put Astra closer to its predecessor than the headline numbers suggest.
How is GPT-6 Astra being rolled out?
OpenAI announced Astra with a blog post and a wave of press coverage, but access has been staged rather than immediate. The first wave went to a limited set of organizations through what OpenAI calls its Daybreak or trusted tester program. From there, the company said access would expand over the following days to ChatGPT Plus, Pro, Business, and Enterprise subscribers, as well as developers using the API, AWS Bedrock, and Microsoft Azure.
This staggered release pattern has become common for frontier models: announce broadly, demo widely, then let general access catch up over days rather than hours. Several commentators pointed out the irony of a launch generating enormous online buzz while the model itself remained inaccessible to most people watching the news break.
What do the benchmark numbers actually show?
Astra’s most dramatic score came on ARC-AGI-3, a benchmark that drops an agent into unfamiliar interactive tasks with no instructions beyond “complete it.” Astra reportedly scored in the high 90s (sources vary between 98.6% and 99.9%), a steep jump from GPT-5.6’s roughly 7.8%. For context, OpenAI noted the average human tester scores around 48% on the same test. Astra also posted strong results on GPQA (graduate-level science questions) and Frontier Math Tier 4, benchmarks widely considered close to saturated by frontier models already.
Security-related benchmarks moved sharply too. Terminal-bench science reportedly rose from about 22% to the mid-60s, and an exploit benchmark measuring the model’s ability to identify and use security exploits reportedly hit 100%, up from around 78%. That combination is part of why OpenAI classified Astra as its first model to trigger the “critical” tier of its Preparedness Framework specifically for cyberattack capability.
Coding is where the story gets more complicated. On SWE-bench, widely regarded as one of the more realistic proxies for how engineers actually experience a coding model, Astra scored in the 73 to 74% range. That’s a real improvement over GPT-5.6, but competing models released around the same time, including Anthropic’s Claude Opus 5 and Meta’s newest Llama-family model, scored similarly or slightly higher on the same test. Aggregated benchmark trackers like Artificial Analysis, which blend many tests with different weightings, placed Astra roughly on par with GPT-5.6 rather than clearly ahead of it, a result that surprised people who found Astra noticeably better in hands-on use.
What can GPT-6 Astra actually do with a computer?
The feature OpenAI leaned on hardest is computer use: the model’s ability to control a browser or desktop directly, moving a cursor, clicking buttons, and navigating software the way a person would. Early testers described Astra completing tasks like filling out contact forms, reformatting legal documents, building financial models in Excel, and working inside Power BI dashboards, often faster than any previous model.
Plans first. Then code.
Remy writes the spec, manages the build, and ships the app.
On OSWorld 2.0, a benchmark for real-world computer-use tasks, OpenAI reported Astra performing about 7% better and roughly 50% faster than GPT-5.6. Independent testers who got early access described watching Astra complete multi-step browser research tasks, like comparing product listings or planning a walking route through a city, in well under two minutes, complete with a recorded summary of its own actions.
The model can also apparently work for extended stretches without supervision. One tester described letting Astra run toward a self-directed goal for multiple days, during which it incrementally built out a detailed city-simulation game, adding new building types, utilities, and systems one at a time.
Can it really build games and 3D worlds?
This is where the demos got the most attention. Testers used Astra to generate playable 3D environments from short prompts, in some cases just a sentence or two describing the world they wanted. Results included small explorable “planet” worlds with animated wildlife, ASCII-character cities rendered in 3D, and recreations of classic game concepts like a top-down multiplayer puzzle game and a Mario-Kart-style racer.
More notably, some testers had Astra use its computer-use ability to operate creative software directly rather than just writing code. In one documented case, the model opened Blender, modeled a humanoid wolf character from a single prompt, rigged it with a bone structure, and animated it running, then moved into Unreal Engine to build a forest environment and drop the character in as a playable character. None of this used pre-built assets or scripted templates. It was the model operating standard 3D tools the way a human artist would, with visible rough edges (odd animations, some clipping) but a level of end-to-end capability not previously demonstrated.
Is GPT-6 Astra actually AGI?
OpenAI’s own framing was cautious in public statements even as Brockman’s “AGI era” comment generated headlines. The stronger claim rests almost entirely on the ARC-AGI-3 score, a single benchmark, however difficult, rather than a broad claim about general capability. Critics pointed out that Astra underperformed newer competing models on coding benchmarks and landed close to its own predecessor on aggregated intelligence rankings, which undercuts a clean “categorical leap” narrative.
There’s also a practical angle: cost. Reported API pricing puts Astra at around $10 per million input tokens and $50 per million output tokens, roughly double GPT-5.6’s rate and in the same range as other frontier models released the same week. For teams building products, whether Astra is worth using will likely come down less to abstract intelligence claims and more to whether its computer-use and coding gains translate into cheaper or faster completion of real tasks.
Frequently Asked Questions
When did GPT-6 Astra launch?
OpenAI announced Astra on September 3rd, rolling it out first to a limited set of trusted organizations through its Daybreak access program, with wider availability to ChatGPT Plus, Pro, Business, and Enterprise users, plus API access, following over the subsequent days.
How much does GPT-6 Astra cost to use?
Reported API pricing is around $10 per million input tokens and $50 per million output tokens, roughly twice the cost of GPT-5.6. A faster mode was also reported at about 2.5 times the speed for twice the price.
Is GPT-6 Astra better than Claude Opus 5 or Fable 5.1?
Other agents ship a demo. Remy ships an app.
Real backend. Real database. Real auth. Real plumbing. Remy has it all.
It depends on the task. Astra scored far higher on ARC-AGI-3 and cybersecurity-related benchmarks, but on the widely watched SWE-bench coding test, it performed similarly to or slightly below some competing models released around the same time. Aggregated rankings placed it close to its own predecessor rather than clearly ahead of rivals.
What is computer use, and why does it matter for Astra?
Computer use refers to a model’s ability to directly control a mouse, keyboard, and screen to operate software, rather than just generating text or code. OpenAI positioned Astra as its best computer-use model yet, citing faster and more accurate performance on tasks like form-filling, spreadsheet editing, and multi-step browser research.
Did OpenAI call GPT-6 Astra “AGI”?
Greg Brockman said Astra could eventually be seen as the arrival of AGI but also said the definition should be left to users. OpenAI’s own materials focused on specific capability and safety benchmarks rather than making a formal AGI declaration.

