Skip to main content
MindStudio
Pricing
BlogAbout
My Workspace
GPT-6 Astra benchmarksARC-AGI-3DeepSWE score

GPT-6 Astra Benchmarks: How It Really Compares to Fable 5.1 and Gemini

GPT-6 Astra's ARC-AGI-3, DeepSWE, and Frontier Math scores compared against Claude Fable 5.1 and Gemini 3 Flash, with the numbers that actually matter.

Edited by Luis Chavez-Mattos, Director of Product RSS
GPT-6 Astra Benchmarks: How It Really Compares to Fable 5.1 and Gemini

What are GPT-6 Astra’s actual benchmark scores?

GPT-6 Astra, OpenAI’s newest frontier model released in September, posted a reported 99.9% on ARC-AGI-3 (some outlets cited 98.6%), 97.6% on Frontier Math Tier 4, and roughly 73-74% on DeepSWE, a coding benchmark. It also hit 100% on Exploit Bench and 64.6% on Terminal Bench Science. Those numbers put it at or near the top of most public leaderboards, though on DeepSWE specifically, Gemini 3 Flash and Meta’s Muse Spark 1.3 scored close to or slightly above it, which complicates the “best model across the board” narrative OpenAI is pushing.

TL;DR

  • ARC-AGI-3, a benchmark designed to test how well an agent learns unfamiliar interactive tasks, went from 7.8% on GPT-5.6 to roughly 99% on Astra, compared to an average human score of 48%.
  • DeepSWE, widely seen as the coding benchmark that best matches real developer experience, landed around 73-74% for Astra, which is only a couple points above GPT-5.6 and actually below Gemini 3 Flash (73.7-74%) and Meta’s Muse Spark 1.3 (75.4%) in some rankings.
  • Frontier Math Tier 4 came in at 97.6% for Astra versus 87.8% for Claude Fable 5.1, a clear and meaningful gap in advanced math reasoning.
  • Cost is steep: Astra runs $10 per million input tokens and $50 per million output tokens on the API, with a “fast mode” offering 2.5x speed for 2x price, and early testers estimated real task costs around $167 on aggregated cost-per-task tracking.
  • Computer and browser use is where testers reported the biggest qualitative jump, with OS World 2.0 scores about 7% higher and 50% faster than GPT-5.6, and alignment testing showing 0% unauthorized task completion versus 48.2% for the prior model.
  • Aggregated benchmarks tell a different story than raw scores, with Artificial Analysis ranking Astra only fifth overall, roughly tied with GPT-5.6, despite testers saying it feels like a much bigger leap in practice.
  • Astra reportedly contributed to two small advances in prime number research, a rare case of a released model output being cited as a genuine, if narrow, contribution to open math problems.

Everyone else built a construction worker.
We built the contractor.

🦺
CODING AGENT
Types the code you tell it to.
One file at a time.
🧠
CONTRACTOR · REMY
Runs the entire build.
UI, API, database, deploy.

How does Astra compare to Claude Fable 5.1 on hard benchmarks?

The clearest win for Astra over Anthropic’s Claude Fable 5.1 shows up in Frontier Math Tier 4, where Astra scored 97.6% against Fable’s 87.8%, a 10-point gap that suggests a real advantage in multi-step mathematical reasoning. Astra also reportedly beat Fable 5.1 on Bench CAD, a benchmark testing 3D object creation, by more than 10 points (95.9% versus roughly 85%), which lines up with tester reports that Astra is unusually strong at spatial reasoning and building 3D assets.

On DeepSWE, the coding benchmark most engineers watch closely, the gap narrows. Astra scored around 73%, ahead of Fable 5.1’s reported 67% in one comparison, but that’s a smaller margin than the math and CAD results suggest. Cost is where Fable 5.1 looks weakest: multiple testers flagged it as the most expensive model per task to run, with one creator reporting it burned through 40% of a weekly usage quota on a single task. If Astra can match or beat Fable’s quality while consuming fewer resources, that’s arguably a bigger practical difference for builders than any single benchmark score.

How does Astra compare to Gemini 3 Flash and other models?

This is where the picture gets less flattering for OpenAI’s launch narrative. On DeepSWE, Gemini 3 Flash scored around 73.7-74%, essentially matching or slightly beating Astra’s 73-74%. Meta’s Muse Spark 1.3, which launched around the same time, reportedly scored 75.4% on the same benchmark, putting it ahead of Astra on the metric most developers treat as a proxy for real-world coding feel.

That inconsistency showed up clearly in an informal “Bucky Bench” SVG-drawing test one creator ran across models. Muse Spark 1.3, despite outperforming Astra on DeepSWE, produced a noticeably worse image when asked to draw a cartoon figure using code, scoring around 5.7 out of 10 by an AI-judge rubric versus a clear first-place result for Astra. That mismatch between a leaderboard number and hands-on output quality is a recurring theme with Astra’s launch: the benchmark suite doesn’t always predict how a model performs on a given creative or coding task.

Why does Astra score so high on ARC-AGI-3?

ARC-AGI-3 is built to be hard for AI systems specifically. It drops an agent into an unfamiliar interactive task, usually something game-like, with no instructions beyond “complete it,” and measures how quickly the system learns the rules through trial and error. Earlier models struggled badly here. GPT-5.6 scored just 7.8%, and Claude Opus 5 scored 30.2%. Astra’s reported 99.9% (some coverage says 98.6%) is a near-complete saturation of the test, especially striking against an average human score of 48%.

Worth noting: Astra reportedly took this test with the assistance of OpenAI’s response API harness, meaning the scoring setup may include tooling support that isn’t identical to a bare model attempt. That doesn’t erase the result, but it’s a caveat worth keeping in mind before treating the number as a clean apples-to-apples comparison with older model scores.

Is GPT-6 Astra worth the cost for developers?

REMY IS NOT
  • a coding agent
  • no-code
  • vibe coding
  • a faster Cursor
IT IS
a general contractor for software

The one that tells the coding agents what to build.

At $10 per million input tokens and $50 per million output tokens, Astra sits solidly in frontier pricing territory, more expensive than GPT-5.6 but reportedly cheaper to run than Claude Fable 5.1 for comparable tasks. The fast mode option, offering 2.5x speed for 2x the price, is a better ratio than the roughly 1.5x speed for 2x price typically seen on other fast-mode offerings.

For coding specifically, the DeepSWE numbers suggest Astra is competitive but not a clear leader over Gemini 3 Flash or Muse Spark 1.3. Where it does appear to separate itself is computer and browser use: OS World 2.0 scores improved about 7% while running roughly 50% faster than GPT-5.6, and testers reported it completing multi-step browser tasks (form filling, spreadsheet work, research and comparison tasks) in a fraction of the time older models needed. If your use case leans heavily on agentic browser or desktop automation rather than pure code generation, the practical gains look larger than the raw coding benchmark suggests.

What do the benchmarks miss?

Artificial Analysis, which aggregates multiple benchmarks into a single weighted score, ranked Astra fifth overall, close to a tie with GPT-5.6. That’s a notable disconnect from firsthand tester accounts describing Astra as a clear step up in usability, particularly for 3D world generation, game prototyping, and long-running agentic tasks. Multiple testers built playable 3D environments, animated game characters in Blender, and assembled Unreal Engine scenes using natural language prompts and Astra’s computer-use capabilities, results that don’t map cleanly onto any single leaderboard number.

Alignment testing is another dimension benchmarks alone don’t capture. OpenAI reported that under adversarial evaluation conditions created after a prior security incident, Astra completed zero percent of unauthorized actions when given difficult or intentionally ambiguous tasks with guardrails, compared to 48.2% for GPT-5.6. That’s a meaningful safety data point that sits outside the usual math and coding scorecards.

Frequently Asked Questions

What is GPT-6 Astra?

GPT-6 Astra is OpenAI’s frontier model released in September, trained on what the company describes as its largest training run to date. It was rolled out first to trusted testers and enterprise partners before wider availability to Plus, Pro, Business, and Enterprise ChatGPT users and API customers.

Did GPT-6 Astra beat every model on every benchmark?

No. While it scored very high on ARC-AGI-3, Frontier Math Tier 4, Exploit Bench, and Bench CAD, it did not lead on DeepSWE, the coding benchmark many developers consider the best real-world proxy. Gemini 3 Flash and Meta’s Muse Spark 1.3 both scored at or above Astra on that specific test in various rankings.

How much does GPT-6 Astra cost to use?

Through the API, Astra is priced at $10 per million input tokens and $50 per million output tokens, with a fast mode available at roughly 2.5x speed for 2x the price. Early cost-per-task estimates put it above GPT-5.6 but reportedly below Claude Fable 5.1 for comparable workloads.

Is a 99.9% ARC-AGI-3 score believable?

The score is credited with an important caveat: it was achieved using OpenAI’s response API harness, which may provide tooling support beyond a bare model call. The result is still a dramatic jump from GPT-5.6’s 7.8% and from the human average of 48%, but the harness detail matters for anyone comparing it strictly to other models’ scores.

Does GPT-6 Astra qualify as AGI?

Other agents start typing. Remy starts asking.

YOU SAID "Build me a sales CRM."
01 DESIGN Should it feel like Linear, or Salesforce?
02 UX How do reps move deals — drag, or dropdown?
03 ARCH Single team, or multi-org with permissions?

Scoping, trade-offs, edge cases — the real work. Before a line of code.

OpenAI executives have used the term loosely around this launch, but the claim is contested. The model shows major gains in narrow domains like computer use, agentic coding, and interactive world building, but strong scores on specific benchmarks don’t settle the broader definitional debate over what counts as general intelligence.

Editorial standards

Presented by MindStudio

No spam. Unsubscribe anytime.