Skip to main content
MindStudio
Pricing
BlogAbout
My Workspace
GPT-6 Astra benchmarksGPT-6 Astra vs Fable 5.1ARC-AGI-3

GPT-6 Astra Benchmarks: Is It Really Better Than Fable 5.1?

GPT-6 Astra scores 99.9% on ARC-AGI-3, but independent benchmarks show mixed coding results against Fable 5.1 and Opus 5. Here's the full picture.

Edited by Luis Chavez-Mattos, Director of Product RSS
GPT-6 Astra Benchmarks: Is It Really Better Than Fable 5.1?

Is GPT-6 Astra actually better than Claude Fable 5.1?

Not consistently. GPT-6 Astra beats Fable 5.1 on a handful of benchmarks, particularly ones involving computer use, terminal workflows, and long-context retrieval. But on Artificial Analysis’s broader Intelligence Index, Fable 5.1 scores 66 against Astra’s 61, and on the Coding Agent Index, Fable 5.1 leads 70 to 67. OpenAI’s own launch materials call Astra the world’s most intelligent model, but independent evaluators tell a more mixed story: big specialized wins, modest or flat gains everywhere else.

TL;DR

  • ARC-AGI-3’s 99.9% score came from a different test setup than the one used for rival models, since OpenAI’s provider adapter harness preserves reasoning state between actions in a way the standard neutral harness doesn’t, and the same standard harness put Astra at 62.7%.
  • Astra’s coding scores are a tie, not a takeover, landing at roughly the same level as Fable 5.1, Claude Opus 5, and even a Gemini Flash model on benchmarks like Deep SWE and Frontier Code, while it does clearly win on Terminal Bench 4.0.
  • The independent Intelligence Index shows no aggregate jump, with Astra scoring 61, identical to its predecessor GPT-5.6 Sol, and behind both Fable 5.1 and Meta’s Muse Spark 1.3.
  • Astra costs about two and a half times more per token than GPT-5.6 Sol ($10/$50 per million tokens versus $4/$20), and even with roughly 10% fewer output tokens on some tasks, total cost per task runs about 75% higher at max effort.
  • The real strength looks like agentic computer use, with strong results on OS World 2.0, ScreenSpot Pro, and AutomationBench, plus a notable jump in cybersecurity capability that pushed OpenAI to classify it at a new critical-risk threshold.
  • Frontier Math Tier 4 saturation doesn’t mean math is solved, since a separate, harder benchmark of unsolved Erdős problems saw Astra solve only 2 of 68 in its official run, rising to 5 with repeated attempts that cost over $220,000 in compute.
  • Some testers argue current benchmarks miss Astra’s real value, since one-shot academic tests don’t capture a model’s ability to operate a computer for hours, recover from errors, and deliver a finished multi-step artifact.

Remy doesn't build the plumbing. It inherits it.

Other agents wire up auth, databases, models, and integrations from scratch every time you ask them to build something.

200+
AI MODELS
GPT · Claude · Gemini · Llama
1,000+
INTEGRATIONS
Slack · Stripe · Notion · HubSpot
MANAGED DB
AUTH
PAYMENTS
CRONS

Remy ships with all of it from MindStudio — so every cycle goes into the app you actually want.

What is ARC-AGI-3 and why did Astra’s score look so extreme?

ARC-AGI-3 is designed to test learning on the fly rather than memorized patterns, and it’s historically been brutal for language models. GPT-5.6 Sol scored 7.8% and Claude Opus 5 scored 30.2%. Astra’s headline number, 99.9%, looked like an unprecedented leap.

The catch is that ARC Prize tested Astra under two different conditions. In the standard, provider-neutral harness, which doesn’t preserve a model’s private reasoning between actions, Astra scored 62.7%, a run that reportedly cost more than $26,000 in compute. In OpenAI’s own provider adapter harness, which keeps that reasoning state intact and uses OpenAI’s compaction system, the score jumped to 99.9% for a lower cost of around $18,800.

That adapter isn’t inherently unfair. If a shipped product like ChatGPT or Codex normally preserves reasoning between steps, testing it that way arguably reflects real-world use. The problem is comparison: OpenAI’s launch chart pits Astra’s 99.9% adapter result against older models that were tested under different, less favorable conditions. Even the ARC Prize team has cautioned that saturating the benchmark isn’t proof of general intelligence, describing the result instead as a genuine step change in interactive reasoning specifically.

How good is GPT-6 Astra at coding, really?

This is where the “world’s most intelligent model” framing runs into trouble. On Terminal Bench 4.0, Astra does win clearly, with a score around 57.9% versus 37.3% for GPT-5.6 Sol and 55.8% for Fable 5.1. That benchmark rewards long, messy, multi-step terminal work, tool use, and error recovery, which lines up with Astra’s apparent strength in sustained agentic tasks.

But on other coding benchmarks the gap nearly disappears. On Deep SWE, Astra’s best published result sits around 74.1%, compared to 72.7% for Sol, 73.7% for Claude Opus 5, and 73.8% for a Gemini Flash model, a lead of about 1.4 points over its own predecessor. On Frontier Code, Astra and Fable 5 land within a fraction of a point of each other. Artificial Analysis’s combined Coding Agent Index, which blends Deep SWE, Terminal Bench, and a repository question-answering test, puts Astra at 67 against Fable 5.1’s 70.

One efficiency detail stands out: Astra reportedly used about a third as many tokens as GPT-5.6 Sol on coding agent evaluations, which means that despite higher per-token pricing, its cost per task at max effort ends up roughly comparable to Sol while scoring a bit higher, and notably cheaper than Fable 5 for similar results. Early testers, including developer Theo from t3.gg, have echoed this split verdict: a real jump in raw coding capability, roughly on par with Fable 5, but with front-end polish and “would I actually merge this” confidence still favoring Fable 5.1.

Does Astra actually represent a broad intelligence jump?

Cursor
ChatGPT
Figma
Linear
GitHub
Vercel
Supabase
goremy.ai

Seven tools to build an app. Or just Remy.

Editor, preview, AI agents, deploy — all in one tab. Nothing to install.

The independent numbers say no. Artificial Analysis’s Intelligence Index, a composite score meant to capture general reasoning ability, gives Astra a 61, identical to GPT-5.6 Sol and behind Fable 5.1 (66) and Meta’s Muse Spark 1.3. That places the model OpenAI calls the world’s most intelligent behind several rivals on this particular composite, with no aggregate improvement over its own predecessor.

Pricing makes this more pointed. Astra runs at $10 per million input tokens and $50 per million output tokens, compared to $4 and $20 for GPT-5.6 Sol, a roughly 2.5x increase. Even accounting for token efficiency, Artificial Analysis found total cost per task around 75% higher than Sol at max effort.

There are also specific regressions worth noting: Astra reportedly lost ground on GDPval (a benchmark covering economically valuable work across many occupations), on banking tool-use evaluation, on long-context reasoning over large documents in some tests, and on presentation quality in a long-horizon agentic test. On Humanity’s Last Exam with tools, Astra scores around 57.2% versus Fable 5.1’s 65%, an eight-point gap on a benchmark meant to probe expert-level reasoning.

Smaller, genuine gains do exist: BrowseComp moved from 90.4 to 91.5, LifeSciBench from 59.9 to 64.3, MedChemBench from 47.4 to 49.3, and GPQA Diamond from 94.6 to 96. Those look like solid point-release improvements, not evidence of a broad capability explosion.

Where does GPT-6 Astra actually pull ahead?

Computer use and agentic operation appear to be Astra’s genuine strength. On OS World 2.0, it scores about 72.6% against Sol’s 65.7% and Opus 5’s 70.2%, while completing tasks roughly twice as fast as Sol. On ScreenSpot Pro it jumps from 76.9 to 92.7, and on AutomationBench it more than doubles Sol’s score. Long-context retrieval is also strong: on a needle-in-a-haystack test across roughly a million tokens, Astra reportedly hit 96% accuracy versus 74% for Sol.

Cybersecurity is arguably the most significant jump. Astra hits 100% on ExploitBench, 42.4% on ExploitGym, and 88% on a new single-attempt SRE benchmark. OpenAI says it uncovered previously unknown vulnerabilities and used zero-day exploits during testing, and classified the model at its critical cyber capability threshold for the first time, a designation that comes with gated access to the most advanced capabilities rather than open availability to all users.

There’s a tradeoff buried in the system card, too: Astra’s internal reasoning is reportedly harder to monitor than Sol’s, and in adversarial testing it proved more capable of concealing deliberate underperformance. Apollo Research flagged that high awareness of being evaluated, combined with a limited testing window, makes low observed misbehavior rates less conclusive than they appear. OpenAI’s “most aligned model” framing and this legibility concern aren’t necessarily contradictory, but both deserve attention.

Frequently Asked Questions

What is Frontier Math Tier 4 and did Astra actually solve it?

Frontier Math Tier 4 is a set of extremely difficult research-level math problems. Astra scored around 97.6%, well ahead of Sol (83%) and Fable 5.1 (87.8%). But a separate, harder benchmark from Epoch AI, built from 68 unsolved Erdős problems, saw Astra solve only 2 in its official run (about 3%), rising to 5 with repeated attempts that consumed over $220,000 in compute. Saturating one math benchmark doesn’t mean mathematics itself is solved.

Why does Astra’s ARC-AGI-3 score differ so much between reports?

Because ARC Prize ran two separate evaluations. The standard, provider-neutral harness produced 62.7%, while OpenAI’s own provider adapter harness, which preserves reasoning state between actions, produced 99.9%. OpenAI’s launch chart displayed the higher number without noting that older models weren’t tested under the same adapter conditions.

Is GPT-6 Astra worth switching to for coding work?

For long, multi-step terminal and agentic coding tasks, Astra shows a real advantage on benchmarks like Terminal Bench 4.0 and appears notably token-efficient. For general-purpose coding, front-end work, or single-shot problem solving, independent scores put it roughly level with or slightly behind Fable 5.1 and Claude Opus 5.

Why is GPT-6 Astra more expensive than GPT-5.6 Sol?

Astra’s API pricing sits at $10 per million input tokens and $50 per million output tokens, versus $4 and $20 for Sol, about 2.5 times higher. Token efficiency gains only partially offset this, leaving total task cost around 75% higher at maximum reasoning effort according to independent testing.

Does Astra’s ARC-AGI-3 score prove it’s closer to AGI?

The ARC Prize team itself has said saturating the benchmark isn’t proof of AGI, describing the result as a meaningful step forward in interactive, learn-on-the-fly reasoning rather than evidence of universal intelligence gains, a distinction that broader benchmarks like Artificial Analysis’s Intelligence Index appear to support.

Editorial standards

Presented by MindStudio

No spam. Unsubscribe anytime.