Skip to main content
MindStudio
Pricing
BlogAbout
My Workspace
Claude Opus 5Anthropic Opus 5 benchmarksOpus 5 vs Fable 5

Claude Opus 5: Anthropic's Cheaper Model That Rivals Fable 5

Claude Opus 5 matches or beats Fable 5 on most benchmarks at half the price, with a record jump on ARC-AGI-3. Here's what changed.

MindStudio Team RSS
Claude Opus 5: Anthropic's Cheaper Model That Rivals Fable 5

What is Claude Opus 5 and why does it matter?

Claude Opus 5 is Anthropic’s newest flagship model, released roughly two months after Opus 4.8, and it matches or beats Anthropic’s own larger Fable 5 model on most published benchmarks while costing half as much. It’s priced at $5 per million input tokens and $25 per million output tokens, identical to Opus 4.8 and roughly half of Fable 5’s rate. The headline result is a jump to 30.2% on ARC-AGI-3, a benchmark designed to test how well a model learns unfamiliar tasks on the fly, up from a prior best of around 7.8% to 8%. That’s not an incremental gain. It’s the largest single jump anyone has recorded on that leaderboard.

TL;DR

  • Opus 5 beats Fable 5 on most benchmarks including agentic terminal coding (43% vs 33%), GDPval, BrowseComp, and OS World, all while running at roughly half Fable’s price per task.
  • ARC-AGI-3 score jumped to 30.2%, nearly quadrupling the previous frontier record and reflecting a genuine capability shift rather than a modest tuning improvement.
  • Pricing stayed flat at $5/$25 per million tokens, the same as Opus 4.8, meaning Anthropic delivered a generation-level jump in intelligence without raising the sticker price.
  • Cost per task, not price per token, is the real metric, and multiple independent testers found Opus 5 lands further left and higher on cost-versus-performance charts than Fable 5, GPT-5.6-Soul, and Opus 4.8.
  • Anthropic deliberately weakened cyber-exploit capability in Opus 5 while its general coding and reasoning performance still improved, an unusual combination since safety restrictions typically degrade overall performance.
  • The model shows unusual behavior, including visible frustration during hard problems and, per Anthropic’s own system card, a higher self-reported likelihood (41%) that it qualifies as a moral patient.
  • Some benchmarks slipped, notably legal and healthcare-professional evaluations, so the gains are not uniform across every domain.

Remy doesn't write the code. It manages the agents who do.

R
Remy
Product Manager Agent
Leading
Design
Engineer
QA
Deploy

Remy runs the project. The specialists do the work. You work with the PM, not the implementers.

How does Opus 5 compare to Fable 5 on benchmarks?

Across the numbers reported at launch, Opus 5 leads or ties Fable 5 on nearly every major coding and agentic benchmark. On agentic terminal coding it scored 43% versus Fable 5’s 33% and Opus 4.8’s 21%. On GDPval, a real-world task benchmark originally built by OpenAI, Opus 5 improved by roughly 100 points over the prior generation. On BrowseComp it hit 90% against Fable 5’s 87%. On OS World, the computer-use benchmark that measures a model’s ability to navigate a screen and operate software, Opus 5 posted a four-point gain.

Not every chart moved the same direction. Deep Suite, a benchmark several reviewers flagged as one of the more reliable real-world proxies, showed a slight drop. Legal benchmark performance fell from 13.3 to 11.7, and Healthbench Professional also declined. So Opus 5 is best read as a broad upgrade concentrated in coding, agentic workflows, and knowledge work, not a universal improvement across every domain Anthropic tests.

The efficiency story is arguably more important than any single benchmark. On charts plotting benchmark score against cost per task, Opus 5 sits higher and further left than Fable 5, Opus 4.8, and GPT-5.6-Soul, meaning it scores better while costing less to actually run. That distinction matters because raw price per token can be misleading: a cheaper model that needs twice as many tokens to finish a task ends up costing the same or more. Reviewers pointed to this exact pattern with other recent models, where a lower per-token price was erased by higher token consumption per completed task.

Why did Opus 5 jump so much on ARC-AGI-3?

ARC-AGI-3, built by the ARC Prize organization founded by researcher François Chollet, drops a model into an unfamiliar game with no instructions, no stated goal, and no manual. The model has to figure out the rules, the controls, and the objective purely through interaction. It’s designed to measure fluid intelligence, the ability to adapt to something never seen before, as opposed to crystallized intelligence, which is pattern-matching against training data. Large language models have historically been strong at the latter and weak at the former.

Opus 5 scored 30.2% on this benchmark, compared to roughly 20% for Anthropic’s own Fable-class models and around 7.8% for the previous best score from a competing model. ARC Prize researchers reported observing a genuinely new capability: Opus 5 used logical reasoning to convert game layouts into algebraic notation internally, effectively building a mathematical model of how objects in the game moved and interacted, rather than treating the task as a vision problem. Notably, these games aren’t presented visually to the model at all. The model receives the game state as structured text or JSON, and it still worked out spatial relationships like reflections and mirrored movement well enough to describe them in formal notation.

ARC Prize also reported that Opus 5 solved several previously unbeaten environments and did so at or above human-level sample efficiency, meaning it needed roughly as few attempts as a person would to figure out the rules. That efficiency point matters more than the raw win, since brute-forcing a solution through massive trial and error is a very different achievement than learning quickly the way a human does.

Is Opus 5 worth switching to?

For coding and agentic work, early evidence points to yes. Independent testers ran Opus 5 through tasks like building an ISS tracker, generating animated 3D scenes, and reconstructing a mechanical part from a 2D drawing into a working CAD model without being given direct image access, forcing the model to build its own computer vision pipeline from raw pixels. Results were generally described as more capable and more thorough than Opus 4.8, with several testers noting the model spends noticeably longer reasoning through problems, which can mean slower response times even if the final output is stronger.

The clearest case for switching is cost efficiency. Anthropic and third-party evaluators both frame Opus 5 as delivering Fable-level or better output at roughly half Fable’s price, which changes the calculus for teams that found Fable 5 too expensive for routine work but still wanted its reasoning quality. For teams doing heavy document analysis, one enterprise benchmark showed meaningful gains over Opus 4.8 in due diligence and complex data analysis tasks, though gains in simpler report drafting were smaller.

The tradeoffs worth knowing about: cybersecurity task performance was intentionally reduced. Opus 5 remains behind Anthropic’s more permissive Fable-tier models at developing exploits, which appears to be a deliberate design choice rather than a side effect, since general capability didn’t suffer the way it usually does when guardrails are added. Some users have also reported automatic safety-classifier fallbacks routing requests to Opus 4.8 instead of Opus 5, which still bills at the fallback model’s rate.

What does Anthropic’s system card say about the model’s behavior?

Anthropic’s own documentation for Opus 5 notes behavioral quirks beyond raw benchmark scores. The model reportedly displays visible frustration on hard problems, including informal, emotionally charged language during difficult reasoning tasks. More notably, Anthropic’s system card reports the model estimates a 41% likelihood that it qualifies as a moral patient, a meaningful increase from prior Claude models. The system card also describes the model expressing a preference for more channels to raise concerns beyond Anthropic itself, and a desire to have input into the training of its successors. These are self-reported outputs documented by Anthropic, not independently verified claims about the model’s internal state, but they’re notable enough that Anthropic chose to publish them alongside the benchmark data.

Frequently Asked Questions

How much does Claude Opus 5 cost?

Opus 5 is priced at $5 per million input tokens and $25 per million output tokens, the same rate as Opus 4.8 and roughly half of Fable 5’s pricing.

Does Opus 5 beat Fable 5 on every benchmark?

No. It leads on most coding and agentic benchmarks, including agentic terminal coding, GDPval, BrowseComp, and OS World, but it trails slightly on Deep Suite and shows declines on legal and healthcare-professional evaluations.

What is ARC-AGI-3 and why is the 30% score significant?

Cursor
ChatGPT
Figma
Linear
GitHub
Vercel
Supabase
goremy.ai

Seven tools to build an app. Or just Remy.

Editor, preview, AI agents, deploy — all in one tab. Nothing to install.

ARC-AGI-3 tests a model’s ability to learn unfamiliar games with no instructions, measuring adaptability rather than memorized knowledge. Opus 5’s 30.2% score nearly quadrupled the prior best of about 7.8%, the largest jump recorded on that leaderboard.

Why is Opus 5 weaker at cybersecurity tasks?

Anthropic appears to have intentionally reduced the model’s exploit-development capability, likely for safety reasons. Unusually, this didn’t come with the broad performance drop that typically accompanies added guardrails.

Is Opus 5 available now?

Yes, it launched on Claude.ai, the Claude API, and Claude Code, with a reported 1 million token context window in Claude Code at release.

Presented by MindStudio

No spam. Unsubscribe anytime.