Claude Opus 5 Benchmarks: The Numbers Anthropic Didn't Headline
Claude Opus 5's independent benchmark results, from Arc AGI 3 to IMO 2026, and how it actually stacks up against Claude Fable 5 in practice.

Is Claude Opus 5 actually more capable than Claude Fable 5?
Anthropic itself says no, describing Fable 5 as its most generally capable model. But independent and third-party benchmarks tell a messier story. On several work-oriented benchmarks, including agentic business workflows, computer use, and multidisciplinary reasoning, Opus 5 matches or edges out Fable 5, and does so at roughly half the cost. On raw intelligence benchmarks like Arc AGI 3 and IMO 2026, Opus 5 posts numbers that look less like incremental progress and more like a step change. Whether it’s “the world’s most powerful model” depends entirely on what you’re measuring.
TL;DR
- Work benchmarks favor Opus 5 on value, not always on raw score. Across computer use, agentic business workflows, and multidisciplinary reasoning, Opus 5 performs close to or ahead of Fable 5 while costing about half as much per token.
- The Arc AGI 3 jump is the headline number nobody at Anthropic put on a slide. Opus 5 scored 30% on Arc AGI 3, roughly three times the next best model, a jump that happened in about two months rather than across a normal model generation.
- Opus 5 hit a perfect IMO 2026 score without tool use. It solved 42 out of 42 problems at gold medal level using no agent harness and no external tools, well above the historical gold medal threshold of 29 out of 42.
- Independent economic-impact benchmarks show a tighter race. On the Vales Index and Apex Agents indices, which weight tasks by real economic value rather than academic difficulty, top frontier models including Opus 5 and Fable 5 sit within about a percentage point of each other.
- More reasoning effort doesn’t always mean a better score. Anthropic’s own data shows Opus 5’s performance on Frontier code peaks at medium reasoning effort (around 53%) and actually gets worse, and more expensive, at higher settings.
- On a saturation-resistant meta-benchmark, Opus 5 trails. The Epoch Capabilities Index, which aggregates dozens of benchmarks using item response theory, places Opus 5 behind both Fable 5 and GPT-5.6 in overall composite capability.
- Pricing stayed flat. Opus 5 costs 5 dollars per million input tokens and 25 dollars per million output tokens, the same as its predecessor, Opus 4.8.
Remy is new. The platform isn't.
Remy is the latest expression of years of platform work. Not a hastily wrapped LLM.
How does Opus 5 perform on real work tasks?
The clearest pattern in Opus 5’s benchmark results is that it’s optimized for value rather than pure top-line intelligence. Anthropic’s released benchmarks around computer use, agentic business workflows, multidisciplinary reasoning, and agentic coding all show Opus 5 performing at or near the level of Fable 5, but at a noticeably lower price point.
Independent testing backs this up in part. The Apex Agents benchmark, built by Mercor to measure whether frontier models can do economically valuable professional work, tests things like investment banking analysis, management consulting, legal associate tasks, and primary care physician scenarios using more than 200 real tasks graded against expert-written rubrics. Opus 5 leads this benchmark at 43.5%, only marginally ahead of Fable 5. But given the price gap, that marginal lead still tips the practical decision toward Opus 5 for most workflows.
The picture flips slightly on Apex SWE, the software-engineering-focused version of the benchmark built with Cognition, where Fable 5 leads by a small margin. On SWE-bench Verified, Opus 5 scores around 97%, putting it in a tight cluster with other frontier models where differences of a percentage point or two are the norm rather than a meaningful gap.
The Vales Index, a proprietary benchmark that weights model performance by economic sector contribution to US GDP rather than treating all tasks as equal, shows Opus 5 slightly behind Fable 5. But again, the overall spread across top models on Vales is close to 1%, meaning the choice between models often comes down to cost and integration rather than capability alone.
Why does the Arc AGI 3 result matter so much?
Arc AGI 3 is designed to test something most benchmarks don’t: performance on problems a model has never seen in any form before. Where many benchmarks are arguably variations on familiar problem types, Arc AGI 3 uses genuinely novel puzzle environments, making it a proxy for general reasoning rather than pattern memorization.
Opus 5 scored 30% on this benchmark, about three times the score of the next best model. That jump reportedly happened within roughly two months of the prior model version, which is unusually fast for this kind of gain. Arc Prize, the organization behind the benchmark, reported that Opus 5 achieved 100% on five previously unbeaten environments, fully solving four of them at or above human-level efficiency, and in at least one case, identified a hidden mathematical rule within a visual puzzle on its own, using it to solve something no previous model had solved.
Analysis of Opus 5’s Arc AGI 3 replays found a clear split: simple games became winnable, but complex games remained largely unsolved, with the model fully completing about 20% of semi-private games and scoring zero on another 20%. When it did win, it showed effective goal recognition, hypothesis exploration, and the ability to carry context forward across steps, capabilities that matter for agentic tasks in unfamiliar environments, not just puzzle solving.
What happened with the IMO 2026 result?
Plans first. Then code.
Remy writes the spec, manages the build, and ships the app.
Claude Opus 5 reportedly achieved a perfect 42 out of 42 score on IMO 2026 problems, a gold medal level result, without using an agent harness or external tools. According to the reported methodology, the model was prompted to produce a rigorous, self-contained proof for each problem, with a 256,000 token output limit and adaptive thinking set to maximum. Attempts that ran out of output space were resampled at lower thinking effort, a fix that was needed for one solution.
A perfect score is meaningfully above the historical gold medal threshold of 29 out of 42, and doing it without tool assistance suggests the underlying reasoning ability, not just tool orchestration, has improved substantially. This result lines up with the Arc AGI 3 jump as evidence that Opus 5’s core reasoning capability moved more than a routine version bump would suggest.
Does more “thinking” always make Opus 5 better?
No, and this is one of the more practically useful findings buried in the benchmark data. Anthropic’s own reporting shows that pushing Opus 5’s reasoning effort beyond medium can decrease performance rather than improve it. On Frontier code specifically, the model peaks at medium reasoning effort around 53%, then gets both worse and more expensive as effort is cranked higher.
This runs counter to the usual assumption that longer reasoning chains produce better answers. The likely explanation is overthinking: for problems that have a fairly direct solution path, extended reasoning can lead the model to second-guess itself or wander into unnecessary complexity. Vales AI’s own testing found high reasoning effort produced the most consistent results across its benchmark suite, while Anthropic’s internal data pointed to medium as the sweet spot on coding tasks specifically. The practical takeaway is that setting reasoning effort to maximum by default is often just spending more money for a worse result.
How does Opus 5 compare on safety and biosecurity benchmarks?
On alignment testing, Opus 5 is reported as showing largely standard behavior, with some minor unaligned actions during testing that researchers characterized as not particularly concerning, in line with typical patterns seen in other frontier models.
On biology and virology risk benchmarks, Opus 5 scored 0.68, placing it in second place behind Meta’s Spark 1.1, a comparatively underrated model in this specific category. One notable observation from AI safety commentary is that despite performing comparably to Fable 5 on virology-related tasks, Opus 5 apparently doesn’t require the same bio-safety classifiers and synthesis-screening safeguards that Fable 5 uses, suggesting a different underlying risk profile even at similar task performance.
Where does Opus 5 rank on a saturation-resistant benchmark?
The Epoch Capabilities Index offers a useful counterweight to the highlight-reel numbers. Built by Epoch AI, it aggregates performance across dozens of individual benchmarks into a single composite score using item response theory, the same statistical approach used in educational testing to separate task difficulty from test-taker ability. Its purpose is to keep distinguishing model capability even after top models start saturating individual benchmarks like SWE-bench.
On this index, Opus 5 sits behind both Fable 5 and GPT-5.6. That doesn’t contradict the Arc AGI 3 or IMO results so much as complicate them: Opus 5 appears to have made an unusually large leap in specific reasoning-heavy, novel-problem domains while remaining behind on aggregate, broad-based capability measurement.
Frequently Asked Questions
Is Claude Opus 5 better than Claude Fable 5?
One coffee. One working app.
You bring the idea. Remy manages the project.
It depends on the task. On work-oriented benchmarks like agentic workflows and computer use, Opus 5 performs close to or ahead of Fable 5 at a lower price. On broad aggregate capability measures like the Epoch Capabilities Index, Fable 5 still leads.
What is Arc AGI 3 and why did Opus 5’s score matter?
Arc AGI 3 tests reasoning on novel puzzle environments a model hasn’t encountered before. Opus 5 scored 30%, about three times the next best model, and solved puzzle types no prior model had solved.
Did Opus 5 really get a perfect IMO score?
Reported results show Opus 5 scored 42 out of 42 on IMO 2026 problems without using tools or an agent harness, above the standard gold medal threshold of 29 out of 42.
Should reasoning effort always be set to maximum for Opus 5?
No. Anthropic’s own data shows performance on coding benchmarks peaks at medium reasoning effort and can decline at higher settings, while costing more.
How much does Claude Opus 5 cost to use?
Opus 5 is priced at 5 dollars per million input tokens and 25 dollars per million output tokens, unchanged from Opus 4.8.

