Skip to main content
MindStudio
Pricing
BlogAbout
My Workspace
Nex-N2.5 benchmarksNex-N2.5 vs ClaudeOSWorld benchmark

Nex-N2.5 Benchmarks vs Claude Opus 5, GPT-5.6, and Kimi K3

Nex-N2.5-Pro and Max benchmark scores on SWE-Bench Pro, OSWorld, and BrowseComp, compared against Claude Opus 5, GPT-5.6, and Kimi K3.

Edited by Luis Chavez-Mattos, Director of Product RSS
Nex-N2.5 Benchmarks vs Claude Opus 5, GPT-5.6, and Kimi K3

What is Nex-N2.5 and how does it compare to Claude Opus 5?

Nex-N2.5 is an open-weight family of agentic AI models from Nex-AGI, released in three sizes: mini, Pro, and Max. On published benchmarks, Nex-N2.5-Max trails Claude Opus 5 on most coding tasks (65.7 vs 79.2 on SWE-Bench Pro) but beats it outright on BrowseComp (92.6 vs 90.8) and comes close on agentic benchmarks like Toolathlon Verified and AutomationBench. The gap is real on hard coding evals, but narrow to nonexistent on web browsing and tool-use tasks.

TL;DR

  • Nex-N2.5-Max, the largest model in the family, runs on a 1.6-trillion-parameter Mixture-of-Experts foundation and represents Nex-AGI’s first full post-training pass at trillion-parameter scale.
  • On SWE-Bench Pro, Nex-N2.5-Max scores 65.7 against Claude Opus 5’s 79.2, a sizable gap that puts Nex behind GPT-5.6, Kimi K3, and GLM-5.3 too on this particular test.
  • On BrowseComp, a web-research benchmark, Nex-N2.5-Max actually posts the top score in the table at 92.6, ahead of Claude Opus 5, GPT-5.6, and Kimi K3.
  • On Toolathlon Verified, an agentic tool-use benchmark, Nex-N2.5-Max ties Kimi K3 at 76.5, just half a point behind Claude Opus 5’s 76.5-adjacent lead (Claude and Kimi are tied at the top; Nex-Max sits at 74.7).
  • The mini and Pro variants are meaningfully weaker than Max across the board, which matters if you’re picking a model based on hosted pricing tiers or self-hosting hardware limits.
  • On computer-use benchmarks like OSWorld-Verified and OSWorld-2, Nex-N2.5-Pro lags behind Claude Opus 5, GPT-5.6, and Kimi K3, though it leads on the narrower OSWorld-G grounding test.
  • All three sizes ship as open weights on Hugging Face and ModelScope, with hosted access also available through OpenRouter, unlike the closed Claude, GPT, and Kimi models it’s benchmarked against.

Remy is new. The platform isn't.

Remy
Product Manager Agent
THE PLATFORM
200+ models 1,000+ integrations Managed DB Auth Payments Deploy
BUILT BY MINDSTUDIO
Shipping agent infrastructure since 2021

Remy is the latest expression of years of platform work. Not a hastily wrapped LLM.

What benchmarks did Nex-AGI use to test Nex-N2.5?

Nex-AGI evaluated Nex-N2.5 across four categories: coding, agentic workflows, computer use, and multimodal understanding. The coding suite includes Terminal-Bench 2.1, SWE-Bench Pro, and DeepSWE v1.1, all run through Nex-AGI’s own NexAU evaluation harness. The agentic suite covers AutomationBench, Toolathlon Verified, GDPval-AA v2, Job Bench, and BrowseComp. Computer-use and browser-use tasks (OSWorld-Verified, OSWorld-2, WebTest, WebArena-Verified, OSWorld-G) run through a separate harness called NexCUA, which Nex-AGI says will be open-sourced.

Comparison models in the published tables include Claude Opus 5, GPT-5.6 Sol, Kimi-K3, GLM-5.3, DeepSeek-V4-Pro-0813, and Qwen3.8-Max. Scores for those models come from a mix of official leaderboards, published provider reports, and Nex-AGI’s own evaluation runs where no public number existed. That mixed sourcing is worth keeping in mind: not every number in the comparison table was measured under identical conditions.

How does Nex-N2.5-Max perform on coding benchmarks?

Coding is where Nex-N2.5-Max shows the clearest weakness relative to closed frontier models. On Terminal-Bench 2.1, it scores 86.1, behind Claude Opus 5 (89.1), GPT-5.6 (88.8), Kimi-K3 (88.3), and GLM-5.3 (88.2), though ahead of Qwen3.8-Max (86.6). On SWE-Bench Pro, a harder real-world software engineering benchmark, the gap widens: Nex-N2.5-Max hits 65.7 against Claude Opus 5’s 79.2, and even trails Qwen3.8-Max (67.7). DeepSWE v1.1 shows the same pattern, with Nex-Max at 65.6 versus Claude at 73.7 and GPT-5.6 at 72.7.

The pattern across all three coding benchmarks is consistent: Claude Opus 5 leads, GPT-5.6 and Kimi-K3 sit close behind it, and Nex-N2.5-Max trails by roughly 7 to 14 points depending on the test. For teams doing heavy autonomous software engineering work, that gap is large enough to matter.

Where does Nex-N2.5 actually lead the field?

The agentic and web-research benchmarks tell a different story. On BrowseComp, which tests an agent’s ability to research and answer questions using live web browsing, Nex-N2.5-Max posts 92.6, the highest score in the published comparison, edging out Kimi-K3 (91.2), Claude Opus 5 (90.8), and GPT-5.6 (90.4).

On Toolathlon Verified, a benchmark for verified tool-use tasks, Nex-N2.5-Max ties with Kimi-K3 at 76.5 for second place, just behind Claude Opus 5 at the same score (the table lists both Claude and Kimi at 76.5 as joint leaders). On AutomationBench, Nex-N2.5-Max scores 50.2, essentially matching Claude Opus 5’s 50.3 and beating GPT-5.6, Kimi-K3, and GLM-5.3.

On OSWorld-G, a grounding-focused computer-use test, Nex-N2.5-Pro (not even the largest variant) scores 87.4, the top result in that row, ahead of Qwen3.8-Max (84.9) and GLM-5.3-Flash (83.3). That’s a notable result given Pro is a smaller model than Max.

Where Nex-N2.5 falls short in the agentic category is Job Bench and GDPval-AA v2, both of which measure broader knowledge-work and productivity task completion. Claude Opus 5 leads comfortably on both: 65.7 versus Nex-Max’s 53.6 on Job Bench, and 1831 versus 1713 on GDPval-AA v2.

Is Nex-N2.5 competitive on computer use and multimodal tasks?

Remy doesn't write the code. It manages the agents who do.

R
Remy
Product Manager Agent
Leading
Design
Engineer
QA
Deploy

Remy runs the project. The specialists do the work. You work with the PM, not the implementers.

Mixed. On OSWorld-Verified, a benchmark for controlling a full desktop environment, Nex-N2.5-Pro scores 82.2, below Claude Opus 5 (83.4), GPT-5.6 (83.2), Kimi-K3 (84.8), and Qwen3.8-Max, which tops the row at 86.1. On the harder OSWorld-2 variant, the gap widens: Nex-N2.5-Pro hits 56.4 against Claude Opus 5’s 68.3.

On web-agent tasks, results are similarly split. WebArena-Verified has Kimi-K3 on top at 71.6, with Nex-N2.5-Pro at 67.6. Vision2Web, which evaluates GUI agents across frontend, webpage, and website categories, shows GPT-5.6 clearly ahead at 79.8 versus Nex-N2.5-Pro’s 68.2. SWE-MM, a multimodal software engineering benchmark, again favors Claude Opus 5 (59.4) over Nex-N2.5-Pro (38.2).

The one area where Nex-N2.5-Pro leads outright is OSWorld-G, the grounding benchmark mentioned above. Nex-AGI’s own notes flag that grounding coordinates in these tests are normalized to a 0 to 1000 scale and run through their NexCUA harness, which is planned for open-source release but isn’t public yet, so independent replication of these specific numbers isn’t currently possible.

Should you use Nex-N2.5 instead of a closed model?

That depends on what you’re optimizing for. If the priority is raw coding capability on hard software engineering tasks, the benchmark data points toward Claude Opus 5, GPT-5.6, or Kimi-K3 over Nex-N2.5, based on the SWE-Bench Pro and DeepSWE gaps. If the workload leans toward web research, browsing, and tool-calling agents, Nex-N2.5-Max’s BrowseComp and Toolathlon scores put it in the same tier as the closed frontier models, sometimes ahead.

The bigger differentiator may not be the benchmark scores at all: Nex-N2.5 ships as open weights on Hugging Face and ModelScope, with deployment instructions for multi-node H200 clusters (Max) down to dual H100 setups (mini). Claude Opus 5, GPT-5.6, and Kimi-K3 are closed, accessed only through hosted APIs. For teams that need to self-host, audit model weights, or fine-tune on private infrastructure, that’s a structural advantage no benchmark table captures, even where Nex-N2.5 trails on raw scores.

Frequently Asked Questions

What is the difference between Nex-N2.5-mini, Pro, and Max?

Mini and Pro build on the multimodal foundation of the earlier Nex-N2 model, with improvements to computer use, browsing, and visual grounding. Max is a separate, larger foundation model built on a 1.6-trillion-parameter Mixture-of-Experts architecture, and it’s text-only rather than natively multimodal. Across nearly every benchmark, Max outperforms Pro, which outperforms mini.

Does Nex-N2.5 beat Claude Opus 5 on any benchmark?

Yes, on BrowseComp (web research), Nex-N2.5-Max scores 92.6 versus Claude Opus 5’s 90.8. It also roughly matches Claude on AutomationBench (50.2 vs 50.3) and Toolathlon Verified (74.7 vs 76.5). It trails Claude clearly on SWE-Bench Pro, DeepSWE, Job Bench, GDPval-AA v2, OSWorld-2, and SWE-MM.

Can I run Nex-N2.5 on my own hardware?

Nex-AGI publishes deployment instructions using a customized sglang server. Nex-N2.5-Max requires a multi-node setup (documented for two nodes of 8 H200 GPUs each). Pro runs on a single node of 8 H100 GPUs. Mini runs on a single node with 2 H100 GPUs. All three are available as open weights on Hugging Face and ModelScope.

What is BrowseComp and why does it matter here?

VIBE-CODED APP
Tangled. Half-built. Brittle.
AN APP, MANAGED BY REMY
UIReact + Tailwind
APIValidated routes
DBPostgres + auth
DEPLOYProduction-ready
Architected. End to end.

Built like a system. Not vibe-coded.

Remy manages the project — every layer architected, not stitched together at the last second.

BrowseComp is a benchmark that tests an agent’s ability to research and answer questions by browsing the live web, often requiring multi-step search and information synthesis. It’s one of the few benchmarks in this comparison where Nex-N2.5-Max posts the top score among all listed models, suggesting the model’s web-browsing and context-compaction strategies are competitive with, or ahead of, closed frontier models in this specific task type.

Is Nex-N2.5 good for coding agents?

It depends on the task’s difficulty. On Terminal-Bench 2.1, Nex-N2.5-Max scores close to the leaders (86.1 vs Claude’s 89.1). But on harder, real-world engineering benchmarks like SWE-Bench Pro, the gap to Claude Opus 5, GPT-5.6, and Kimi-K3 grows substantially, suggesting Nex-N2.5 is more competitive on shorter or more contained coding tasks than on complex, multi-file software engineering work.

Editorial standards

Presented by MindStudio

No spam. Unsubscribe anytime.