Skip to main content
MindStudio
Pricing
BlogAbout
My Workspace
Qwen 3.8 MaxQwen benchmarksopen weight LLM

Qwen 3.8 Max Benchmarks: Where It Really Ranks vs Claude and GPT-5.6

Qwen 3.8 Max claims to trail only Gemini. Real DeepSWE and GPQA scores show a more mixed picture against GPT-5.6 and Opus.

MindStudio Team RSS
Qwen 3.8 Max Benchmarks: Where It Really Ranks vs Claude and GPT-5.6

What is Qwen 3.8 Max and what does it claim?

Qwen 3.8 Max is a 2.4 trillion parameter open weight language model from Alibaba’s Qwen team. When it was first teased, Qwen positioned it as second only to Gemini among available models. Now that it’s actually released with published benchmarks, that claim holds up in some areas and falls apart in others, depending entirely on which test you’re looking at. On raw knowledge and reasoning tasks it’s genuinely close to the top tier. On software engineering, the gap to the leading closed models is still wide.

TL;DR

  • Qwen 3.8 Max is a 2.4 trillion parameter open weight model, a big jump from its predecessor Qwen 3.7 Max.
  • On the DeepSWE benchmark, a widely watched coding test, it scored 56.6, well behind Gemini 5 at 70, GPT-5.6 Sol at 73, and even Opus 4.8 at 59.
  • On GPQA (Google Proof Question and Answers), it scored 92.6, putting it on par with Gemini and just below GPT-5.6 Sol, which is a genuinely strong showing.
  • The jump from Qwen 3.7 Max to 3.8 Max on DeepSWE (21.6 to 56.6) is one of the largest generation-over-generation gains reported this cycle, even if it still lags the top closed models.
  • At 2.4 trillion parameters, this is not a model you’ll run on a consumer GPU; it’s a cloud-scale open weight release, not a local one.
  • The “second only to Gemini” marketing line appears to be cherry-picked from favorable benchmarks rather than a claim that holds across the board.
  • For serious coding work, GPT-5.6 Sol, Gemini 5, and Opus still outperform it on the metric developers care about most.

Remy is new. The platform isn't.

Remy
Product Manager Agent
THE PLATFORM
200+ models 1,000+ integrations Managed DB Auth Payments Deploy
BUILT BY MINDSTUDIO
Shipping agent infrastructure since 2021

Remy is the latest expression of years of platform work. Not a hastily wrapped LLM.

How does Qwen 3.8 Max compare on coding benchmarks?

The benchmark most relevant to developers right now is DeepSWE, a software engineering test that correlates closely with how a model actually performs when asked to write, debug, or extend real code. This is the number people watch because it tends to predict how a model feels in day to day coding work, more than generic reasoning scores do.

Qwen 3.8 Max scored 56.6 on DeepSWE. For context:

  • Gemini 5: 70
  • GPT-5.6 Sol: 73
  • Opus 4.8: 59
  • Qwen 3.7 Max (previous generation): 21.6

That 21.6 to 56.6 jump is a real improvement, more than doubling the previous generation’s score. But it still lands below Opus 4.8, and well below both Gemini 5 and GPT-5.6 Sol. If coding is your primary use case, the DeepSWE numbers say Qwen 3.8 Max is a strong open weight option, not a frontier-matching one. It closes a lot of the gap with the previous generation without actually closing the gap with the current leaders.

How does it perform on reasoning and knowledge tasks?

The picture changes when you look at GPQA, a benchmark built from graduate-level science questions designed to be resistant to simple lookup or pattern matching. This is where Qwen 3.8 Max actually backs up part of its marketing claim.

It scored 92.6 on GPQA, putting it essentially level with Gemini and just a step behind GPT-5.6 Sol. That’s a legitimately strong result for an open weight model and suggests Qwen 3.8 Max is genuinely competitive when the task is answering hard questions and reasoning through problems, rather than writing and executing code.

This split matters. A model can be excellent at structured reasoning and knowledge recall while still lagging on the messier, more open-ended work of software engineering, where planning, tool use, and iterative debugging all come into play. Qwen 3.8 Max looks like exactly that kind of model: a strong generalist reasoner that isn’t yet a top-tier coding agent.

Is the “second only to Gemini” claim accurate?

Only partially, and only if you pick your benchmark carefully. On GPQA, yes, Qwen 3.8 Max is close to Gemini and ahead of most other competitors mentioned. On DeepSWE, the coding benchmark that most developers actually care about, it’s behind Gemini, GPT-5.6 Sol, and Opus 4.8. That’s a meaningfully different story than “second only to Gemini” implies.

This is a common pattern in AI model announcements: companies lead with the benchmark that makes their model look best, and it’s up to independent scrutiny to check whether that result generalizes. In this case, Qwen’s claim is defensible for reasoning tasks and misleading if you assume it applies to coding as well. Anyone choosing a model based on the marketing headline alone would be making a decision based on incomplete information.

Can you actually run Qwen 3.8 Max yourself?

REMY IS NOT
  • a coding agent
  • no-code
  • vibe coding
  • a faster Cursor
IT IS
a general contractor for software

The one that tells the coding agents what to build.

Technically yes, but not the way most people think of “running an open weight model.” At 2.4 trillion parameters, Qwen 3.8 Max is nowhere near small enough to run on a consumer GPU or even a high-end workstation. Open weight in this case means the weights are published and available for anyone with sufficient infrastructure, not that it’s practical to self-host on a laptop or single GPU rig.

For most people who want to actually try it, the accessible route is through Qwen’s own hosted chat interface, where you can log in and use Qwen 3.8 Max directly in the cloud. If your goal is running a model locally for privacy or cost reasons, Qwen 3.8 Max isn’t the right fit. Smaller open weight models like Google’s Gemma line or the GPT-OSS family are built for that use case, trading raw capability for something that actually fits on hardware you own.

That leaves Qwen 3.8 Max in an odd middle position: too large to self-host, but competing in the cloud against models like GPT-5.6 and Gemini that currently outperform it on the coding benchmark most people care about. For teams that specifically want an open weight model for licensing, fine-tuning, or data control reasons, it’s a legitimate option. For teams that just want the best coding assistant available and don’t care about openness, it’s not currently the top pick.

Is Qwen 3.8 Max worth using over Claude or GPT-5.6?

It depends on the job. For reasoning-heavy tasks, question answering, and general knowledge work, Qwen 3.8 Max’s GPQA score suggests it holds up well against the best closed models available. For software engineering, the DeepSWE numbers say otherwise: GPT-5.6 Sol, Gemini 5, and Opus 4.8 all outperform it, some by a meaningful margin.

The practical takeaway is that “open weight” and “state of the art” aren’t the same claim, and Qwen 3.8 Max’s release is a reminder to check the specific benchmark that matches your actual use case rather than trusting a single headline number.

Frequently Asked Questions

What is Qwen 3.8 Max’s parameter count?

Qwen 3.8 Max is a 2.4 trillion parameter open weight model, a significant scale-up from previous Qwen generations.

How does Qwen 3.8 Max score on DeepSWE compared to GPT-5.6 and Gemini?

It scored 56.6 on DeepSWE, behind Gemini 5 (70), GPT-5.6 Sol (73), and Opus 4.8 (59), though far ahead of its predecessor Qwen 3.7 Max (21.6).

Is Qwen 3.8 Max good at reasoning tasks?

Yes. It scored 92.6 on GPQA, putting it roughly on par with Gemini and just behind GPT-5.6 Sol, making it one of the stronger open weight models for reasoning and knowledge tasks.

Can I run Qwen 3.8 Max on my own computer?

Not practically. At 2.4 trillion parameters it requires cloud-scale infrastructure, not a consumer GPU. It’s accessible through Qwen’s hosted chat interface instead.

Is Qwen 3.8 Max really “second only to Gemini”?

That claim holds up on select reasoning benchmarks like GPQA, but not on coding benchmarks like DeepSWE, where several closed models currently outperform it.

Presented by MindStudio

No spam. Unsubscribe anytime.