Skip to main content
MindStudio
Pricing
BlogAbout
My Workspace
Text Generation Model

Qwen3 235B

Qwen3 235B is a 235-billion-parameter mixture-of-experts text generation model from Qwen with a 262,144-token context window.

PublisherQwen
TypeText
Context Window262,144 tokens
ReleasedJuly 2025
Input$0.09/MTok
Output$0.55/MTok
ProviderDeepInfra

Large mixture-of-experts model for complex tasks

Qwen3 235B is a large language model developed by Qwen, the AI research team at Alibaba Cloud. It uses a mixture-of-experts (MoE) architecture with 235 billion total parameters and 22 billion active parameters per forward pass, released in July 2025 under the model identifier Qwen3-235B-A22B-Instruct-2507. The model is served through DeepInfra and supports a context window of up to 262,144 tokens, making it suitable for tasks involving long documents or extended conversations.

Qwen3 235B is designed for instruction-following tasks including reasoning, coding, summarization, and multilingual text generation. The MoE design activates only 22 billion parameters at inference time, which allows the model to handle complex tasks while keeping compute requirements lower than a dense model of equivalent total parameter count. It is well suited for developers and researchers working on applications that require long-context understanding or nuanced instruction following.

What Qwen3 235B supports

Long Context Window

Supports up to 262,144 tokens of context, enabling processing of long documents, codebases, or extended multi-turn conversations in a single pass.

Mixture-of-Experts Architecture

Uses a sparse MoE design with 235B total parameters but only 22B active per forward pass, reducing per-token compute relative to a dense model of the same size.

Instruction Following

Fine-tuned as an instruct model to follow complex, multi-step user instructions across a wide range of task types.

Code Generation

Capable of writing, explaining, and debugging code across multiple programming languages as part of its general instruction-following training.

Multilingual Text Generation

Supports text generation across multiple languages, consistent with the broader Qwen3 model family's multilingual training scope.

Reasoning Tasks

Handles multi-step reasoning problems in areas such as mathematics, logic, and analytical question answering.

Ready to build with Qwen3 235B?

Get Started Free

Benchmark scores

Scores represent accuracy — the percentage of questions answered correctly on each test.

BenchmarkWhat it testsScore
MMLU-ProExpert knowledge across 14 academic disciplines76.2%
GPQA DiamondPhD-level science questions (biology, physics, chemistry)61.3%
MATH-500Undergraduate and competition-level math problems90.2%
AIME 2024American math olympiad problems32.7%
LiveCodeBenchReal-world coding tasks from recent competitions34.3%
HLEQuestions that challenge frontier models across many domains4.7%
SciCodeScientific research coding and numerical methods29.9%

Common questions about Qwen3 235B

What is the context window size for Qwen3 235B?

Qwen3 235B supports a context window of 262,144 tokens, which also matches its maximum response size.

How many parameters does Qwen3 235B have?

The model has 235 billion total parameters. Due to its mixture-of-experts architecture, only 22 billion parameters are active during each forward pass.

Who developed Qwen3 235B and when was it released?

Qwen3 235B was developed by Qwen, the AI team at Alibaba Cloud, and released in July 2025.

What provider serves Qwen3 235B on MindStudio?

Qwen3 235B is served through DeepInfra on MindStudio.

Does Qwen3 235B support image or video inputs?

Based on the available metadata, Qwen3 235B is a text generation model and image or video input support is not confirmed for this version.

What is the pricing for Qwen3 235B?

Pricing information for Qwen3 235B is not published in the current metadata. Check MindStudio or DeepInfra directly for current pricing details.

What people think about Qwen3 235B

Community reception on r/LocalLLaMA has been broadly positive, with the original Qwen3 release thread accumulating nearly 2,000 upvotes and over 430 comments. Users frequently highlight the model's coding and reasoning benchmark scores, as well as the efficiency of its MoE architecture that activates only 22B parameters at inference time.

A notable discussion thread from August 2025 focused on the addition of 1M token context support for the 2507 variants, which generated significant interest among users working with long documents. Some community members also noted the distinction between the instruct and thinking variants, with discussion around which version to use for different task types.

View more discussions →

Parameters & options

Max Temperature1
Max Response Size262,144 tokens

Start building with Qwen3 235B

No API keys required. Create AI-powered workflows with Qwen3 235B in minutes — free.