Nemotron 3 Super 120B
Nemotron 3 Super 120B is a text generation model from Nvidia with a 1,000,000-token context window and selectable reasoning.
Large sparse MoE model with 1M token context
Nemotron 3 Super 120B (NVIDIA-Nemotron-3-Super-120B-A12B) is a large language model developed by Nvidia and released in March 2026. It uses a mixture-of-experts architecture with 120 billion total parameters but only 12 billion active parameters per forward pass, which allows it to handle large workloads while keeping compute requirements lower than a dense model of equivalent total size. The model supports a context window of up to one million tokens, making it suited for tasks that require processing very long documents or extended conversations.
The model is designed for text generation tasks and includes a selectable reasoning mode, giving users the option to toggle reasoning behavior depending on the task at hand. With a maximum response size of 16,384 tokens, it can produce detailed, long-form outputs in a single generation. It is available through DeepInfra on MindStudio and is well suited for use cases such as document summarization, long-context question answering, and complex instruction following.
What Nemotron 3 Super 120B supports
Long Context Window
Processes up to 1,000,000 tokens in a single context, enabling analysis of very long documents or extended multi-turn conversations.
Selectable Reasoning
Offers a toggle to enable or disable reasoning mode, letting users choose between standard generation and step-by-step reasoning depending on the task.
Mixture-of-Experts Architecture
Uses 120 billion total parameters with only 12 billion active per forward pass, reducing compute cost compared to a fully dense model of the same scale.
Long-Form Text Generation
Generates responses of up to 16,384 tokens per completion, supporting detailed reports, summaries, and extended structured outputs.
Instruction Following
Trained to follow complex, multi-step instructions in a chat format, making it suitable for agentic workflows and detailed task completion.
Ready to build with Nemotron 3 Super 120B?
Get Started FreeBenchmark scores
Scores represent accuracy — the percentage of questions answered correctly on each test.
| Benchmark | What it tests | Score |
|---|---|---|
| AIME 2025 | American math olympiad problems (2025) | 90.2% |
| GPQA Diamond | PhD-level science questions (biology, physics, chemistry) | 82.7% |
| SWE-bench Verified | Real GitHub issues requiring multi-file code fixes | 60.5% |
Common questions about Nemotron 3 Super 120B
What is the context window size for Nemotron 3 Super 120B?
Nemotron 3 Super 120B supports a context window of 1,000,000 tokens, allowing it to process very long documents or extended conversations in a single session.
How many parameters does this model have?
The model has 120 billion total parameters but activates only 12 billion per forward pass, as indicated by its full model ID (NVIDIA-Nemotron-3-Super-120B-A12B), reflecting its mixture-of-experts design.
What is the maximum response length?
The model can generate up to 16,384 tokens in a single response.
What does the reasoning toggle do?
The reasoning input is a selectable option that allows users to switch reasoning behavior on or off, enabling step-by-step reasoning for complex tasks or standard generation for simpler ones.
When was Nemotron 3 Super 120B released?
Nemotron 3 Super 120B was released in March 2026 and became available on MindStudio on March 16, 2026.
Does the model support image or video inputs?
No. Based on the available metadata, Nemotron 3 Super 120B does not support image or video inputs and is limited to text-based interactions.
What people think about Nemotron 3 Super 120B
Community reception on r/LocalLLaMA and r/singularity has been generally positive, with users highlighting the model's 1M token context window, fast inference due to its 12B active parameter design, and suitability for local deployment on Blackwell hardware. The hybrid SSM LatentMoE architecture and open-weight availability were frequently cited as notable attributes.
A significant early concern was the model's original license, which contained clauses users described as restrictive; this was resolved when NVIDIA updated the license shortly after release. Discussions also noted the absence of vision capabilities as a trade-off compared to some contemporaries, and users debated use cases where long context and speed matter more than multimodal support.
Nvidia updated the Nemotron Super 3 122B A12B license to remove the rug-pull clauses
Qwen3.5 122b vs. Nemotron 3 Super 120b: Best-in-class vision Vs. crazy fast + 1M context (but no vision). Which one are you going to choose and why?
Nemotron 3 Super release soon?
Nemotron-3-Super-120B-A12B NVFP4 inference benchmark on one RTX Pro 6000 Blackwell
Nvidia Nemotron 3 Super is here — 120B total / 12B active, Hybrid SSM Latent MoE, designed for Blackwell
Documentation & links
Parameters & options
Explore similar models
Start building with Nemotron 3 Super 120B
No API keys required. Create AI-powered workflows with Nemotron 3 Super 120B in minutes — free.