Skip to main content
MindStudio
Pricing
BlogAbout
My Workspace
Text Generation Model

Nemotron 3 Nano 30B

Nemotron 3 Nano 30B is a text generation model from Nvidia with a 1,000,000-token context window and mixture-of-experts architecture.

PublisherNvidia
TypeText
Context Window1,000,000 tokens
ReleasedDecember 2025
Input$0.05/MTok
Output$0.20/MTok
ProviderDeepInfra

Large sparse model with 1M token context

Nemotron 3 Nano 30B is a large language model developed by Nvidia, released in December 2025 under the full identifier nvidia/Nemotron-3-Nano-30B-A3B. The model uses a mixture-of-experts (MoE) architecture, activating approximately 3 billion parameters per forward pass despite having 30 billion total parameters, which reduces compute requirements relative to dense models of the same size. It supports a context window of up to 1,000,000 tokens, making it suited for tasks involving very long documents or extended conversations.

Nemotron 3 Nano 30B is a text-only model designed for chat and instruction-following use cases. It includes a configurable reasoning input, allowing users to toggle reasoning behavior at inference time. With a maximum response size of 16,384 tokens and its large context capacity, it is well suited for summarization of lengthy documents, multi-turn dialogue, and tasks that require retaining information across long inputs.

What Nemotron 3 Nano 30B supports

Long Context Window

Supports up to 1,000,000 tokens of context, enabling processing of very long documents or extended multi-turn conversations in a single pass.

Configurable Reasoning

Exposes a reasoning toggle as a select input at inference time, allowing users to enable or disable chain-of-thought reasoning behavior per request.

Mixture-of-Experts Architecture

Uses a sparse MoE design that activates roughly 3 billion of its 30 billion parameters per forward pass, reducing active compute per token.

Instruction Following

Trained for chat and instruction-following tasks, supporting structured dialogue and multi-step task completion.

Extended Response Output

Generates responses of up to 16,384 tokens per request, supporting detailed answers, long-form writing, and document-length outputs.

Ready to build with Nemotron 3 Nano 30B?

Get Started Free

Common questions about Nemotron 3 Nano 30B

What is the context window size for Nemotron 3 Nano 30B?

Nemotron 3 Nano 30B supports a context window of 1,000,000 tokens, allowing very long documents or conversations to be processed in a single request.

How many parameters does Nemotron 3 Nano 30B activate per request?

Despite having 30 billion total parameters, the model uses a mixture-of-experts architecture and activates approximately 3 billion parameters per forward pass, as indicated by the 'A3B' in its full model name nvidia/Nemotron-3-Nano-30B-A3B.

What is the maximum response length this model can generate?

The model supports a maximum response size of 16,384 tokens per request.

Does Nemotron 3 Nano 30B support image or video inputs?

No. Based on the available metadata, Nemotron 3 Nano 30B is a text-only model and does not support image or video inputs.

What is the reasoning input option and how does it work?

The model exposes a 'Reasoning' select input that can be configured at inference time. This allows users to toggle reasoning behavior, such as chain-of-thought processing, on or off depending on the task.

When was Nemotron 3 Nano 30B released?

Nemotron 3 Nano 30B was released in December 2025 by Nvidia.

What people think about Nemotron 3 Nano 30B

Community reception on r/LocalLLaMA was largely positive, with the release announcement gathering over 850 upvotes and 180 comments, reflecting strong interest in the model's hybrid MoE architecture and its 1M token context support. Users highlighted the efficiency of the 3.5B active parameter design and the strong math and coding benchmark scores as notable characteristics.

Some community members focused on practical local deployment, with a dedicated thread benchmarking the model using Vulkan and RPC backends to assess real-world performance on consumer hardware. Concerns and discussion points included inference speed on non-NVIDIA hardware and how the model performs outside of benchmark conditions.

View more discussions →

Parameters & options

Max Temperature1
Max Response Size16,384 tokens
ReasoningSelect

When enabled, the model generates a reasoning trace before providing a final answer. This improves performance on complex tasks like math, coding, and multi-step reasoning, but results in longer responses and higher token usage.

Default: false
DisabledEnabled

Start building with Nemotron 3 Nano 30B

No API keys required. Create AI-powered workflows with Nemotron 3 Nano 30B in minutes — free.