Naive-N0.5-Flash: What's Inside This 309B MoE Coding Model?
Naive-N0.5-Flash packs 309B params, 15.5B active, native 1M context via sparse attention, and MIT-licensed weights aimed at coding and AI R&D.

What is Naive-N0.5-Flash?
Naive-N0.5-Flash is an open-weight mixture-of-experts (MoE) language model built by NaiveAI for coding and AI research and development tasks. It has 309 billion total parameters but only activates 15.5 billion per token, and it natively handles context windows up to 1 million tokens without relying on traditional full attention. The weights and inference code are released under the MIT license, with an API also planned.
TL;DR
- Naive-N0.5-Flash is a 309B-parameter MoE model with only 15.5B active parameters per forward pass, keeping inference costs closer to a much smaller dense model.
- It achieves a native 1M-token context window using a hybrid of Sliding-Window Attention and DeepSeek Sparse Attention (DSA), with zero full-attention layers anywhere in the network.
- The model builds on Xiaomi’s MiMo-V2.5 base model and went through 3.25 trillion tokens of continued training to adapt to its new sparse attention design.
- NaiveAI’s own inference stack, called NaiveRT, claims up to 2,000 tokens per second in “Ultrafast” mode versus 50 tokens per second in standard mode.
- The model card reports benchmark comparisons against models like GLM-5.3, Kimi-K3, Qwen-3.8-Max, and Claude Opus 5.5 across coding and AI R&D tasks, though these are self-reported.
- Weights are released under the MIT license, and planned API pricing is listed at $0.10 per million input tokens, $0.40 per million output tokens, and $0.01 per million cached tokens.
- Running it locally requires FP8-capable NVIDIA GPUs and roughly 315 GB of memory just to hold the weights.
Other agents start typing. Remy starts asking.
Scoping, trade-offs, edge cases — the real work. Before a line of code.
How does the hybrid attention architecture work?
The headline feature of Naive-N0.5-Flash is that it reaches a 1-million-token context window without a single full-attention layer. Most long-context transformers rely on at least some global attention layers to preserve information across distant tokens, but those layers get expensive fast as context grows, since their decoding cost scales with sequence length.
Naive-N0.5-Flash avoids that entirely. The 48-layer network is organized into eight six-layer modules. A typical module runs five Sliding-Window Attention (SWA) layers followed by one DeepSeek Sparse Attention (DSA) layer, and the very first layer of the first module is also DSA instead of SWA. That gives the model 39 SWA layers and 9 DSA layers overall, a roughly 5-to-1 ratio.
SWA layers use a small 128-token window, so their per-token cost stays flat no matter how long the input gets. DSA layers handle long-range dependencies differently: a lightweight 16-head indexer scores the entire history, and then the backbone only computes full attention over the top 2,048 tokens that the indexer flags as most relevant. The full key-value cache is still retained and the indexer still scans everything, but the expensive attention computation itself only touches a small, selected subset of tokens. Both attention types also use sink bias, a technique that helps stabilize attention to the earliest tokens in a sequence.
One architectural deviation from the original DeepSeek Sparse Attention design is worth noting: Naive-N0.5-Flash swaps out DeepSeek’s multi-head latent attention (MLA) in favor of grouped-query attention (GQA) with four KV groups. NaiveAI says this choice, along with the rest of the architecture, was shaped through its own AI-driven R&D process rather than purely manual design.
What does “built on MiMo-V2.5” actually mean?
Naive-N0.5-Flash isn’t trained from scratch. It starts from Xiaomi’s open-weight MiMo-V2.5 base model, which NaiveAI describes as architecturally simple but strong in world knowledge and research-oriented reasoning. NaiveAI then performed continued pretraining to graft the hybrid SWA-DSA attention structure onto that base and push its coding and AI R&D capabilities further.
That continued training ran across 3.25 trillion tokens in three stages:
- Indexer Warmup (50 billion tokens) to get the new DSA indexer calibrated before heavy training begins.
- Sparse Attention Training (3 trillion tokens), the bulk of the process, which adapts the model’s weights to operate under the new sparse attention regime.
- Learning Rate Decay (200 billion tokens), a final annealing phase typical of large-scale pretraining runs.
This staged approach matters because swapping a model’s attention mechanism after pretraining is non-trivial. A model trained under dense or different sparse-attention assumptions doesn’t automatically know how to use a brand-new indexer and sparse selection pattern well. The multi-stage process is essentially retraining the model’s internal habits around a fundamentally different way of reading context.
How fast is it, and what does “NaiveRT” change?
Remy doesn't build the plumbing. It inherits it.
Other agents wire up auth, databases, models, and integrations from scratch every time you ask them to build something.
Remy ships with all of it from MindStudio — so every cycle goes into the app you actually want.
NaiveAI pairs the model with its own inference engine, NaiveRT, which it says was developed through AI-centered research and engineering rather than conventional manual optimization. NaiveRT combines three techniques: mega-kernel fusion (merging multiple GPU operations into single kernels to cut overhead), Programmatic Dependent Launch (a scheduling technique for overlapping GPU work), and speculative decoding (predicting multiple future tokens and verifying them in parallel to speed up generation).
The reported throughput numbers are 50 tokens per second per user in “Standard” mode and up to 2,000 tokens per second in “Ultrafast” mode. That’s a substantial spread, and it’s worth remembering these are vendor-reported figures from the model card rather than independently verified numbers. The exact tradeoffs between the two modes (likely batching, latency, or cost differences) aren’t detailed in the public documentation.
What do the benchmark claims show?
The model card presents two sets of benchmark comparisons: one on coding and agentic software engineering tasks, and one on AI research and systems optimization tasks.
On the coding side, Naive-N0.5-Flash is compared against a wide field of current models including GLM-5.3 and GLM-5.3-Flash, Kimi-K3, Qwen-3.8-Max, Hy4-preview, DeepSeek-V4.1-Flash, Step-5-preview, GPT-5.6-Sol, Claude Opus 5 and Opus 5.5, and others, across benchmarks like SWE-Bench Pro, Terminal-Bench 2.1, ALE-CLI, FrontierSWE v1, and ProgramBench. On the AI R&D side, it’s evaluated on MLE-bench-30, PaperBench, SOL-ExecBench, NanoChat AutoResearch, and NanoGPT SpeedRun, again against a mix of frontier models.
A few caveats are worth flagging for anyone reading these numbers closely. First, the comparison scores are pulled from each competing lab’s own blog posts, model cards, or public leaderboards rather than being run by a neutral third party under identical conditions. Second, Naive-N0.5-Flash’s own evaluations were run using Claude Code 2.1.207 as the harness, with a 1M-token context window, temperature 1.0, and top-p 0.95, exposing only basic file I/O and Bash tools. Third, some benchmarks (SOL-ExecBench, NanoChat AutoResearch, NanoGPT SpeedRun) were scored using NaiveAI’s own in-house “AutoResearch” harness, which isn’t an externally standardized tool. None of this means the results are wrong, but it does mean they should be read as a vendor’s self-reported comparison rather than an independent audit.
Is Naive-N0.5-Flash practical to run yourself?
For most individuals and small teams, running Naive-N0.5-Flash locally is not realistic. The model card specifies that the weights alone occupy approximately 315 GB, and that’s before accounting for additional memory needed during inference (KV cache, activations, and so on). It also requires FP8-capable NVIDIA GPUs, which narrows the hardware pool to recent, high-end accelerators typically found in data centers rather than consumer workstations.
The practical path for most developers will be the planned API, where NaiveAI has listed pricing of $0.10 per million input tokens, $0.40 per million output tokens, and $0.01 per million cached-token reads. That pricing structure, with cheap cache reads, suggests the service is tuned for workflows that repeatedly reuse large context windows, like iterative coding sessions or long research threads, where caching previously processed context saves money.
For teams that do have the GPU budget, the model is distributed through Hugging Face Transformers (version 5.17.0 or later) with standard AutoModelForCausalLM loading, and supports FP8 mixed-precision inference with the recommended sampling settings of temperature 1.0 and top-p 0.95.
Frequently Asked Questions
How many parameters does Naive-N0.5-Flash actually use per token?
Seven tools to build an app. Or just Remy.
Editor, preview, AI agents, deploy — all in one tab. Nothing to install.
It has 309 billion total parameters across its mixture-of-experts layers, but only 15.5 billion are active for any given token, which is what keeps inference cost and latency closer to a mid-sized dense model rather than a 309B dense one.
Does Naive-N0.5-Flash use full attention anywhere?
No. Every one of its 48 transformer layers uses either Sliding-Window Attention (local, fixed-cost) or DeepSeek Sparse Attention (indexed, top-k selection), with no global full-attention layers in the network.
What base model is Naive-N0.5-Flash built on?
It continues pretraining from Xiaomi’s open-weight MiMo-V2.5 base model, adapting it to the new hybrid sparse attention architecture over 3.25 trillion additional training tokens.
Can I run Naive-N0.5-Flash on a single consumer GPU?
No. The weights alone take up roughly 315 GB, and the model requires FP8-capable NVIDIA GPUs, which effectively limits practical deployment to multi-GPU server setups.
Is Naive-N0.5-Flash free to use?
The model weights and inference code are released under the MIT license, so they’re free to download and modify. A paid API is also planned, with per-token pricing listed on the model card for input, output, and cached tokens.



