DeepSeek Sparse Attention
DeepSeek Sparse Attention Articles
Browse 4 articles about DeepSeek Sparse Attention.

Naive-N0.5-Flash: Specs and Setup for the 309B MoE Coding Model
How to run Naive-N0.5-Flash locally: hardware needs, FP8 quantization, context window, and setup steps for this 309B MoE coding model.
Naive-N0.5-Flashrun Naive-N0.5-Flash locallyopen weight coding model

Naive-N0.5-Flash: What's Inside This 309B MoE Coding Model?
Naive-N0.5-Flash packs 309B params, 15.5B active, native 1M context via sparse attention, and MIT-licensed weights aimed at coding and AI R&D.
Naive-N0.5-FlashNaiveAI modelDeepSeek Sparse Attention

Naive-N0.5-Flash Benchmarks: Coding and AI R&D Scores Explained
Naive-N0.5-Flash benchmark results across coding, agentic, and AI R&D evals, with context on how it stacks up against rival open-weight models.
Naive-N0.5-Flash benchmarksNaive-N0.5-Flash vs GLM-5.3AI R&D benchmark

Naive-N0.5-Flash: Inside the 309B MoE Model With 1M Native Context
Naive-N0.5-Flash pairs a 309B MoE architecture with hybrid SWA-DSA attention for native 1M-token context, built on MiMo-V2.5.
Naive-N0.5-FlashMoE model1M context window