MoE model
MoE model Articles
Browse 4 articles about MoE model.

MiMo-V2.6-Pro-RL: Xiaomi's 1T-Parameter Agentic Model, Explained
Xiaomi's MiMo-V2.6-Pro-RL is a 1T-parameter MoE model trained with large-scale RL. Here's how it compares to Claude Opus 5 and GPT-5.6.
MiMo-V2.6-Pro-RLXiaomi MiMo1T parameter model

DeepSeek-V4.1-Flash: How 890-Byte KV Cache Compression Works
DeepSeek-V4.1-Flash compresses KV cache to 890 bytes/token via a Causal Encoder-Decoder design. Full benchmarks vs GPT-5.6, Opus 5, K3, GLM-5.3.
DeepSeek-V4.1-FlashKV cache compressionDeepSeek benchmark

Tencent Hy4 Preview: Full Specs and How It Stacks Up to GLM 5.3, Kimi K3
Tencent's open-weight Hy4 preview MoE model edges out GLM 5.3 and Kimi K3 in blind engineering evals. Specs, architecture, and benchmarks explained.
Tencent Hy4 previewHy4 benchmarksGLM 5.3 vs Hy4

NVIDIA Nemotron 3.5 Lightning: A 30B MoE Built for Agent Grunt Work
NVIDIA's Nemotron 3.5 Lightning is a 30B-A3B open MoE model built for fast, cheap agent execution. Here's what its architecture and benchmarks mean.
Nemotron 3.5 LightningNVIDIA open modelMoE model