mixture of experts model Articles
Browse 6 articles about mixture of experts model.

Xing4.0-29B-A4B: China Telecom's Agent-Focused MoE Model, Explained
Xing4.0-29B-A4B is China Telecom's new 29B MoE model trained on Ascend NPUs, benchmarked against Gemma4-26B-A4B and Qwen3.6-35B-A3B.

DeepSeek V4.1 Flash Benchmarks vs Opus 5 and GPT-5.6: What's Real?
DeepSeek V4.1 Flash matches Opus 5 and GPT-5.6 on paper, but hands-on coding tests expose a gap between benchmark scores and real output.

Tencent Hy4 Preview: A 770B MoE Model That Edges Out GLM-5.3
Tencent's Hy4 preview is a 770B-parameter, 49B-active MoE model with 1M context that beat GLM-5.3 and Kimi K3 in blind evals.

Qwen3.8-Flash-Next: Inside the Qwen 4 Architecture Preview
Qwen3.8-Flash-Next previews Qwen 4's architecture: hybrid gated-delta and sparse attention, engram embeddings, 125B params, 6B active.

Ornith 1.5 35B-A3B: Local Deployment, VRAM, and Real-World Tests
Ornith 1.5 35B-A3B is a mixture-of-experts model with 3B active params. Here's how it runs locally on an A100, and how it handles agentic and reasoning tests.

Poolside's Laguna S 2.1: A 118B Open Model for Local Agentic Coding
Poolside's open-weight Laguna S 2.1 runs agentic coding locally at 80+ tokens/sec on DGX Spark, matching models ten times its size.