open weight LLM
open weight LLM Articles
Browse 5 articles about open weight LLM.

DeepSeek-V4-Pro-0813: What's New and How It Stacks Up
DeepSeek-V4-Pro-0813 adds DSpark speculative decoding and beats its preview on coding, agentic, and tool-use benchmarks versus GLM 5.2, Kimi K3, and Opus 4.8.
DeepSeek V4 ProDeepSeek-V4-Pro-0813DSpark speculative decoding

dots3-note Preview: Inside the 280B Multimodal MoE Model
dots3-note preview is a 280B-parameter, 16B-active multimodal MoE model with 512K context. Here's what it is and how it works.
dots3-notedots studio AI modelmultimodal MoE model

Qwen3.8-2.4T-A95B: Specs, Architecture, and Benchmarks Explained
Qwen3.8-2.4T-A95B specs: 2.4T total/95B active MoE parameters, 262K context, and benchmark scores versus Opus 4.8 and GPT 5.6.
Qwen3.8-2.4T-A95BQwen3.8-MaxQwen3.8 benchmarks

Qwen 3.8 Max Benchmarks: Where It Really Ranks vs Claude and GPT-5.6
Qwen 3.8 Max claims to trail only Gemini. Real DeepSWE and GPQA scores show a more mixed picture against GPT-5.6 and Opus.
Qwen 3.8 MaxQwen benchmarksopen weight LLM

ThinkingCap: The Qwen Fine-Tune That Cuts Reasoning Tokens 46%
ThinkingCap fine-tunes Qwen 3.6 27B to cut chain-of-thought tokens by 46% while holding benchmark accuracy, a big deal for local coding setups.
ThinkingCap modelQwen 3.6 27Blocal coding AI