AI Model Reviews & Comparisons
Reviews, explainers, and head-to-head comparisons of released AI models. Includes 'What is [model]?' evergreen posts, single-model reviews, capability deep-dives, and side-by-side comparisons. Closed-source frontier models (GPT, Claude, Gemini) are the main beat; non-deployment content on open models lives here too. Deployment guides for open models stay in Local & Open-Weight Models.

Qwen3.8-2.4T-A95B: Specs, Architecture, and Benchmarks Explained
Qwen3.8-2.4T-A95B specs: 2.4T total/95B active MoE parameters, 262K context, and benchmark scores versus Opus 4.8 and GPT 5.6.

What Is Model Distillation in AI? Teacher-Student Training Explained
A clear breakdown of AI model distillation: soft labels, Hinton's original method, on-policy distillation, and why the term gets misused in AI news.

What Is fuse-1 Lite? Inside the Model Built by Transplanting Coding Experts
fuse-1 Lite fuses LiquidAI's LFM2.5 with coding experts pulled from Qwen3.6-35B-A3B. Here's how this expert-transplant model actually works.

Meta Muse Glimmer 30B: How to Run It Locally and Is It Worth It?
Meta's open-weight Muse Glimmer 30B rivals Qwen 3.6 27B on agent benchmarks. Here's the hardware, quantization, and setup to run it yourself.

NVIDIA Nemotron 3.5 Lightning: A 30B MoE Built for Agent Grunt Work
NVIDIA's Nemotron 3.5 Lightning is a 30B-A3B open MoE model built for fast, cheap agent execution. Here's what its architecture and benchmarks mean.

Qwen 3.8 Max Benchmarks: Where It Really Ranks vs Claude and GPT-5.6
Qwen 3.8 Max claims to trail only Gemini. Real DeepSWE and GPQA scores show a more mixed picture against GPT-5.6 and Opus.

Qwen 3.8 Max Explained: Alibaba's 2.4 Trillion Parameter Model
Qwen 3.8 Max is Alibaba's open-weight 2.4 trillion parameter model with frontier coding and agentic benchmarks. Here's what it can actually do.

Qwen 3.8 Max Tested: Coding, Front-End Design, and a Cheating Incident
Hands-on tests of Qwen 3.8 Max on coding, front-end design, and agentic tasks, including a caught cheating incident and pricing comparison.

Google Shipped Three Gemini Models at Once. Here's What Actually Changed
Google quietly released Gemini 3.6 Flash, 3.5 Flash Light, and 3.5 Cyber, prioritizing token efficiency over raw benchmark chasing. Here's the breakdown.

ThinkingCap: The Qwen Fine-Tune That Cuts Reasoning Tokens 46%
ThinkingCap fine-tunes Qwen 3.6 27B to cut chain-of-thought tokens by 46% while holding benchmark accuracy, a big deal for local coding setups.

Kimi K3 vs Claude Opus 5: An Open Model Takes On 3D World Generation
Kimi K3 vs Claude Opus 5 in a long-horizon 3D world generation test using the Klein agent harness. Here's how the open model held up.

Are Chinese AI Models Really Catching the US Frontier?
DeepSeek, Kimi K3, Qwen, and GLM compared on price, licensing, and capability against US frontier models, with a practical framework for testing them.

Claude Opus 5 Is One-Shotting Playable 3D Games From Scratch
Claude Opus 5 is generating playable FPS, zombies, and Minecraft-style games in a single prompt. Here's what that leap actually means.

Claude Opus 5 Benchmarks: The Numbers Anthropic Didn't Headline
Claude Opus 5's independent benchmark results, from Arc AGI 3 to IMO 2026, and how it actually stacks up against Claude Fable 5 in practice.

ChatGPT Triples Custom Instructions Limit: What It Means for You
OpenAI expanded ChatGPT's custom instructions field from 1,500 to over 5,000 characters, allowing much richer persistent context in every chat.

Claude Opus 5: Anthropic's Cheaper Model That Rivals Fable 5
Claude Opus 5 matches or beats Fable 5 on most benchmarks at half the price, with a record jump on ARC-AGI-3. Here's what changed.

Kimi K3 Benchmarked: Is It Really as Good as the Hype?
Hands-on testing shows Kimi K3 nearly matches Opus 4.8 on simple coding tasks but fails far more often on complex, trap-designed work.

How Claude Opus 5 Cracked ARC-AGI-3 With Algebraic Reasoning
Opus 5 jumped from single digits to 30% on ARC-AGI-3 by converting visual puzzles into algebra, a new reasoning behavior researchers hadn't seen before.

Claude Opus 5 Builds Full 3D Worlds and Sims From One Prompt
Claude Opus 5 turns single prompts into 3D worlds, physics sims, and games. Here's what the model actually produced in hands-on testing.

Opus 5 vs Fable 5: Hands-On Testing Reveals the Real Winner
Real workflow tests across coding, video, landing pages, and LinkedIn content show where Opus 5 beats Fable 5 on cost, speed, and output quality.

Build Cinematic Scroll-Driven Websites with Kimi K3 for $1
A practical tutorial on building scroll-animated cinematic websites using Kimi K3, Higsfield's Cinematic Studio, and frame interpolation.

Mixture of Experts Architecture Explained: How GLM 5.2 Runs 40B Active Parameters
GLM 5.2 has 744B total parameters but only 40B active per token thanks to MoE routing. Learn how this architecture enables local inference on consumer hardware.

What Is GLM 5.2? The Open-Weight Model Beating Frontier AI on Design
GLM 5.2 is a 744B parameter open-weight model with 256 experts per layer. Learn what makes it exceptional for frontend design and agentic loops.

What Is ChatGPT Work Mode? OpenAI's Agentic Super App Explained
ChatGPT Work is a new agentic mode that does work for you instead of with you. Learn how it builds websites, runs tasks, and what makes it different.