AI Model Reviews & Comparisons
Reviews, explainers, and head-to-head comparisons of released AI models. Includes 'What is [model]?' evergreen posts, single-model reviews, capability deep-dives, and side-by-side comparisons. Closed-source frontier models (GPT, Claude, Gemini) are the main beat; non-deployment content on open models lives here too. Deployment guides for open models stay in Local & Open-Weight Models.

Gemini 3.7 Flash Pricing: Where to Get It Free or Cheap Right Now
Gemini 3.7 Flash is free in Antigravity and AI Studio, with a limited-time discount on OpenRouter. Here's every access point and price.

Ornith-1.5-35B-A3B: A Self-Improving MoE Model for Coding Agents
Ornith-1.5-35B-A3B is a 35B mixture-of-experts model trained via self-improvement loops that beats larger models on coding benchmarks.

What Is GLM 5.3 and Zcode? ZAI's Coding Model Explained
GLM 5.3 is ZAI's post-trained coding model, and Zcode is its dedicated agentic dev app. Here's how they work and what sets them apart.

DeepSeek's Files API: Upload Images Once, Reuse Them by ID
DeepSeek's Files API lets you upload an image once and reference it by ID across requests, avoiding repeated uploads and wasted tokens.

DeepSeek V4-Flash Vision: What It Nails and Where It Fails
Hands-on tests of DeepSeek's first vision model show strong chart reading and OCR but real mistakes on handwriting and math symbols.

Dots.3 Note Preview: Xiaohongshu's New MoE Model, Tested
Dots.3 Note Preview is a 280B MoE multimodal model from Xiaohongshu's AI lab. Here's what its specs, benchmarks, and hands-on tests show.

GLM 5.3 Review: Coding, Security, and Creative Tests Put to the Test
A hands-on GLM 5.3 review covering live coding refactors, a defensive security audit, and creative HTML generation against a real Dockerized app.

Ornith 1.5 35B-A3B Benchmarks: How It Stacks Up Against Qwen3.6
Ornith 1.5 35B-A3B benchmark breakdown vs Qwen3.6-35B, Gemma 4-31B and Muse Glimmer-30B across SWE-bench, Terminal-Bench, HLE and agentic tests.

DeepSeek-V4-Pro-0813: What's New and How It Stacks Up
DeepSeek-V4-Pro-0813 adds DSpark speculative decoding and beats its preview on coding, agentic, and tool-use benchmarks versus GLM 5.2, Kimi K3, and Opus 4.8.

DeepSeek-V4-Pro-0813 Benchmarks: How It Stacks Up Against Opus and Kimi K3
DeepSeek-V4-Pro-0813 benchmark scores across Terminal Bench, HLE, and Cybergym, compared against GLM-5.2, Kimi K3, Opus-4.8, and Fable-5.

GPT-5.6 Soul Ultrafast: 14x Speed via Cerebras Explained
OpenAI's Ultrafast mode runs GPT-5.6 Soul on Cerebras chips at up to 14-15x normal speed, turning long agent tasks into short ones.

GPT-5.6 Soul vs Claude Opus 5: Who Survives a 7-Day Crisis?
A settlement survival test pits GPT-5.6 Soul against Claude Opus 5 in the same crisis scenario, with sharply different survival outcomes.

What Is Grokbot? xAI's New Agentic Assistant Explained
Grokbot from xAI turns Grok 4.6 into named AI agents that plug into Gmail and Slack, learn tasks by watching you, and run on schedules.

Qwen3.8-27B Explained: Hybrid Attention, 262K Context, New Benchmarks
Qwen3.8-27B pairs gated Delta Net linear attention with full attention, scaling to 262K context. Here's what changed and how it benchmarks.

Qwen3.8 27B Vision and Multilingual Test: How Good Is It Really?
Hands-on testing of Qwen3.8 27B's vision accuracy on real photos and artwork, plus multilingual translation across 80 languages, warts included.

What Is Tencent's WorldClaw? AI-Generated 3D Worlds Explained
Tencent's WorldClaw turns text prompts into editable 3D worlds with separate assets, built on GPT Image 2, Meta's SAM 3, and Hunyuan 3D.

What Is Grok Bot? xAI's Install-and-Go AI Agent Explained
Grok Bot is xAI's no-code AI agent platform where named bots share one cloud computer. Here's how it works and who it's for.

GLM-5.3 Benchmark Results: How It Stacks Up Against Opus 5, Fable 5
GLM-5.3 scores 91% on an independent coding benchmark, topping Opus 5 and Kimi K3 while matching Fable 5 on tough 3D and UI tasks.

GLM-5.3: ZAI's New Model Shows Unexpected Cybersecurity Skills
ZAI's GLM-5.3 pairs frontier coding benchmarks with surprising cybersecurity gains, arriving first through a coding plan ahead of open weights.

Grok 4.6 Explained: xAI's Flagship Nears GPT-5.6 and Claude Levels
Grok 4.6 lands near GPT-5.6 and Claude Opus 5 on major benchmarks at roughly half the cost. Here's what changed and why it matters.

North MicroVision OCR Accuracy: What Real Tests Show
Hands-on testing of Cohere's North MicroVision shows strong invoice and French OCR, shaky handwriting results, and weak Urdu and Indonesian accuracy.

Did Claude Make Progress on the Riemann Hypothesis? Here's What Happened
Anthropic's unreleased Claude research model pushed a key bound on the Riemann Hypothesis past prior human results, guided by encouragement alone.

DeepSeek V4 Pro 0813: Benchmark Results and Hands-On Test
DeepSeek V4 Pro 0813 benchmarked against GPT, Claude, and Gemini rivals, with official scores, independent coding tests, and pricing breakdown.

Qwen3.8-2.4T-A95B Benchmarks vs Opus 4.8 and GPT-5.6 Sol
Qwen3.8-2.4T-A95B benchmark results compared against Opus 4.8, Fable 5, and GPT-5.6 Sol across coding-agent and general-agent tests.