speculative decoding Articles
Browse 7 articles about speculative decoding.

Thinking Cap vs Swift 1.5 vs Qwen Pi: Best Qwen3.8-27B Fine-Tune?
Three teams fine-tuned Qwen3.8-27B to think less without losing accuracy. Here's how Thinking Cap, Swift 1.5, and Qwen Pi actually compare.

How to Run Qwen 3.8 Locally With the Superlinked Inference Engine
A hands-on guide to installing the Superlinked Inference Engine and running Qwen 3.8 27B locally with GPU-tuned profiles and speculative decoding.

What Is DFlash 2? Speculative Decoding Explained for Qwen3.8-27B
DFlash 2 is a block-diffusion draft model that speeds up Qwen3.8-27B inference up to 3.4x. Here's how it works and what it beats.

DFlash 2: Run Qwen3.8-27B at 2x Speed with Speculative Decoding
DFlash 2 speeds up Qwen3.8-27B inference roughly 2x on a single A100 using speculative decoding in SGLang, with no output quality loss.

NVIDIA Nemotron 3.5 Lightning: A 30B MoE Built for Agent Grunt Work
NVIDIA's Nemotron 3.5 Lightning is a 30B-A3B open MoE model built for fast, cheap agent execution. Here's what its architecture and benchmarks mean.

Poolside's Laguna S 2.1: A 118B Open Model for Local Agentic Coding
Poolside's open-weight Laguna S 2.1 runs agentic coding locally at 80+ tokens/sec on DGX Spark, matching models ten times its size.

Local LLM Speed Test: GPT-OSS, Qwen3.6 and Hermes on 128GB Unified Memory
Real token-per-second benchmarks for GPT-OSS 120B, Qwen3.6 MoE, and Hermes agents running locally on 128GB unified memory hardware.