speculative decoding
speculative decoding Articles
Browse 4 articles about speculative decoding.

DFlash 2: Run Qwen3.8-27B at 2x Speed with Speculative Decoding
DFlash 2 speeds up Qwen3.8-27B inference roughly 2x on a single A100 using speculative decoding in SGLang, with no output quality loss.
DFlash 2speculative decodingQwen3.8-27B

NVIDIA Nemotron 3.5 Lightning: A 30B MoE Built for Agent Grunt Work
NVIDIA's Nemotron 3.5 Lightning is a 30B-A3B open MoE model built for fast, cheap agent execution. Here's what its architecture and benchmarks mean.
Nemotron 3.5 LightningNVIDIA open modelMoE model

Poolside's Laguna S 2.1: A 118B Open Model for Local Agentic Coding
Poolside's open-weight Laguna S 2.1 runs agentic coding locally at 80+ tokens/sec on DGX Spark, matching models ten times its size.
Laguna S 2.1Poolside AI modelDGX Spark local LLM

Local LLM Speed Test: GPT-OSS, Qwen3.6 and Hermes on 128GB Unified Memory
Real token-per-second benchmarks for GPT-OSS 120B, Qwen3.6 MoE, and Hermes agents running locally on 128GB unified memory hardware.
local LLM benchmarkGPT-OSS tokens per secondQwen3.6 MoE