vLLM speculative decoding
vLLM speculative decoding Articles
Browse 4 articles about vLLM speculative decoding.

How to Deploy IQuest-Q1 with SGLang or vLLM
A practical guide to serving IQuest-Q1 locally with SGLang or vLLM, covering MTP speculative decoding, Docker setup, and tool-calling.
deploy IQuest-Q1SGLang setupvLLM speculative decoding

How to Run OrcaSAQ2 27B on a 16GB GPU
OrcaSAQ2 27B compresses a 54GB model to 12.3GB with near-BF16 fidelity, letting a single 16GB GPU run vLLM agents at up to 90 tok/s.
run OrcaSAQ2 locally27B model 16GB GPUOrcaSAQ2 vLLM

Run DFlash 2 Speculative Decoding with vLLM and SGLang
How to set up DFlash 2, a lossless draft model for Qwen3.8-27B, using vLLM or SGLang for up to 3.4x faster inference speeds.
DFlash 2 setupvLLM speculative decodingSGLang draft model

How to Run Qwen3.8-27B with DFlash2 on vLLM or SGLang
Set up speculative decoding for Qwen3.8-27B with the DFlash2 draft model using vLLM or SGLang, with H200 benchmark numbers included.
run DFlash2 locallyvLLM speculative decodingSGLang DFlash