OrcaSAQ2 27B: Run Qwen3.8-27B in 12GB With 3-Bit Quantization
OrcaSAQ2 27B shrinks Qwen3.8-27B from 54GB to 12.3GB using 3-bit mixed-precision quantization, with setup steps for vLLM and MTP decoding.

What is OrcaSAQ2 27B?
OrcaSAQ2 27B is a quantized version of Qwen3.8-27B that shrinks the checkpoint from 54GB (BF16) down to 12.3GB using a 3-bit mixed-precision scheme, while claiming only a 0.02% increase in perplexity and 93.2% token-level agreement with the original model. It’s released by OrcaRouter under Apache-2.0 and is built to run through vLLM, making a 27-billion-parameter reasoning model deployable on a single 16GB GPU instead of a multi-GPU server.
TL;DR
- OrcaSAQ2 27B compresses Qwen3.8-27B from a 54GB BF16 checkpoint to 12.3GB, a 77.2% reduction in storage footprint.
- The quantization method uses a sensitivity-aware mixed-precision scheme averaging 3.21 bits per weight, rather than a flat 3-bit cast across every layer.
- On WikiText-2, OrcaSAQ2 reports a perplexity of 5.6482 versus 5.6468 for BF16, a difference of just 0.02%, alongside 93.2% top-1 token agreement and a mean KL divergence of 0.031.
- The model reports 70.0% on SWE-bench Verified and 58.4% on Terminal-Bench 2.1, scores that sit near much larger frontier models despite the small checkpoint, though these are cited as reference points rather than controlled apples-to-apples comparisons.
- It supports 262K context, tool calling, a “thinking” reasoning mode, and MTP (multi-token prediction) speculative decoding for faster single-stream generation.
- Serving is done through vLLM with a dedicated kernel package, and the model exposes a standard OpenAI-compatible API once launched.
- The release is text-only; the vision tower from the base model is not included in this checkpoint.
How does OrcaSAQ2 compress a 27B model to 12.3GB?
OrcaSAQ2 uses what OrcaRouter calls a sensitivity-aware mixed-precision quantization system. Instead of quantizing every weight to a uniform 3-bit format, the method allocates precision unevenly across the network, giving more bits to components that are more sensitive to error and fewer bits to components that tolerate compression well. The result is an average of 3.21 bits per weight across the decoder, landing the total checkpoint at 12.3GB, down from the 54GB BF16 original, a 77.2% reduction.
The company has not disclosed the calibration strategy, precision allocation rules, or packing techniques behind the method. That’s a real limitation for anyone trying to reproduce or audit the approach: the numbers on the model card are self-reported by OrcaRouter, measured against their own evaluation path, and there’s no independent third-party benchmark included here. Treat the fidelity figures as vendor claims worth testing against your own workload rather than settled fact.
Does 3-bit quantization actually hold up, or is this just a smaller, dumber model?
The headline argument from OrcaRouter is that perplexity alone doesn’t tell you whether a compressed model still behaves like the original when it matters. On the standard measure, WikiText-2 perplexity moves from 5.6468 (BF16) to 5.6482 (OrcaSAQ2), a 0.02% delta over 16,376 predicted tokens. That’s a genuinely small gap on a token-prediction benchmark.
But top-1 agreement, meaning how often the quantized model picks the exact same next token as the BF16 original, sits at 93.2%. That means roughly 1 in 14 token decisions differs from what the full-precision model would have produced. The mean KL divergence between the two output distributions is 0.031. None of that is catastrophic, but it’s also not “identical.” OrcaRouter is explicit about this: the model card states that quantization is not mathematically lossless and that 93.2% agreement means some token decisions differ from BF16. It also cautions that a low PPL delta does not guarantee identical downstream performance.
The stronger test, according to the model card’s own framing, is long-horizon agent behavior. In a multi-step agent loop (plan, act, observe, decide, recover, repeat), a single divergent token can change a tool call, which changes environment state, which changes every subsequent decision. Short single-turn benchmarks can mask this kind of compounding error. That’s why the release leans on agent benchmarks rather than perplexity alone to make its case.
How does it perform on coding and agent benchmarks?
OrcaRouter reports 70.0% on SWE-bench Verified and 58.4% on Terminal-Bench 2.1 for OrcaSAQ2 27B. For context, the same tables list Claude Sonnet 4.6 at 79.6% and Gemini 3 at 76.2% on SWE-bench Verified, and Gemini 3.1 Pro paired with Terminus 2 at 70.7% on Terminal-Bench 2.1. OrcaSAQ2’s scores land close to or above some larger, better-resourced systems like Qwen3-Coder-480B-A35B (69.6% on SWE-bench Verified) and GPT-4.1 (54.6%).
The model card itself flags the obvious caveat: agent benchmark scores depend heavily on the surrounding scaffold, the specific tools available, the reasoning token budget, timeouts, and the execution environment. Comparing a 27B model in a 12GB checkpoint against Claude Opus running inside Claude Code is not a controlled, model-only comparison. These figures are best read as “this compressed 27B model is competitive with systems that require far more memory and compute,” not as a definitive leaderboard ranking.
What hardware do you need to run OrcaSAQ2 27B?
Built like a system. Not vibe-coded.
Remy manages the project — every layer architected, not stitched together at the last second.
The core pitch is that a 12.3GB checkpoint fits comfortably on a 16GB GPU, which the original 54GB BF16 weights simply cannot do without multi-GPU sharding. OrcaRouter’s own benchmarks were measured under a 15.7GiB GPU memory cap, leaving a few gigabytes for KV cache and runtime overhead after the weights are loaded.
Actual usable context depends on several variables: vLLM overhead, KV-cache configuration, whether MTP is enabled, batch size, and CUDA graph settings. The model card suggests around 32K tokens of interactive context as a practical starting point on a 16GB card, even though the architecture technically supports up to 262K tokens. Fitting the full context window would require considerably more VRAM for the KV cache alone.
What is MTP speculative decoding and is it worth turning on?
MTP (multi-token prediction) is a speculative decoding technique built into the model that predicts multiple tokens ahead per forward pass, then verifies them, which can speed up generation without changing output quality. OrcaRouter’s benchmarks show single-stream decode throughput jumping from 65.3 tok/s with MTP off to 90.1 tok/s with MTP on, a 38% improvement.
The tradeoff shows up under concurrency. With multiple simultaneous streams, MTP off actually wins: 332-333 tok/s aggregate across 8-16 streams versus 219-220 tok/s with MTP on, because MTP consumes extra compute and KV-cache capacity (14,563 tokens of KV pool versus 29,354 tokens with it off). In practice, this means MTP is a good default for single-user, latency-sensitive use like a coding assistant or interactive terminal agent, but should probably be disabled for high-throughput batch serving.
How do you install and serve it?
The setup uses standard vLLM tooling plus a dedicated kernel package for the quantization format:
pip install -U vllm huggingface_hub
pip install git+https://github.com/Continuum-AI-Corp/OrcaSAQ2-kernel
hf download orcarouter/OrcaSAQ2-27B --local-dir ./OrcaSAQ2-27B
vllm serve ./OrcaSAQ2-27B \
--served-model-name OrcaSAQ2-27B \
--max-model-len 262144 \
--reasoning-parser qwen3 \
--enable-auto-tool-choice \
--tool-call-parser qwen3_xml \
--speculative-config '{"method":"qwen3_next_mtp","num_speculative_tokens":2}'
Once running, it exposes an OpenAI-compatible chat completions endpoint, so existing code using the openai Python client can point base_url at the local server and swap in OrcaSAQ2-27B as the model name. Recommended sampling settings are temperature 1.0, top_p 0.95, top_k 20, with thinking mode enabled by default.
Frequently Asked Questions
What base model does OrcaSAQ2 27B use?
It’s a quantized version of Qwen/Qwen3.8-27B, a 64-layer hybrid-attention model (48 Gated DeltaNet layers plus 16 full-attention layers) with a 5120 hidden size and a 248,320-token vocabulary. The architecture is labeled Qwen3_5ForCausalLM.
Does OrcaSAQ2 support vision or multimodal input?
No. This checkpoint is text-only. The vision tower from the base model was not included in the release.
Can I run OrcaSAQ2 27B without vLLM?
Not according to the model card. It explicitly requires the OrcaSAQ2 vLLM integration and a dedicated kernel package installed via pip, so standard transformers-based inference paths aren’t the intended route.
How does OrcaSAQ2 compare to just running a smaller model at full precision?
The model card doesn’t include that comparison directly, but the argument implicit in the release is that a compressed 27B model retains more reasoning and agentic capability than a smaller model trained or distilled to fit the same memory budget. Whether that holds for a specific use case is worth testing directly, since the published benchmarks come from OrcaRouter’s own evaluation pipeline.
Is the quantization method open source?
Plans first. Then code.
Remy writes the spec, manages the build, and ships the app.
No. OrcaRouter describes the sensitivity-aware mixed-precision method as proprietary and has not disclosed the calibration strategy, precision allocation approach, or packing techniques. The resulting weights and license (Apache-2.0, inherited from Qwen3.8-27B) are open, but the method itself is not.



