OrcaSAQ2 27B Benchmarks: How a 3-Bit Model Rivals Claude on SWE-bench
OrcaSAQ2 27B hits 70% on SWE-bench Verified and 58.4% on Terminal-Bench 2.1 from a 12.3 GB checkpoint. Here's what the numbers actually mean.

What is OrcaSAQ2 27B and why does its benchmark data matter?
OrcaSAQ2 27B is a quantized version of Qwen3.8-27B, compressed from a 54 GB BF16 checkpoint down to 12.3 GB using a proprietary “sensitivity-aware mixed-precision” method from OrcaRouter. What makes it notable isn’t the compression ratio alone, it’s that the model reportedly scores 70.0% on SWE-bench Verified and 58.4% on Terminal-Bench 2.1, landing in the same range as much larger, closed frontier models like Claude Sonnet 4.5 and Gemini 2.5 Pro, while running on roughly a quarter of the memory.
TL;DR
- OrcaSAQ2 27B shrinks Qwen3.8-27B from 54 GB to 12.3 GB (a 77.2% reduction) while reporting only a +0.02% perplexity increase versus the original BF16 weights.
- On SWE-bench Verified, the model scores 70.0%, ahead of Qwen3-Coder-480B-A35B (69.6%) and Gemini 2.5 Pro (63.8%), and behind Claude Sonnet 4.5 (77.2%) and Claude Sonnet 4.6 (79.6%).
- On Terminal-Bench 2.1, it scores 58.4%, close to Claude Sonnet 4.6 running Claude Code (58.5%) and ahead of Gemini 3 Flash with Gemini CLI (56.9%).
- Fidelity to the original model is measured with 93.2% token-level Top-1 agreement and a mean KLD of 0.031, both computed on WikiText-2 against the BF16 reference.
- The checkpoint supports 262K context, tool calling, thinking mode, and MTP speculative decoding, and is served through vLLM.
- With MTP speculative decoding enabled, single-stream decode throughput reportedly reaches 90.1 tokens/second on a 16 GB GPU, a 38% gain over MTP disabled.
- OrcaRouter frames the benchmarks as public reference points, not strict apples-to-apples comparisons, since agent scaffolds, tools, and reasoning budgets differ across models.
Everyone else built a construction worker.
We built the contractor.
One file at a time.
UI, API, database, deploy.
How does OrcaSAQ2 achieve 3-bit quantization without wrecking the model?
The technique is called “sensitivity-aware mixed-precision quantization.” Instead of applying a uniform bit-width across every weight, the method allocates precision unevenly, giving more bits to parameters that are more sensitive to error and fewer bits to those that tolerate compression well. The result is a decoder that averages 3.21 bits per weight (bpw), down from the original 16-bit BF16 format.
OrcaRouter has not disclosed the calibration strategy, precision allocation rules, or packing techniques behind the method. That’s a real gap for anyone trying to reproduce the approach independently. What’s published instead is the output: fidelity metrics measured against the BF16 reference on WikiText-2, using 16,376 predicted tokens as the evaluation set.
The headline fidelity numbers are:
- Perplexity: 5.6468 (BF16) versus 5.6482 (OrcaSAQ2), a +0.02% delta.
- Top-1 token agreement: 93.2%, meaning the quantized model picks the identical next token as BF16 about 93 times out of 100.
- Mean KLD (Kullback-Leibler divergence) of 0.031, a measure of how far the quantized model’s output probability distribution drifts from the original.
These are standard proxies for “did quantization break the model,” and they look strong. But they’re static, single-token measurements. That’s precisely why the more interesting question is what happens over a long agent trajectory.
Why do long-horizon agent benchmarks matter more than perplexity?
Perplexity measures how well a model predicts the next token in a fixed piece of text. It’s cheap to compute and easy to reproduce, but it says nothing about what happens when a model has to act, observe the result, and decide what to do next, repeatedly, over dozens or hundreds of steps.
That’s the core argument OrcaRouter makes for testing on long-horizon agent benchmarks: a small perplexity gap can hide compounding errors. If a quantized model picks the wrong token 6.8% of the time (implied by the 93.2% Top-1 agreement figure), and one of those wrong tokens happens to break a tool call or misread a file path, the entire trajectory can go sideways. The error doesn’t stay contained to one token, it changes the state of the environment, which changes every decision downstream.
This is the plan-act-observe-decide-recover loop that coding agents, terminal agents, and browser agents all run on. A model can have near-perfect perplexity and still fail a long agentic task if its errors happen to land on decision points rather than filler tokens. Conversely, strong performance on SWE-bench Verified and Terminal-Bench, both of which require multi-step reasoning, tool use, and recovery from mistakes, is a much harder signal to fake than a marginal perplexity score.
What do the SWE-bench and Terminal-Bench numbers actually show?
SWE-bench Verified measures a model’s ability to resolve real GitHub issues by generating patches that pass hidden test suites. It’s one of the most widely cited benchmarks for coding agents because it requires understanding a codebase, localizing a bug, and producing a working fix, not just writing plausible-looking code.
OrcaSAQ2 27B’s reported score of 70.0% places it:
- Behind Claude Sonnet 4.6 (79.6%), Claude Sonnet 4.5 (77.2%), and Gemini 3 (76.2%)
- Ahead of Qwen3-Coder-480B-A35B (69.6%), Gemini 2.5 Pro (63.8%), and GPT-4.1 (54.6%)
Other agents start typing. Remy starts asking.
Scoping, trade-offs, edge cases — the real work. Before a line of code.
That’s a notable result for a 27B-parameter model running from a 12.3 GB checkpoint, especially against Qwen3-Coder-480B-A35B, a model with roughly 18x the total parameters.
Terminal-Bench 2.1 tests agents on real command-line tasks executed inside a terminal environment, a harder proxy for autonomous execution than single-turn code generation. OrcaSAQ2’s 58.4% sits almost exactly between Claude Sonnet 4.6 running Claude Code (58.5%) and Gemini 3 Flash running Gemini CLI (56.9%), and ahead of GPT-5.4 paired with Terminus 2 (54.8%).
One caveat matters here: Terminal-Bench scores are reported per model-and-agent-stack pairing, not per model alone. Claude Opus 4.6 scores differently depending on whether it’s run through Claude Code (70.1%) or Terminus 2 (63.8%), a nearly 7-point swing from harness choice alone. OrcaRouter is explicit that these public numbers shouldn’t be read as a strict model-only ranking, since scaffold, tool access, and reasoning budget all shift the outcome.
Is OrcaSAQ2 worth using instead of a larger hosted model?
The answer depends on what you’re optimizing for. If you need the single highest score on SWE-bench Verified, Claude Sonnet 4.6 and Gemini 3 are still ahead by a meaningful margin (roughly 6 to 10 points). If you need a model that fits on a single 16 GB GPU, keeps 262K context support, and gets within striking distance of frontier coding and terminal agent performance, OrcaSAQ2 is a more unusual proposition: a locally-servable model competing with API-only frontier systems on agentic tasks.
The practical deployment numbers back this up. The 12.3 GB checkpoint leaves roughly 3.7 GB of headroom on a 16 GB GPU for KV cache and runtime overhead, enough for a starting point of about 32K interactive context before tuning. Served through vLLM with MTP speculative decoding enabled, single-stream decode throughput reaches 90.1 tokens/second, a 38% improvement over running without MTP, though MTP trades away KV cache capacity and batched throughput in exchange (220-333 tok/s across 8-16 concurrent streams with MTP off, versus 219-220 tok/s with it on).
For teams building coding assistants, terminal agents, or tool-heavy workflows where self-hosting, cost control, or context length matter more than shaving off the last few benchmark points, that tradeoff is worth evaluating directly against production workloads rather than public leaderboard numbers.
Frequently Asked Questions
What base model is OrcaSAQ2 27B built on?
It’s a quantized version of Qwen/Qwen3.8-27B, using the Qwen3_5ForCausalLM architecture with 64 layers, a hybrid attention design (48 Gated DeltaNet layers plus 16 full-attention layers), and a vocabulary of 248,320 tokens. It inherits the Apache-2.0 license from the base model.
How much smaller is OrcaSAQ2 than the original model?
The checkpoint drops from 54 GB in BF16 to 12.3 GB, a 77.2% reduction, using an average decoder precision of 3.21 bits per weight instead of 16-bit.
Does quantization hurt OrcaSAQ2’s accuracy?
On WikiText-2, perplexity rises only 0.02% relative to BF16 (5.6468 to 5.6482), and the model agrees with BF16’s top predicted token 93.2% of the time. That’s not lossless, roughly 6.8% of token decisions differ from the original model, but it’s a small enough gap that downstream agent benchmarks remain competitive with much larger models.
Can OrcaSAQ2 run on consumer hardware?
- ✕a coding agent
- ✕no-code
- ✕vibe coding
- ✕a faster Cursor
The one that tells the coding agents what to build.
Yes. The 12.3 GB checkpoint is designed to fit on a single 16 GB GPU, served through vLLM with an OpenAI-compatible API, tool calling, and thinking mode enabled by default. Actual usable context depends on GPU memory overhead, batch size, and whether MTP speculative decoding is enabled.
Are the SWE-bench and Terminal-Bench comparisons fair to other models?
Not strictly. These are public reference scores gathered from different sources, each using its own agent scaffold, tool access, and reasoning budget. Terminal-Bench in particular shows the same base model scoring differently depending on which agent harness (Claude Code versus Terminus 2, for example) it’s paired with, so the numbers are directional rather than a controlled head-to-head test.



