Underdog Saluki 27B: Tool Calling Gains vs Full-Size Qwen3.8 Losses
Underdog Saluki 27B compresses Qwen3.8-27B to under 8GB. Here's where it beats the full model on tool calling and where it falls short.

What is Underdog Saluki 27B?
Underdog Saluki 27B 1.0 is a compressed version of Qwen3.8-27B, shrunk from 54 GB down to 7.89 GB using an IQ2-mix quantization format, while specifically tuning to preserve tool-calling accuracy. On the Berkeley Function Calling Leaderboard tasks used for its benchmark suite, it actually scores higher than the full-size model on both single and parallel tool calls, though it gives up ground on competition math and some coding tasks. It runs on stock llama.cpp with no custom runtime required.
TL;DR
- Underdog Saluki 27B compresses Qwen3.8-27B by roughly 7x (54 GB to 7.89 GB) while retaining 96% of the parent model’s performance across nine benchmarks.
- On tool calling, the compressed model scored 88 out of 120 on BFCL v4 tasks versus 84 for the full-size model, and 42 versus 35 on parallel tool calls.
- The model loses the most ground on competition math, keeping only about 82 to 85% of the full model’s AIME 2025 and 2026 scores.
- Instruction following (IFEval, IFBench) held steady or slightly improved, but the model struggles specifically with letter-level puzzles like palindromes and alphabetical ordering.
- It ships as a standard GGUF file that runs on unmodified llama.cpp, with an optional vision add-on under 1 GB for image support.
- The release is Apache 2.0 licensed, built on Qwen3.8-27B and a prior GGUF quantization effort from ISTA-DASLab, with benchmark tasks sourced from BFCL.
How does Underdog Saluki 27B compare to the full-size model?
The benchmark table tells a split story. Saluki wins clearly on tool calling: 88 versus 84 out of 120 BFCL v4 tasks, and 42 versus 35 on the 100-task parallel calling set. Instruction following also edges upward, with IFEval at 93.5 versus a public score of 91.5 for the full model, and IFBench at 72.7 versus 71.0.
Everywhere else, the full-size model holds an advantage. SWE-bench Verified drops from 33 to 30 fixed issues out of 50. MBPP+ falls from 83.9 to 78.0. MuSR drops more sharply, from 79.6 to 67.5. The widest gap shows up in competition math: AIME 2025 average-at-4 falls from 96.7 to 79.2, and AIME 2026 from 94.6 to 80.0. That’s the model retaining only 82 to 85% of its parent’s math performance, the single largest regression in the whole suite.
The model card frames the overall picture as 96% average retention across nine benchmarks, which is consistent with the pattern: small losses or gains on most tasks, one steep drop on extended multi-step math reasoning.
Why does a compressed model outperform the full-size one on tool calling?
This is the detail worth sitting with. Compression usually costs accuracy across the board, so beating the parent model on any benchmark is unusual. The model card attributes this to deliberate tuning aimed at keeping tool calling intact rather than just shrinking weights uniformly. In practice, that likely means the quantization and any fine-tuning choices prioritized the structured-output and function-call formatting behaviors that BFCL tests, possibly at some cost to other capabilities.
It’s also worth noting the test conditions: tool-calling numbers were measured with thinking turned off and temperature set to zero, which removes a major source of run-to-run variance. Berkeley Function Calling Leaderboard v4 tasks are designed to test whether a model picks the right function, fills parameters correctly, and formats the call properly. These are narrower, more mechanical tasks than open-ended math reasoning, which may make them more tractable to preserve under aggressive quantization than something like AIME problem-solving, which depends on long chains of exact intermediate reasoning.
The model card also flags that about a fifth of parallel-call responses have small formatting slips, so the win on tool calling isn’t flawless, just a net positive on the aggregate pass count.
What are the practical tradeoffs of running Saluki instead of the full model?
The headline tradeoff is size versus math accuracy. At 7.89 GB, Saluki fits on hardware that the 54 GB full-size Qwen3.8-27B simply can’t touch without aggressive offloading or multi-GPU setups. For workloads centered on tool use, agents, and instruction following, the benchmark numbers suggest you’re not giving up accuracy to get that size reduction, you may even gain a little.
For workloads that lean on multi-step mathematical or logical reasoning, the gap is real. A roughly 15 to 18 point drop on AIME-style problems is a meaningful regression if your use case involves complex quantitative reasoning. The model card also calls out a specific weak spot: letter-level instruction puzzles such as palindrome checks, vowel-counting rules, or alphabetical ordering. These are tasks that require precise character-level tracking, which seems to be a casualty of the compression process in a way that broader instruction-following isn’t.
Built like a system. Not vibe-coded.
Remy manages the project — every layer architected, not stitched together at the last second.
Vision support is bolted on separately rather than baked into the main file. If image input matters, you need the extra mmproj file (928 MB in F16 or 629 MB in Q8_0), which still keeps the total footprint well under the full model’s size.
Is Underdog Saluki 27B worth using over the full-size Qwen3.8-27B?
It depends on what the workload actually needs. If the primary use case is agentic: calling APIs, parsing structured outputs, following formatted instructions, chaining tool calls, the benchmark data supports using the compressed model without hesitation. It scores as well or better than its parent on exactly those tasks, at a fraction of the memory footprint, which matters for anyone deploying on constrained hardware or running multiple model instances concurrently.
If the workload leans on competition-level math, dense coding benchmarks like SWE-bench, or tasks requiring long multi-hop reasoning (MuSR), the full-size model is still meaningfully stronger. The 7.89 GB version keeps the shape of the parent model’s behavior but loses sharpness on the hardest reasoning tasks.
The model card’s own framing, 96% average retention across the full benchmark suite, is a reasonable way to think about it: most capabilities survive the compression close to intact, with tool calling as a specific bright spot and competition math as the clear weak point.
Frequently Asked Questions
What quantization format does Underdog Saluki 27B use?
It uses an IQ2-mix GGUF quantization, a 2-bit mixed-precision format designed to run on stock llama.cpp without any custom runtime or patches.
How much smaller is Saluki than the full Qwen3.8-27B model?
The compressed model is 7.89 GB compared to 54 GB for the full-size Qwen3.8-27B, a reduction of roughly 7x in file size.
Does Underdog Saluki 27B support image input?
Not by default. The core file is text-only. An optional vision add-on file (928 MB in F16 or 629 MB in Q8_0 precision) can be loaded alongside it via the --mmproj flag to enable image inputs through the OpenAI chat API.
What benchmark was used to measure tool-calling accuracy?
Underdog Bench, a set of 120 tasks drawn from the Berkeley Function Calling Leaderboard (BFCL v4), frozen before testing began. Tests were run with thinking disabled and temperature set to zero.
What is the biggest weakness of the compressed model?
Competition math. On AIME 2025 and AIME 2026 benchmarks, Saluki retains only about 82 to 85% of the full model’s score, the largest gap of any benchmark in the comparison. It also underperforms on letter-level instruction tasks like palindrome or alphabetical-order puzzles.

