Underdog Saluki 27B: Run a 27B Tool-Calling Model on 8GB VRAM
Underdog Saluki 27B compresses Qwen3.8-27B into a 7.89GB GGUF that keeps tool-calling intact, runnable on stock llama.cpp.

What is Underdog Saluki 27B?
Underdog Saluki 27B is a quantized GGUF build of Qwen3.8-27B, compressed from 54GB down to 7.89GB while largely preserving the base model’s tool-calling ability. It’s released by Conway Research (credited to “Underdog”) under Apache 2.0, built on Qwen’s Qwen3.8-27B and ISTA-DASLab’s Qwen3.8-27B-GSQ-RCO-GGUF. The point of the release isn’t just smaller file size, it’s that a 27B-class model with function-calling chops can now run on a single consumer GPU with 8GB of VRAM, using nothing more exotic than stock llama.cpp.
TL;DR
- Underdog Saluki 27B shrinks Qwen3.8-27B from 54GB to a 7.89GB GGUF using IQ2-mix quantization, small enough to fit on an 8GB consumer GPU.
- On the model’s own Underdog Bench (120 tasks from BFCL v4), Saluki scores 88 out of 120 versus 84 for the full-size model, meaning quantization didn’t hurt tool-calling and arguably improved consistency on this test.
- Parallel tool calling also holds up: Saluki passes 42 of 100 BFCL v4 parallel tasks compared to 35 for the uncompressed model.
- Across nine benchmarks total, Saluki retains about 96% average performance of the full model, with the biggest drop showing up in competition math (AIME scores fall from the 90s into the high 70s/low 80s).
- It runs on stock llama.cpp with the
--jinjaflag for chat templating, no custom runtime or forks needed, and exposes an OpenAI-compatible chat API on port 8080. - An optional vision add-on (628MB to 928MB) bolts on image understanding; the base GGUF is text-only.
- Weak spots include letter-level instruction puzzles (palindromes, vowel counting, alphabetizing) and occasional formatting slips in about a fifth of parallel tool-call responses.
Other agents start typing. Remy starts asking.
Scoping, trade-offs, edge cases — the real work. Before a line of code.
How was Qwen3.8-27B compressed to under 8GB?
The model card doesn’t walk through the quantization recipe in detail, but the lineage is clear from the tags and base model references: Saluki is built on top of ISTA-DASLab’s Qwen3.8-27B-GSQ-RCO-GGUF, itself a quantized variant of Qwen’s Qwen3.8-27B. Saluki applies what it calls an “IQ2-mix” scheme, a form of low-bit (around 2-bit) quantization that mixes precision levels across the model’s weights rather than flattening everything uniformly. That’s the standard approach serious GGUF quantizers use to avoid the accuracy collapse that naive 2-bit compression usually causes: keep more bits where they matter (attention layers, certain projections) and squeeze harder where the model tolerates it.
The result is a single file, Underdog-Saluki-27B-1.0-IQ2-mix.gguf, at 7.89GB, down from the 54GB of the full-precision 27B model. That’s roughly a 7x size reduction. For context, a 27B parameter model at full 16-bit precision would need on the order of 54GB just to hold the weights, which is why it normally requires multiple high-end GPUs or heavy offloading. At under 8GB, Saluki fits the VRAM budget of popular consumer cards.
Does the compression hurt tool-calling performance?
Not according to the benchmarks published with the model. Underdog built its own evaluation, “Underdog Bench,” using 120 tasks drawn from the Berkeley Function Calling Leaderboard (BFCL v4), frozen before testing any model to avoid cherry-picking. With thinking mode off and temperature at 0, Saluki passed 88 of 120 tasks. The full-size Qwen3.8-27B, run through the same harness, passed 84. For comparison, a smaller model called Bonsai 2 (5.95GB) passed 70.
On parallel tool calls, a harder test of handling multiple simultaneous function calls, Saluki ran 100 BFCL v4 parallel tasks with the official checker and passed 42, against 35 for the full model.
These numbers come with caveats the model card itself states plainly: 120 tasks is a modest sample size, and a gap of a few tasks either direction is likely just run-to-run variation rather than a meaningful signal. The honest takeaway is that compression didn’t meaningfully degrade tool-calling, and might even land within noise of being slightly better on this particular test set. It’s not evidence that quantized models generically beat full-precision ones at function calling.
What does it give up?
Tool calling held up, but not every capability survived compression equally. Across nine benchmarks, Saluki retains about 96% of the full model’s performance on average, but that average hides a wide spread.
Instruction-following actually improved slightly: IFEval (93.5 vs 91.5 public baseline) and IFBench (72.7 vs 71.0) both edge up. SWE-bench Verified, a coding benchmark built from real GitHub issues, holds close (30 vs 33 of 50 issues fixed).
The real cost shows up in competition math. AIME 2025 drops from 96.7 to 79.2 (average of 4 runs), and AIME 2026 falls from 94.6 to 80.0. That’s roughly 82 to 85% retention on hard math reasoning, the steepest decline of anything tested. MuSR (multi-step reasoning) also takes a hit, down from 79.6 to 67.5.
Remy is new. The platform isn't.
Remy is the latest expression of years of platform work. Not a hastily wrapped LLM.
The model card is specific about where it struggles beyond raw scores: letter-level instruction puzzles, like generating palindromes, applying vowel rules, or sorting letters alphabetically, trip it up more than general instruction-following does. And roughly one in five parallel tool-call responses comes back with small formatting errors, which matters if you’re piping output straight into a parser without validation.
How do you run it locally?
The setup uses stock llama.cpp, no forks or custom builds required. Download the GGUF file with the Hugging Face CLI:
huggingface-cli download ConwayResearch/Underdog-Saluki-27B-1.0 Underdog-Saluki-27B-1.0-IQ2-mix.gguf --local-dir .
Then launch it with llama-server:
llama-server -m Underdog-Saluki-27B-1.0-IQ2-mix.gguf --jinja -ngl 99 -fa on -c 32768
The --jinja flag turns on the Qwen3.8 chat template, which handles both tool-call formatting and the thinking/reasoning mode. The server speaks the OpenAI chat API on port 8080, so anything built against OpenAI’s client libraries should connect with minimal changes.
For image input, download one of the vision add-ons, either the 928MB F16 version or the smaller 629MB Q8_0 version, and pass it with --mmproj:
llama-server -m Underdog-Saluki-27B-1.0-IQ2-mix.gguf --mmproj mmproj-Underdog-Saluki-27B-1.0-F16.gguf --jinja -ngl 99 -fa on -c 32768
For general use, reasoning, and instruction-following, the recommended settings keep thinking on (the default) with temperature 0.6, top_p 0.95, top_k 20. For fast, direct tool calls where you don’t need the model to reason out loud, switch thinking off via "chat_template_kwargs": {"enable_thinking": false} and drop temperature to 0.
Is Underdog Saluki 27B worth using?
For anyone building agents or tool-calling workflows on local hardware, it’s a reasonable pick if 8GB of VRAM is your ceiling and you don’t need competition-grade math reasoning. The core selling point, that a 27B model can be compressed by 7x while matching or slightly beating the full model on function-calling benchmarks, is a real and useful result, backed by a benchmark methodology (frozen task set, consistent harness) that’s more rigorous than typical marketing claims.
Where it’s not the right choice: workloads leaning on complex multi-step reasoning or math competitions, where the gap between compressed and full model is largest. It’s also not a drop-in for strict letter-level text manipulation tasks. And if you’re building pipelines that assume perfectly formatted tool-call output every time, budget for validation logic given the roughly 20% formatting slip rate on parallel calls.
The license is Apache 2.0, inherited from both Qwen3.8-27B and the ISTA-DASLab quantization it builds on, so commercial use isn’t restricted.
Frequently Asked Questions
What base model is Underdog Saluki 27B built on?
It’s a quantized GGUF of Qwen3.8-27B, built via ISTA-DASLab’s Qwen3.8-27B-GSQ-RCO-GGUF intermediate quantization. Both are released under Apache 2.0.
How much VRAM does Underdog Saluki 27B need?
The main GGUF file is 7.89GB, designed to fit within 8GB of VRAM on a consumer GPU. Add 629MB to 928MB more if you load the optional vision add-on.
Can Underdog Saluki 27B process images?
Not by default. The core model is text-only. Image support requires downloading a separate mmproj vision add-on file and loading it alongside the main model with the --mmproj flag in llama.cpp.
How does its tool-calling performance compare to the full-size model?
On Underdog Bench (120 BFCL v4 tasks), it passed 88 versus 84 for the uncompressed 54GB model, and on 100 parallel-call tasks it passed 42 versus 35. The quantized version held up well on function calling specifically, even as it lost ground on math-heavy benchmarks.
What are its biggest weaknesses?
Competition math (AIME) scores drop the most, retaining roughly 82 to 85% of the full model’s performance. It also struggles more with letter-level puzzles like palindromes and alphabetical sorting, and about a fifth of its parallel tool-call outputs have minor formatting issues.





