Run Underdog Saluki 27B Locally with Llama.cpp: Full Setup Guide
Download, configure, and serve Underdog Saluki 27B, a sub-8GB Qwen3.8-27B GGUF quant, with llama.cpp and tool calling intact.

What is Underdog Saluki 27B?
Underdog Saluki 27B is a 2-bit GGUF quantization of Qwen3.8-27B, compressed to 7.89 GB from the original 54 GB, built specifically to preserve tool calling and function calling accuracy. It runs on stock llama.cpp with no custom fork or patches required, and it scores higher than the full-size model on the Berkeley Function Calling Leaderboard-derived benchmark its maker used for testing. For anyone who wants a 27B-class model with agentic capabilities on a single consumer GPU, it’s a practical option.
TL;DR
- Underdog Saluki 27B is a 7.89 GB GGUF quant of Qwen3.8-27B, roughly one-seventh the size of the original 54 GB model.
- It’s released by Underdog (Conway Research on Hugging Face) under Apache 2.0, built on Qwen3.8-27B and ISTA-DASLab’s Qwen3.8-27B-GSQ-RCO-GGUF quantization work.
- On a 120-task tool calling benchmark drawn from BFCL v4, it scored 88 versus 84 for the full-size model, and 42 versus 35 on parallel tool calls.
- It runs on stock llama.cpp with the
--jinjaflag enabled, which activates the Qwen3.8 chat template for tool calls and thinking. - An optional vision add-on (
mmprojfile, 629 MB to 928 MB) adds image support through the standard OpenAI chat API. - It retains 96% average performance across 9 benchmarks compared to the full model, but loses more ground on competition math (AIME) and letter-level instruction puzzles.
- Total download for text plus vision is under 9 GB, small enough to fit comfortably on a single modern GPU with room to spare.
How do you download and run it?
The model ships as a single GGUF file plus optional vision components, hosted on Hugging Face under ConwayResearch/Underdog-Saluki-27B-1.0. Grab the main weights with the Hugging Face CLI:
huggingface-cli download ConwayResearch/Underdog-Saluki-27B-1.0 Underdog-Saluki-27B-1.0-IQ2-mix.gguf --local-dir .
Then launch it with llama-server:
llama-server -m Underdog-Saluki-27B-1.0-IQ2-mix.gguf --jinja -ngl 99 -fa on -c 32768
That command loads all layers onto GPU (-ngl 99), turns on flash attention (-fa on), and sets a 32768-token context window (-c 32768). The server exposes an OpenAI-compatible chat API on port 8080, so anything built against OpenAI’s client libraries should talk to it with just a base URL change.
The --jinja flag is the one detail easy to miss. It activates the Qwen3.8 chat template, which is what makes tool calls and the thinking/non-thinking toggle work correctly. Skip it and you lose the behavior that makes this quant worth using over a generic chat model.
Why does tool calling survive quantization here?
Most aggressive quantization degrades a model’s ability to follow structured output formats, which is exactly what tool calling and function calling depend on. Underdog Saluki 27B is built from ISTA-DASLab’s Qwen3.8-27B-GSQ-RCO-GGUF quantization, and the “Saluki” tuning on top of it appears aimed specifically at retaining that structured behavior rather than general perplexity.
The numbers back this up. On Underdog Bench, a set of 120 tasks drawn from the Berkeley Function Calling Leaderboard (BFCL v4) and frozen before testing began, the quantized model scored 88 out of 120 with thinking turned off and temperature at 0, compared to 84 for the uncompressed 54 GB original. On parallel tool calls (handling multiple function calls in a single turn), it scored 42 against 35 for the full model. A third model in the same size class, Bonsai 2 at 5.95 GB, scored 70 on the same test, giving some sense of where a smaller, less tool-focused quant lands by comparison.
It’s worth noting these results use a smaller, internally frozen benchmark rather than the full BFCL leaderboard, and the model card is upfront that a handful of points of difference on 120 tasks could reflect run-to-run variation rather than a real capability gap.
How do you add vision support?
The base GGUF file is text-only. To handle images, download one of the two vision add-on files, called mmproj files, which pair with the main model:
huggingface-cli download ConwayResearch/Underdog-Saluki-27B-1.0 mmproj-Underdog-Saluki-27B-1.0-F16.gguf --local-dir .
There’s a choice between the F16 version (928 MB) and a Q8_0 version (629 MB), trading a bit of file size for precision. Pass the add-on at launch with --mmproj:
llama-server -m Underdog-Saluki-27B-1.0-IQ2-mix.gguf --mmproj mmproj-Underdog-Saluki-27B-1.0-F16.gguf --jinja -ngl 99 -fa on -c 32768
Once running, images go through the same OpenAI-style chat API calls as any multimodal model, no special client code needed. Total disk footprint for text plus vision stays under 9 GB even with the larger F16 add-on, which is a small price for image understanding on top of an already compact model.
How should you configure thinking and sampling?
- ✕a coding agent
- ✕no-code
- ✕vibe coding
- ✕a faster Cursor
The one that tells the coding agents what to build.
Qwen3.8 ships with a “thinking” mode on by default, where the model reasons through a problem before producing a final answer. Saluki inherits this behavior and the same per-request switch. The model card lays out two configurations depending on what you’re doing:
For general use, reasoning, and following instructions, leave thinking on (the default) and set temperature to 0.6, top_p to 0.95, and top_k to 20.
For fast, direct tool calls where you don’t want the model deliberating out loud, disable thinking by passing "chat_template_kwargs": {"enable_thinking": false} in the request body, and drop temperature to 0 for deterministic output.
This matters in practice because with thinking on, the model can reason at length before it answers, adding latency that’s wasteful for a straightforward function call but useful for a harder reasoning task.
Is Underdog Saluki 27B worth running over the full-size model?
It depends on what you’re optimizing for. The quant retains 96% average performance across 9 benchmarks relative to the 54 GB original, which is a strong retention rate for a model compressed to roughly 15% of its original size. It actually outperforms the full model on tool calling, parallel tool calls, IFEval, and IFBench, likely because the quantization and tuning process specifically targeted structured output reliability.
The tradeoffs show up on math and some instruction-following edge cases. On AIME 2025 and 2026 competition math benchmarks, Saluki keeps about 82 to 85% of the full model’s score, a real but not catastrophic gap. It’s also weaker on letter-level instruction puzzles like palindrome generation, vowel counting rules, or alphabetical ordering tasks, the kind of fine-grained symbolic manipulation that quantization tends to hurt first. And about a fifth of parallel tool call responses show small formatting slips, which is worth testing for if your pipeline depends on strict output parsing.
For anyone building tool-using agents, coding assistants, or instruction-following systems on hardware that can’t fit a 54 GB model, the tradeoff looks favorable. For heavy competition-math or precision symbolic work, the full-size model or a less aggressive quant is the safer choice.
Frequently Asked Questions
What hardware do you need to run Underdog Saluki 27B?
The model file is 7.89 GB, small enough to run on a single consumer GPU with 8 GB or more of VRAM when fully offloaded with -ngl 99. Adding the vision component brings total VRAM needs up by roughly 0.6 to 0.9 GB depending on which mmproj file you choose.
Does it support tool calling out of the box?
Yes, as long as you launch llama-server with the --jinja flag, which enables the Qwen3.8 chat template responsible for handling tool calls and the thinking toggle. Without it, structured tool call formatting won’t work correctly.
What license is Underdog Saluki 27B released under?
Apache 2.0, the same license as its base models, Qwen3.8-27B and ISTA-DASLab’s Qwen3.8-27B-GSQ-RCO-GGUF quantization. That makes it usable in commercial projects without additional licensing negotiation.
How does it compare to the full 54 GB Qwen3.8-27B model?
It retains about 96% average performance across 9 benchmarks while using roughly 15% of the storage. It beats the full model on tool calling and instruction-following benchmarks but trails on competition math (AIME) and some letter-level instruction tasks.
Can it process images?
Not by default. The core GGUF file is text-only. Image support requires downloading a separate vision add-on file (either 629 MB or 928 MB) and passing it to llama-server with the --mmproj flag.



