Skip to main content
MindStudio
Pricing
BlogAbout
My Workspace
run Qwen3.8-27B locallyVLLM reasoning effortQwen3.8 chat template

Running Qwen3.8-27B Locally: Reasoning Effort, VLLM and Speed Tradeoffs

How to serve Qwen3.8-27B locally with VLLM, tune reasoning effort (low/medium/X-high), and understand its speculative decoding behavior.

Edited by Luis Chavez-Mattos, Director of Product RSS
Running Qwen3.8-27B Locally: Reasoning Effort, VLLM and Speed Tradeoffs

What is reasoning effort on Qwen3.8-27B?

Qwen3.8-27B ships with a built-in setting called reasoning effort, controlled through the chat template when you serve the model. It has three levels: low, medium, and X-high. X-high is the default out of the box, which means the model is, by default, configured to think as much as possible before answering. That default matters a lot if you’re running this locally, because every extra thinking token is time you wait and compute you pay for.

TL;DR

  • Reasoning effort is a chat-template setting on Qwen3.8-27B with three levels (low, medium, X-high), and X-high is the out-of-the-box default, meaning you get maximum thinking length unless you change it.
  • On single-GPU local serving you’re typically decode bound, so thinking tokens map almost directly to wall-clock seconds. Fewer tokens out effectively means a faster answer.
  • Speculative decoding on this model gets faster, not just shorter, at lower reasoning effort. Thinking Cap’s own numbers show tokens-per-step rising from about 2.6 at X-high to over 3 at medium and low.
  • Dropping the base model straight to medium cuts thinking roughly in half but costs about nine points of accuracy, which is why the community has built fine-tunes that try to recover that accuracy.
  • Thinking Cap, from Bottlecap AI, trims about 37% of thinking tokens at X-high for under a one-point accuracy drop, and is meant as a drop-in replacement for the base model at X-high.
  • Swift 1.5, from Yukus AI, is more aggressive, claiming up to 58.5% fewer thinking tokens, with published training data and the widest range of runtime formats (GGUF, MLX, NVFP4, AMD builds).
  • Qwen Pi takes a different approach entirely: instead of shortening X-high, it’s fine-tuned on a minimal coding-agent harness (Pi) so that its low and medium settings become good enough to match the base model’s X-high.

Other agents ship a demo. Remy ships an app.

UI
React + Tailwind ✓ LIVE
API
REST · typed contracts ✓ LIVE
DATABASE
real SQL, not mocked ✓ LIVE
AUTH
roles · sessions · tokens ✓ LIVE
DEPLOY
git-backed, live URL ✓ LIVE

Real backend. Real database. Real auth. Real plumbing. Remy has it all.

How do you actually serve Qwen3.8-27B with VLLM?

The practical path is to serve the model through VLLM and set the reasoning effort inside the chat template that VLLM applies at request time. That template setting is what tells the model how hard to “try” before producing an answer. Low and medium settings produce shorter chains of thought; X-high lets the model restate the question, double-check itself, and reconsider before committing to a final answer; useful behavior for hard problems, wasteful for easy ones.

Because reasoning effort is a template-level switch rather than a different checkpoint, you can serve one set of weights and let different callers request different effort levels depending on the task. A simple factual lookup doesn’t need X-high. A multi-step math or agentic coding task might.

What hardware do you need to run Qwen3.8-27B locally?

Qwen3.8-27B is explicitly positioned as one of the better open models you can serve on “reasonably small hardware,” in contrast to Qwen’s larger “3.8 next” family, which pushes you toward bigger GPUs or multi-GPU setups. For many use cases the smaller 3.8-27B is the more practical local option precisely because it stays within single-GPU territory, while the larger sibling may be better for specific tasks at the cost of needing serious multi-unit hardware.

The key operational detail: when you run this on a single GPU, you are usually decode bound. That means the bottleneck isn’t loading the prompt, it’s generating tokens one at a time. Since reasoning or “thinking” tokens are generated the same way as the final answer tokens, the length of the chain of thought translates almost directly into how many seconds you wait for your result. This is why reasoning effort settings aren’t just an accuracy lever, they’re a latency lever.

Why does lowering reasoning effort make each token faster too?

Qwen3.8-27B uses multi-token prediction heads for speculative decoding, a technique where the model proposes several tokens at once and verifies them, rather than generating strictly one at a time. That speculation works better on shorter, lower-effort reasoning traces than on long X-high traces. Thinking Cap’s model card reports roughly 2.6 tokens accepted per step at X-high, climbing to more than 3 tokens per step at medium and low settings.

The practical implication: cutting reasoning effort doesn’t just reduce the number of tokens generated, it increases the throughput of each token generated. The two effects compound, so a drop from X-high to a well-tuned lower setting can feel considerably faster than the raw token-count reduction would suggest.

Is it worth fine-tuning Qwen3.8-27B to think less?

The tradeoff is real but not fixed. If you simply force the stock model down to medium effort, you remove about half its thinking tokens but give up roughly nine points of accuracy on hard benchmarks. That’s a blunt instrument. Several teams have built fine-tunes specifically to make that tradeoff less punishing, with different strategies and different results.

Everyone else built a construction worker.
We built the contractor.

🦺
CODING AGENT
Types the code you tell it to.
One file at a time.
🧠
CONTRACTOR · REMY
Runs the entire build.
UI, API, database, deploy.

Thinking Cap (Bottlecap AI) deliberately avoided teaching the model anything new. The objective was narrow: reduce how many thinking tokens it takes to land on the same answer at X-high, without touching instruction-following, safety, or general capability. Across 12 benchmarks they report an average 37% reduction in thinking tokens, with average accuracy dropping only from about 86.6 to 85.8. On GPQA Diamond specifically, average thinking tokens dropped from about 12,800 to 7,300. On a long-context retrieval benchmark, thinking dropped 39% while accuracy actually improved slightly. The weak spot is agentic traces, where token counts shorten but only by about 11%. Thinking Cap also sells enterprise versions tuned specifically for medium or low effort, and licensing moved from Apache 2.0 (on the earlier Qwen3.6 version) to a small-business-plus-personal-use license.

Swift 1.5 (Yukus AI) is more aggressive, targeting up to 58.5% fewer thinking tokens. It was built in two stages: Swift 1.0 identified specific “overthinking” tokens in Qwen’s reasoning (phrases like self-questioning or restating the problem) and penalized them directly, then merged that with Thinking Cap. Swift 1.5 scaled up reinforcement learning and on-policy distillation, focused heavily on long-horizon agentic and coding tasks. On LiveCodeBench, it reportedly improves accuracy (from just under 77% to 81.7%) while using fewer tokens (around 8,400 versus 11,200) than the base model. Swift also claims savings hold across every effort level, including a slight accuracy gain at low effort with about 29% fewer tokens. It’s released under a custom license free up to $1 million in revenue, and it’s the most widely packaged of the three, with GGUF, MLX, NVFP4, and AMD builds plus a no-key research API.

Qwen Pi takes the opposite approach. Rather than shortening X-high, it targets making low and medium effort genuinely usable inside Pi, a deliberately minimal open-source coding agent harness (read, write, edit, bash tools, a system prompt under 1,000 tokens). It was built via supervised fine-tuning on successful real Pi sessions only, independent checks for working versus broken code, and reinforcement learning with a custom reward aimed specifically at improving low and medium effort while leaving X-high to focus on correctness. Checkpoints were selected based on actual agent task results, not just training loss. Qwen Pi reportedly lets medium effort match the base model’s X-high performance on Terminal Bench while using about 41% fewer output tokens. It’s released under Apache 2.0.

Frequently Asked Questions

What does reasoning effort actually control in Qwen3.8-27B?

It controls how long the model’s internal chain of thought is before it commits to a final answer. It’s set in the chat template at serving time (for example via VLLM) and has three levels: low, medium, and X-high, with X-high as the default.

Does lowering reasoning effort hurt accuracy?

Yes, by default. Dropping the stock model straight to medium cuts thinking roughly in half but costs about nine points of accuracy on hard benchmarks. Fine-tunes like Thinking Cap and Swift 1.5 exist specifically to reduce that penalty.

Which fine-tune should I use for coding agents?

Swift 1.5 focuses heavily on long-horizon coding and agentic tasks and reports both fewer tokens and higher LiveCodeBench accuracy than the base model. Qwen Pi is narrower still, tuned specifically for the Pi agent harness rather than general coding use.

Why does speculative decoding get better at lower reasoning effort?

Multi-token prediction heads accept more tokens per step on shorter, lower-effort traces. Reported figures go from about 2.6 tokens per step at X-high to over 3 at medium and low, so lower effort settings are faster per token, not just shorter overall.

Are these fine-tuned models free to use?

Licensing varies. Qwen Pi is Apache 2.0. Swift 1.5 uses a custom license free up to $1 million in revenue. Thinking Cap’s X-high release uses a small-business-plus-personal-use license, with separate enterprise versions sold for medium and low effort tuning.

Editorial standards

Presented by MindStudio

No spam. Unsubscribe anytime.