Thinking Cap vs Swift 1.5 vs Qwen Pi: Best Qwen3.8-27B Fine-Tune?
Three teams fine-tuned Qwen3.8-27B to think less without losing accuracy. Here's how Thinking Cap, Swift 1.5, and Qwen Pi actually compare.

What is the fastest Qwen3.8-27B fine-tune for reasoning tasks?
There’s no single winner. Thinking Cap, Swift 1.5, and Qwen Pi all start from the same base weights and all cut the number of reasoning tokens Qwen3.8-27B burns through before answering, but they target different jobs. Thinking Cap is the safest drop-in replacement for general reasoning and knowledge tasks. Swift 1.5 pushes the hardest on coding and agentic workloads. Qwen Pi is narrowly built for one coding harness and isn’t meant for general use at all.
TL;DR
- Qwen3.8-27B ships with three reasoning effort settings (low, medium, x-high) and defaults to x-high, which means most people are running the slowest, most token-hungry version without realizing it.
- Thinking Cap, built by Bottlecap AI, cuts thinking tokens by about 37% on average across 12 benchmarks while accuracy drops less than a full point, making it close to a free lunch for x-high workloads.
- Swift 1.5 from Yucus AI claims up to 58.5% fewer thinking tokens and actually scores higher than the base model on LiveCodeBench, with fewer tokens used to get there.
- Qwen Pi takes the opposite approach: instead of shrinking x-high, it trains the low and medium settings to perform as well as x-high, specifically inside the Pi coding agent harness.
- Token count maps almost directly to wait time on single-GPU local setups, since decoding is the bottleneck, so a 37-58% token cut is a roughly proportional cut in how long you sit waiting for an answer.
- Shorter reasoning traces also decode faster per token, because speculative decoding’s multi-token prediction heads hit higher acceptance rates on short chains than long ones.
- Licensing differs sharply: Thinking Cap uses a small-business-plus-personal license, Swift 1.5 is free under $1M revenue, and Qwen Pi is the only one of the three still under Apache 2.0.
Why does Qwen3.8-27B need a fine-tune for this at all?
Qwen3.8-27B already exposes a built-in “reasoning effort” dial with three settings: low, medium, and x-high. The surprising part is that x-high is the default. That means out of the box, the model is configured to think as hard as possible on every query, generating long chains of thought that restate the question, double-check themselves, and occasionally backtrack (“wait, let me reconsider”) before committing to an answer.
Dialing the setting down to medium or low does cut tokens, but it costs accuracy. According to Bottlecap AI’s own published comparison, dropping from x-high to medium removes about half the thinking tokens but gives up roughly nine points of accuracy. That’s the tradeoff these fine-tunes are trying to break: get the token savings of a lower effort setting without the accuracy penalty that normally comes with it.
This matters more on local hardware than it might in a cloud API. When you’re running on a single GPU, you’re typically decode-bound, meaning the number of tokens generated maps almost one-to-one with wall-clock wait time. A chain of thought that’s twice as long takes roughly twice as long to produce the final answer. There’s a second multiplier too: speculative decoding (used by multi-token prediction heads) tends to accept more tokens per step on short reasoning traces than on long, rambling ones. Thinking Cap’s own benchmarks reportedly show speculative decoding acceptance rising from around 2.6 tokens per step at x-high to over 3 tokens per step at medium or low. Shorter traces aren’t just fewer tokens, each token also arrives faster.
How was Thinking Cap built, and what is it good for?
Thinking Cap comes from Bottlecap AI, a small lab in Prague co-founded by Tomas Mikolov, known for being the first author on the Word2Vec paper and for early work on sequence-to-sequence modeling at Google, both foundational to how modern language models represent and generate text. The team previously released a Thinking Cap version built on an earlier Qwen 3.6 model, and followed community requests to build one for Qwen3.8-27B.
Bottlecap hasn’t published its exact training recipe, but it has been explicit about the objective: reduce how many thinking tokens the model needs to reach the same answer, without teaching it anything new and without changing instruction-following, safety behavior, or general knowledge. The training stays anchored to x-high and leans on harder benchmark categories (math, long-context retrieval, agentic tasks) rather than optimizing for brevity directly. The idea is to teach the model when it’s done, not to teach it to be terse.
The reported numbers: 37% fewer thinking tokens on average across 12 benchmarks, with average accuracy slipping from about 86.6% to 85.8%. On GPQA Diamond specifically, thinking tokens reportedly drop from around 12,800 to about 7,300. On a long-context retrieval benchmark, thinking drops 39% while accuracy actually ticks up slightly. The one soft spot is agentic traces, where token reduction is smaller, around 11%.
One coffee. One working app.
You bring the idea. Remy manages the project.
Practically, Thinking Cap functions as a near drop-in replacement for the stock model at x-high: same behavior, same knowledge, noticeably faster. The license changed from Apache 2.0 (on the earlier 3.6-based version) to a small-business-plus-personal-use license, since Bottlecap also sells enterprise variants tuned for medium and low effort settings.
How does Swift 1.5 differ from Thinking Cap?
Swift 1.5, built by Yucus AI, goes after the same problem more aggressively and with a different method. The earlier Swift 1.0 release identified specific tokens in Qwen’s reasoning chains associated with overthinking (phrases like “wait, let me ask what the user wants” or “let me reconsider”) and fine-tuned the model to penalize those tokens directly rather than penalizing total length. That version also incorporated a transfer component from the original Thinking Cap model, giving Swift a shared lineage with Bottlecap’s work.
Swift 1.5 scales this up with reinforcement learning and on-policy distillation, with the training focus shifted toward long-horizon agentic and coding tasks. On LiveCodeBench, Yucus AI reports the base model scoring just under 77% using around 11,200 tokens, while Swift 1.5 scores 81.7% using only about 8,400 tokens, an increase in accuracy alongside a drop in token count. The team claims the savings hold across all three effort levels, including a version at “low” that scores slightly above the base model while using about 29% fewer tokens.
Swift 1.5 is also the most accessible of the three in terms of deployment: standard GGUF builds for llama.cpp, small variants sized for modest GPUs, MLX builds, an NVFP4 version, an AMD build, and a free research API with no key required. The license is a custom Yucus AI license, free for revenue under $1 million, with a direct contact requirement above that threshold.
What makes Qwen Pi different from the other two?
Qwen Pi is the outlier, both in approach and in licensing (it’s the only one of the three still Apache 2.0). It’s built specifically to run inside Pi, a deliberately minimal open-source coding agent with a small, fixed toolset (read, write, edit, bash) and a system prompt reportedly under a thousand tokens.
Rather than shrinking x-high reasoning, Qwen Pi trains the low and medium settings to perform at x-high levels, specifically within Pi’s narrow harness. The training recipe has three parts: supervised fine-tuning on real Pi sessions, but only successful ones, with working versus broken code verified separately; reinforcement learning using GRPO with a custom reward tied to reasoning efficiency, applied mainly at low and medium settings while x-high is left alone to focus on correctness; and checkpoint selection based on actual agent task outcomes rather than training loss alone, since improving loss doesn’t always translate into better agent behavior.
The reported headline is that Qwen Pi’s medium setting matches the base model’s x-high performance on Terminal-Bench while using about 41% fewer output tokens. The likely reason this works well is that Pi’s minimal toolset and short system prompt produce very uniform session data, which is easier to fine-tune against than the more varied traces from general-purpose agent use. Qwen Pi isn’t meant as a general chat or reasoning model. It’s a preview of a pattern likely to show up more often: models fine-tuned for one specific harness rather than for general use.
Other agents ship a demo. Remy ships an app.
Real backend. Real database. Real auth. Real plumbing. Remy has it all.
Is it worth switching from stock Qwen3.8-27B to one of these?
For anyone running Qwen3.8-27B locally or at scale, yes, in most cases. Thinking Cap gives a near-free accuracy tradeoff for general use at x-high. Swift 1.5 is the strongest pick for coding and agentic work, where it reportedly beats the base model’s accuracy while using fewer tokens. Qwen Pi only makes sense if you’re actually running the Pi agent harness; outside that context it’s not built for your workload.
Frequently Asked Questions
What is Qwen3.8-27B’s default reasoning setting?
It defaults to “x-high,” the most token-intensive of its three built-in reasoning effort levels (low, medium, x-high), meaning most deployments run the slowest configuration unless explicitly changed.
Does reducing reasoning tokens hurt accuracy?
It can. Simply dropping Qwen3.8-27B’s built-in effort setting from x-high to medium removes about half the thinking tokens but costs roughly nine points of accuracy, according to Bottlecap AI’s reported comparisons. The fine-tunes covered here aim to avoid that penalty through targeted training rather than just turning the dial down.
Which fine-tune is best for coding tasks?
Swift 1.5 is the one built with the heaviest focus on coding and agentic benchmarks, reportedly outperforming the base model on LiveCodeBench while using fewer tokens.
Why does token count matter so much for local inference?
On single-GPU setups, generation is typically decode-bound, so the number of tokens produced maps closely to wall-clock wait time. Shorter reasoning chains also tend to benefit more from speculative decoding, compounding the speed gain.
Are any of these fine-tunes free to use commercially?
Qwen Pi is Apache 2.0 licensed. Swift 1.5 is free for revenue under $1 million under its own custom license. Thinking Cap uses a small-business-plus-personal-use license, with enterprise variants sold separately.




