Qwen3.8-27B quantized Articles
Browse 6 articles about Qwen3.8-27B quantized.

OrcaSAQ2 27B: 3-Bit Quantization That Barely Loses Fidelity
OrcaSAQ2 27B shrinks Qwen3.8-27B from 54GB to 12.3GB at 3-bit precision, holding perplexity loss to +0.02% for agent tasks.

Bonsai 2 27B: A 27B Model That Runs in Under 6GB on a Laptop
Bonsai 2 27B uses ternary quantization to shrink a 27B model to under 6GB while keeping 98% of FP16 performance. Here's how it works.

Bonsai 2 27B Benchmarks: Does Ternary Quantization Actually Hold Up?
Bonsai 2 27B claims 98.2% of FP16 intelligence at ~1.72 bits per weight. Here's how its benchmarks compare to conventional 2-bit and 4-bit builds.

Ternary-Bonsai-2-27B: A 27B Model That Fits in 8.6GB
Bonsai 2 27B shrinks a 27B model to 8.6GB using ternary quantization, keeping 98.2% of FP16 benchmark performance intact.

Qwen3-8-27B at 11.8GB: Do GSQ and RCO Quantization Actually Hold Up?
ISTA's Das Lab shrank Qwen3.8-27B to 11.8GB with new GSQ and RCO quantization. Here's what that means and how it performs locally via llama.cpp.

How to Run Qwen3.8-27B-Escha-W2 Locally on a 24GB GPU
Guide to running Escha-W2, a 2-bit quantized Qwen3.8-27B, on a 24GB GPU with SGLang, covering VRAM tuning and 128k context setup.