Skip to main content
MindStudio
Pricing
BlogAbout
My Workspace
Qwen3.8-27B quantized

Qwen3.8-27B quantized Articles

Browse 6 articles about Qwen3.8-27B quantized.

OrcaSAQ2 27B: 3-Bit Quantization That Barely Loses Fidelity

OrcaSAQ2 27B shrinks Qwen3.8-27B from 54GB to 12.3GB at 3-bit precision, holding perplexity loss to +0.02% for agent tasks.

OrcaSAQ2 27B3-bit quantizationQwen3.8-27B quantized

Bonsai 2 27B: A 27B Model That Runs in Under 6GB on a Laptop

Bonsai 2 27B uses ternary quantization to shrink a 27B model to under 6GB while keeping 98% of FP16 performance. Here's how it works.

Bonsai 2 27Bternary quantizationMLX 2-bit model

Bonsai 2 27B Benchmarks: Does Ternary Quantization Actually Hold Up?

Bonsai 2 27B claims 98.2% of FP16 intelligence at ~1.72 bits per weight. Here's how its benchmarks compare to conventional 2-bit and 4-bit builds.

Bonsai 2 27B benchmarksternary weight LLMlow-bit quantization benchmark

Ternary-Bonsai-2-27B: A 27B Model That Fits in 8.6GB

Bonsai 2 27B shrinks a 27B model to 8.6GB using ternary quantization, keeping 98.2% of FP16 benchmark performance intact.

Ternary-Bonsai-2-27Bternary quantization LLM1-bit LLM

Qwen3-8-27B at 11.8GB: Do GSQ and RCO Quantization Actually Hold Up?

ISTA's Das Lab shrank Qwen3.8-27B to 11.8GB with new GSQ and RCO quantization. Here's what that means and how it performs locally via llama.cpp.

Qwen3.8-27B quantizedGSQ RCO quantizationrun Qwen3.8-27B locally

How to Run Qwen3.8-27B-Escha-W2 Locally on a 24GB GPU

Guide to running Escha-W2, a 2-bit quantized Qwen3.8-27B, on a 24GB GPU with SGLang, covering VRAM tuning and 128k context setup.

Escha-W2 local setuprun Qwen3.8-27B RTX 4090SGLang serve.sh