ternary quantization
ternary quantization Articles
Browse 3 articles about ternary quantization.

Bonsai 2 27B: A 27B Model That Runs in Under 6GB on a Laptop
Bonsai 2 27B uses ternary quantization to shrink a 27B model to under 6GB while keeping 98% of FP16 performance. Here's how it works.
Bonsai 2 27Bternary quantizationMLX 2-bit model

Bonsai 2 27B Benchmarks: Does Ternary Quantization Actually Hold Up?
Bonsai 2 27B claims 98.2% of FP16 intelligence at ~1.72 bits per weight. Here's how its benchmarks compare to conventional 2-bit and 4-bit builds.
Bonsai 2 27B benchmarksternary weight LLMlow-bit quantization benchmark

Tencent's Sherry Quantization: How a 1.5TB Model Shrank to 214GB
Tencent's Angel Slim toolkit uses ternary Sherry quantization to compress a 770B model 7x, from 1.5TB to 214GB, with barely any quality loss.
Angel SlimSherry quantizationTencent AI compression