ternary quantization LLM
ternary quantization LLM Articles
Browse 3 articles about ternary quantization LLM.

Run Bonsai 2 27B Locally on a Mac: Ternary Quantization Explained
Bonsai 2 27B compresses a 27B reasoning model to 8.6GB with ternary weights, hitting ~47 tok/s on an M5 Max MacBook via MLX.
Bonsai 2 27Bternary quantization LLMrun 27B model on Mac

Run Bonsai 2 27B Locally on a Mac: Full Setup Guide
How to run Ternary-Bonsai-2-27B, an 8.6GB ternary-quantized 27B model, locally on Apple Silicon with near-FP16 quality and speed.
Bonsai 2 27B localternary quantization LLMrun 27B model on Mac

Ternary-Bonsai-2-27B: A 27B Model That Fits in 8.6GB
Bonsai 2 27B shrinks a 27B model to 8.6GB using ternary quantization, keeping 98.2% of FP16 benchmark performance intact.
Ternary-Bonsai-2-27Bternary quantization LLM1-bit LLM