GGUF quantization
GGUF quantization Articles
Browse 3 articles about GGUF quantization.

MiniCPM5-2B: Running OpenBMB's 2B On-Device Model Locally
MiniCPM5-2B packs 2.5B params, 131K context, and 4B-beating benchmarks. Here's how it works and how to run it via GGUF, MLX, or GPTQ.
MiniCPM5-2Bon-device LLMGGUF quantization

Qwen3.8-4B Distilled: Q4 vs Q6 vs Q8 Quantization Compared
Hands-on benchmark of Q4, Q6, and Q8 GGUF quants for the Qwen3.8-4B distilled model, testing speed, VRAM use, and reasoning depth on llama.cpp.
Qwen3.8-4BGGUF quantizationQ4 vs Q8

How to Run fuse-1 Lite Locally: VRAM, Setup, and Formats
How to run fuse-1 Lite's 5.72B coding MoE model locally via GGUF, MLX, vLLM, or bitsandbytes, with VRAM needs for each backend.
fuse-1 Lite VRAMrun fuse-1 Lite locallyGGUF quantization