Local & Open-Weight Models
Deployment-focused content for open-weight models — running Gemma, Qwen, etc. locally, on phones, laptops, edge devices. Setup guides, hardware requirements, deployment patterns. Single-model reviews and explainers go under AI Model Reviews & Comparisons instead.

How to Fine-Tune Qwen3 27B Locally: LoRA, QLoRA and GGUF Guide
Learn how to fine-tune Qwen3 27B on a single GPU using Unsloth, LoRA/QLoRA, and export to GGUF, with dataset creation steps included.

IBM Granite 4.2 3B vs 8B: Local Reasoning Model Tested
Hands-on test of IBM's Granite 4.2 3B and 8B models locally, checking VRAM use, tool calling, and reasoning on real prompts.

How to Run IBM Granite 4.2 Locally with vLLM
Download and serve IBM's Granite 4.2 models locally with vLLM. VRAM needs and setup steps for the 3B, 8B and 30B reasoning variants.

How Much VRAM Do You Actually Need for Local AI in 2026?
How much VRAM local AI actually needs in 2026, from 24GB cards to 512GB Mac Studios, and why quantization changes the math on every build.

Mac Studio M5 Ultra vs DGX Spark: Which Wins for Local AI?
Comparing the upcoming Mac Studio M5 Ultra 512GB to Nvidia's DGX Spark on bandwidth, prefill speed, and price for local AI inference.

DeepSeek V4 Flash on One RTX 3090: Real Tokens-Per-Second Numbers
Real benchmark results for running DeepSeek V4 Flash and Qwen 3.8 27B locally on a single RTX 3090 or 4090 using FreeToken's desktop app.

How to Install FreeToken and Serve Qwen 3.6 Locally
A hands-on guide to installing FreeToken, benchmarking your GPU/CPU split, serving Qwen 3.6, and connecting coding agents to it locally.

FreeToken Explained: Run 290B+ MoE Models on One Gaming GPU
FreeToken streams only active experts to your GPU, letting massive MoE models like GLM and DeepSeek run locally without a multi-GPU server rig.

Escha-W2: 2-Bit Quantization That Shrinks a 27B Model to 10GB
Escha-W2 compresses Qwen3.8-27B into 10.15GB via 2-bit quantization, matching FP8 quality while fitting 128k context on one 24GB GPU.

Qwen3.8-27B OBLITERATED: How the V3 Abliterated Model Works
Qwen3.8-27B-OBLITERATED V3 removes refusals via complementary abliteration blending. Here's how it works, its MMLU cost, and GGUF options.

Qwen3.8-27B OBLITERATED: How V3 Abliteration Cuts Refusals, Not IQ
Qwen3.8-27B OBLITERATED removes hard refusals and safety-lecture deflections via V3 abliteration, losing just 2.1pp of MMLU score.

Run Qwen3.8-27B-Escha-W2 on a 24GB GPU with SGLang
How to install and tune Escha-W2, a 2-bit quant of Qwen3.8-27B, on a 24GB consumer GPU using SGLang for long context or high throughput.

How to Run Qwen3.8-27B-Escha-W2 Locally on a 24GB GPU
Guide to running Escha-W2, a 2-bit quantized Qwen3.8-27B, on a 24GB GPU with SGLang, covering VRAM tuning and 128k context setup.

Run Qwen3.8-27B-OBLITERATED Locally: GGUF Sizes, VRAM, and Settings
How to run the uncensored Qwen3.8-27B-OBLITERATED model locally: GGUF quant sizes, VRAM needs, and the exact settings that keep it from looping.

How to Run Qwen3.8-27B-OBLITERATED Locally with GGUF Quants
A practical guide to running the uncensored Qwen3.8-27B-OBLITERATED model locally: GGUF quant sizes, VRAM needs, and the exact settings it requires.

How to Run S1 Mini Locally for Clean Dictation Transcripts
S1 Mini cleans messy speech-to-text output locally on under 2GB VRAM. Here's how installation, styling modes, and VRAM usage actually work.

Superlinked Inference Engine (SIE): Run 100+ AI Models on One Server
Superlinked Inference Engine (SIE) is an open-source server that runs embedding, reranking, extraction, and generation models locally on one GPU.

2-Bit vs FP8 Quantization: What the Escha-W2 Benchmarks Show
Escha-W2 compresses a 27B model to 2-bit and matches FP8 on GPQA, LiveCodeBench, and commonsense tests. Here's what that means for local inference.

Qwen 3.8 27B: How to Run This Open Model Locally
Qwen 3.8 27B is an open-weight model that rivals closed frontier systems and runs on consumer GPUs. Here's how to install and use it.

What Is DFlash 2? Speculative Decoding Explained for Qwen3.8-27B
DFlash 2 is a block-diffusion draft model that speeds up Qwen3.8-27B inference up to 3.4x. Here's how it works and what it beats.

Qwen3.8-4B Distilled: How Emprius Shrank a 2.4T Model to 4B
How Emprius distilled Qwen3.8's 2.4 trillion parameter model into a 4B student using 45,000 reasoning traces, and what survived the compression.

Qwen3.8-4B Distilled: Q4 vs Q6 vs Q8 Quantization Compared
Hands-on benchmark of Q4, Q6, and Q8 GGUF quants for the Qwen3.8-4B distilled model, testing speed, VRAM use, and reasoning depth on llama.cpp.

Run DFlash 2 Speculative Decoding with vLLM and SGLang
How to set up DFlash 2, a lossless draft model for Qwen3.8-27B, using vLLM or SGLang for up to 3.4x faster inference speeds.

How to Run Ornith-1.5-9B Locally with vLLM or SGLang
A practical guide to serving Ornith-1.5-9B, a 9B reasoning model with 262K context, on a single GPU using vLLM or SGLang.