Local & Open-Weight Models
Deployment-focused content for open-weight models — running Gemma, Qwen, etc. locally, on phones, laptops, edge devices. Setup guides, hardware requirements, deployment patterns. Single-model reviews and explainers go under AI Model Reviews & Comparisons instead.

How to Run Qwen3.8-9B Distill Locally on a Single GPU
Guide to running Empero's Qwen3.8-9B distilled model locally: required kernels, sampling settings, and 262K context setup on one GPU.

Qwen3.8-27B at 2-Bit Quantization: Does Escha-W2 Actually Hold Up?
Asha Labs shrank Qwen3.8-27B to 2 bits per weight, cutting VRAM needs to 10GB. Here's how the Escha-W2 build performs in real tests.

Qwen3.8-27B OBLITERATED: How This Uncensored Model Actually Works
A breakdown of Qwen3.8-27B-OBLITERATED V2, an abliterated model with a 0% refusal rate that matches or beats stock MMLU scores.

How to Run Qwen3.8-27B AEON Uncensored Locally with vLLM
Setup guide for serving the abliterated Qwen3.8-27B AEON model on a single H200 GPU with vLLM, MTP speculative decoding, and tool calling.

How to Run Qwen3.8-27B with DFlash2 on vLLM or SGLang
Set up speculative decoding for Qwen3.8-27B with the DFlash2 draft model using vLLM or SGLang, with H200 benchmark numbers included.

MacBook M5 Max vs Razer Blade 18 RTX 5090: Which Wins for Devs?
M5 Max MacBook vs Razer Blade 18 RTX 5090 compared on Geekbench, SSD speed, and .NET compile times to see which laptop actually wins.

Ornith 1.5 35B-A3B: Local Deployment, VRAM, and Real-World Tests
Ornith 1.5 35B-A3B is a mixture-of-experts model with 3B active params. Here's how it runs locally on an A100, and how it handles agentic and reasoning tests.

Ornith 1.5 9B: Local Test Results Expose a Benchmark Gap
A hands-on local test of Ornith 1.5 9B on a single GPU reveals agentic and coding tasks breaking down despite strong benchmark claims.

What Is Qwak? Tether's Free Local AI Platform Explained
Qwak is Tether's open-source local AI platform bundling text, RAG, fine-tuning, and image/video generation into one offline install.

Qwen 3.8 27B Benchmarked: Agentic Index, Vision, and Reasoning Tests
Hands-on benchmarks of Qwen 3.8 27B covering its agentic index score, image counting, bounding box drawing, and reasoning effort output quality.

Qwen3.8-9B Distill by Empero: Benchmarks and Local Setup
Empero distilled Qwen3.8's reasoning into a 9B dense model. Here's what the MMLU and GSM8K benchmarks show, and how to run it locally.

Local LLMs in LM Studio: MacBook (MLX) vs Windows GPU Laptop (GGUF)
Compare running local LLMs in LM Studio on Apple Silicon with MLX versus a Windows GPU laptop with GGUF, and which setup fits your workflow.

How to Run Qwen 3 27B Locally with DeepSeek Harness
Set up Qwen 3 27B on your own hardware with DeepSeek Harness. Covers quantization, MLX vs VLM, reasoning effort, and vision capabilities.

DFlash 2: Run Qwen3.8-27B at 2x Speed with Speculative Decoding
DFlash 2 speeds up Qwen3.8-27B inference roughly 2x on a single A100 using speculative decoding in SGLang, with no output quality loss.

What Is Speculative Decoding? How Draft Models Speed Up LLMs
Speculative decoding explained: how a small draft model and a verification pass let large language models generate tokens faster without quality loss.

Antares-1B: Cisco's Tiny Model That Hunts Code Vulnerabilities Locally
Cisco's 1B-parameter Antares model localizes vulnerabilities in code repos and runs on a single GPU. Here's how it works and how to deploy it.

Qwen3.8-27B Ridge Quant: Run the Full Model on 12GB VRAM
Ridge quantization shrinks Qwen3.8-27B from 50GB to 11.7GB by protecting sensitive GatedDeltaNet layers. Here's how it works and performs.

What Is Unsloth? Local LLM Fine-Tuning and Inference Explained
Unsloth is a free, open-source tool for fine-tuning and running open LLMs locally with a point-and-click, ChatGPT-like agent UI.

What Computer Should You Buy for Local AI in 2026?
A practical framework for matching hardware to local AI models, based on VRAM, quantization, and real benchmark testing across Macs, PCs, and mini boxes.

Qwen3.8-27B AEON Uncensored: How This Abliteration Actually Works
A community abliteration of Qwen3.8-27B explains its KL-drift methodology, judge-based refusal testing, and how to run the model via vLLM.

Qwen3.8-9B: Running the Community-Distilled Model Locally
A team distilled Qwen 3.8's reasoning into a 9B model. Here's how it works, what VRAM it needs, and how it holds up in an agentic coding test.

How to Deploy dots3-note Preview Locally with vLLM or SGLang
A hardware and command guide to self-hosting dots3-note preview's FP8 checkpoint on 8-GPU nodes using vLLM or SGLang, with speculative decoding.

Meta Muse Glimmer: A 30B Open Model Built for Your GPU, Not the Frontier
Meta's Muse Glimmer is a 30-billion-parameter open-weight model sized for consumer GPUs like RTX 40/50 cards, not frontier benchmarks.

How to Run Qwen3.8-27B Locally with Ollama, LM Studio, and llama.cpp
Install Qwen3.8-27B GGUF quantizations locally with LM Studio, Ollama, and llama.cpp, plus VRAM comparisons and quantization guidance.