Local & Open-Weight Models
Deployment-focused content for open-weight models — running Gemma, Qwen, etc. locally, on phones, laptops, edge devices. Setup guides, hardware requirements, deployment patterns. Single-model reviews and explainers go under AI Model Reviews & Comparisons instead.

How to Run dots3-note Locally on vLLM or SGLang
A hardware and setup guide to deploying dots3-note-prev-fp8 locally with vLLM or SGLang, covering FP8 quantization and 8-GPU serving.

How to Run DeepSeek V4 Pro Locally with vLLM or SGLang
A practical guide to deploying DeepSeek V4 Pro locally with vLLM or SGLang, including DSpark speculative decoding and hardware configs.

How to Run MiniMax Music 3 Locally for AI Song Generation
A practical guide to installing MiniMax Music 3 locally, covering VRAM needs, the Gradio setup, and what the open music model actually delivers.

North MicroVision 2.4B: Installing Cohere's Local OCR Vision Model
Cohere's North MicroVision is a 2.4B open vision model for OCR and documents. Here's how to install it, its real VRAM use, and where it falls short.

M5 MacBook Air: How Much Faster Is Local AI, Really?
The M5 MacBook Air brings a real memory bandwidth jump for local LLMs. Here's what that means for prompt processing and token generation.

Can You Run Qwen3.8-2.4T-A95B Locally? Hardware Requirements Explained
What it actually takes to self-host Qwen3.8-2.4T-A95B, a 2.4-trillion-parameter MoE model, including VRAM, quantization, and vLLM setup.

What Is NVIDIA SwitchYard? The Open-Source Local AI Model Router
NVIDIA SwitchYard routes agent tasks between local and frontier models automatically. Here's what it does, how routing works, and why it matters.

How to Run fuse-1 Lite Locally: VRAM, Setup, and Formats
How to run fuse-1 Lite's 5.72B coding MoE model locally via GGUF, MLX, vLLM, or bitsandbytes, with VRAM needs for each backend.

How to Run Maple-Preview Locally on a Mac Mini M4
A guide to running DeepGrove's Maple-Preview 20B-A1B ternary model on a Mac mini M4, covering hardware needs and real-world speed.

How to Run Nemotron 3.5 Lightning Locally on Your Own GPU
A practical guide to running and fine-tuning NVIDIA's Nemotron 3.5 Lightning MoE model locally, covering hardware needs, NVFP4, and Unsloth.

Set Up a Local AI Router With SwitchYard and Nemotron Lightning
How to configure NVIDIA's SwitchYard router with a locally hosted Nemotron 3.5 Lightning model to cut API costs without losing task quality.

How to Run Prime Agent Locally With DeepSeek V4 on Your Own Hardware
A hands-on guide to installing Prime Agent, configuring it for a local DeepSeek V4 endpoint, and comparing its performance to Claude Code and Codex.

Run MiniMax H3 Locally: VRAM Guide From 6GB Cards to the 5090
How to run the open-source MiniMax H3 AI video model locally, with VRAM tiers from a 6GB RTX 2060 up to a 5090, plus Mac options.

DeepSeek V4 Flash: The Cheapest Frontier-Level Open Model Yet
DeepSeek V4 Flash jumps to 54% on agentic coding benchmarks and rivals larger models, but harness choice explains much of the gain.

How to Run DeepSeek V4 Flash Locally: Hardware, Quantization, Tests
A practical guide to running DeepSeek V4 Flash locally, covering VRAM needs, quantization tradeoffs, DGX Spark setups, and real agentic coding results.

Self-Hosting Chinese Open-Weight AI Models: A Practical Decision Guide
A practical guide to self-hosting DeepSeek, Qwen, or GLM versus using APIs, covering hardware costs, licensing terms, and data governance risks.

Poolside's Laguna S 2.1: A 118B Open Model for Local Agentic Coding
Poolside's open-weight Laguna S 2.1 runs agentic coding locally at 80+ tokens/sec on DGX Spark, matching models ten times its size.

Local AI Video Generation with ComfyUI: How It Actually Works
How a 128GB local workstation runs ComfyUI with Qwen image and LTX video models to generate unlimited AI content without API fees.

How to Run a 744B AI Model on a Consumer Laptop Using Colibri
Colibri uses three-tier memory and SSD streaming to run GLM 5.2 on consumer hardware. Learn how the hot-cold expert split makes this possible.

AMD's Ryzen AI Developer Center: Local AI Setup Without the ROCm Grind
AMD's Ryzen AI Developer Center ships with playbooks for ComfyUI, LM Studio, and Unsloth fine-tuning, cutting out manual ROCm and driver setup.

AMD Ryzen AI Max Plus 395: 128GB Unified Memory for Local LLMs
AMD's Ryzen AI Max Plus 395 pairs 128GB unified memory with a Radeon 8060S GPU, letting local rigs load 100B+ parameter models without a discrete GPU.

Build a Free Overnight AI Video Pipeline Locally with ComfyUI
Skip per-second cloud video API costs. Learn how a local ComfyUI pipeline with Qwen Image and LTX Video batch-generates clips overnight for free.

Local LLM Speed Test: GPT-OSS, Qwen3.6 and Hermes on 128GB Unified Memory
Real token-per-second benchmarks for GPT-OSS 120B, Qwen3.6 MoE, and Hermes agents running locally on 128GB unified memory hardware.

How to Use AI for Secure Document Processing: Local Models, PII Detection, and Compliance
Learn how to use LM Studio and open-weight models to scan contracts for PII, mask credentials, and process sensitive files without sending data to the cloud.