Skip to main content
MindStudio
Pricing
BlogAbout
My Workspace
Topic

Local & Open-Weight Models

Deployment-focused content for open-weight models — running Gemma, Qwen, etc. locally, on phones, laptops, edge devices. Setup guides, hardware requirements, deployment patterns. Single-model reviews and explainers go under AI Model Reviews & Comparisons instead.

How to Fine-Tune Qwen3 27B Locally: LoRA, QLoRA and GGUF Guide

Learn how to fine-tune Qwen3 27B on a single GPU using Unsloth, LoRA/QLoRA, and export to GGUF, with dataset creation steps included.

fine-tune Qwen3 27BUnsloth tutorialLoRA QLoRA

IBM Granite 4.2 3B vs 8B: Local Reasoning Model Tested

Hands-on test of IBM's Granite 4.2 3B and 8B models locally, checking VRAM use, tool calling, and reasoning on real prompts.

Granite 4.2IBM Granite localGranite 3B 8B benchmark

How to Run IBM Granite 4.2 Locally with vLLM

Download and serve IBM's Granite 4.2 models locally with vLLM. VRAM needs and setup steps for the 3B, 8B and 30B reasoning variants.

run Granite 4.2 locallyvLLM GraniteGranite 4.2 VRAM

How Much VRAM Do You Actually Need for Local AI in 2026?

How much VRAM local AI actually needs in 2026, from 24GB cards to 512GB Mac Studios, and why quantization changes the math on every build.

local AI VRAM requirementshow much VRAM for LLMlocal AI budget

Mac Studio M5 Ultra vs DGX Spark: Which Wins for Local AI?

Comparing the upcoming Mac Studio M5 Ultra 512GB to Nvidia's DGX Spark on bandwidth, prefill speed, and price for local AI inference.

Mac Studio M5 UltraDGX Spark comparisonlocal AI hardware

DeepSeek V4 Flash on One RTX 3090: Real Tokens-Per-Second Numbers

Real benchmark results for running DeepSeek V4 Flash and Qwen 3.8 27B locally on a single RTX 3090 or 4090 using FreeToken's desktop app.

DeepSeek V4 FlashFreeToken desktopRTX 3090 local AI

How to Install FreeToken and Serve Qwen 3.6 Locally

A hands-on guide to installing FreeToken, benchmarking your GPU/CPU split, serving Qwen 3.6, and connecting coding agents to it locally.

install FreeTokenFreeToken tutorialQwen 3.6 local

FreeToken Explained: Run 290B+ MoE Models on One Gaming GPU

FreeToken streams only active experts to your GPU, letting massive MoE models like GLM and DeepSeek run locally without a multi-GPU server rig.

FreeTokenrun large models locallymixture of experts offload

Escha-W2: 2-Bit Quantization That Shrinks a 27B Model to 10GB

Escha-W2 compresses Qwen3.8-27B into 10.15GB via 2-bit quantization, matching FP8 quality while fitting 128k context on one 24GB GPU.

Escha-W2 quantization2-bit LLM quantizationQwen3.8-27B GPU

Qwen3.8-27B OBLITERATED: How the V3 Abliterated Model Works

Qwen3.8-27B-OBLITERATED V3 removes refusals via complementary abliteration blending. Here's how it works, its MMLU cost, and GGUF options.

Qwen3.8-27B abliterateduncensored LLMOBLITERATED model

Qwen3.8-27B OBLITERATED: How V3 Abliteration Cuts Refusals, Not IQ

Qwen3.8-27B OBLITERATED removes hard refusals and safety-lecture deflections via V3 abliteration, losing just 2.1pp of MMLU score.

Qwen3.8-27B OBLITERATEDabliterationuncensored LLM

Run Qwen3.8-27B-Escha-W2 on a 24GB GPU with SGLang

How to install and tune Escha-W2, a 2-bit quant of Qwen3.8-27B, on a 24GB consumer GPU using SGLang for long context or high throughput.

Escha-W2 installSGLang serve.shRTX 3090 4090 5090 LLM

How to Run Qwen3.8-27B-Escha-W2 Locally on a 24GB GPU

Guide to running Escha-W2, a 2-bit quantized Qwen3.8-27B, on a 24GB GPU with SGLang, covering VRAM tuning and 128k context setup.

Escha-W2 local setuprun Qwen3.8-27B RTX 4090SGLang serve.sh

Run Qwen3.8-27B-OBLITERATED Locally: GGUF Sizes, VRAM, and Settings

How to run the uncensored Qwen3.8-27B-OBLITERATED model locally: GGUF quant sizes, VRAM needs, and the exact settings that keep it from looping.

run Qwen3.8-27B locallyGGUF quantization VRAMOllama uncensored model

How to Run Qwen3.8-27B-OBLITERATED Locally with GGUF Quants

A practical guide to running the uncensored Qwen3.8-27B-OBLITERATED model locally: GGUF quant sizes, VRAM needs, and the exact settings it requires.

Qwen3.8-27B-OBLITERATED GGUFrun uncensored LLM locallyllama.cpp Qwen setup

How to Run S1 Mini Locally for Clean Dictation Transcripts

S1 Mini cleans messy speech-to-text output locally on under 2GB VRAM. Here's how installation, styling modes, and VRAM usage actually work.

run S1 Mini locallyinstall S1 Minilocal speech to text

Superlinked Inference Engine (SIE): Run 100+ AI Models on One Server

Superlinked Inference Engine (SIE) is an open-source server that runs embedding, reranking, extraction, and generation models locally on one GPU.

Superlinked Inference EngineSIE open sourceself-hosted embedding server

2-Bit vs FP8 Quantization: What the Escha-W2 Benchmarks Show

Escha-W2 compresses a 27B model to 2-bit and matches FP8 on GPQA, LiveCodeBench, and commonsense tests. Here's what that means for local inference.

2-bit vs FP8quantization benchmarkGPQA Diamond

Qwen 3.8 27B: How to Run This Open Model Locally

Qwen 3.8 27B is an open-weight model that rivals closed frontier systems and runs on consumer GPUs. Here's how to install and use it.

Qwen 3.8 27Brun Qwen locallyopen-weight LLM

What Is DFlash 2? Speculative Decoding Explained for Qwen3.8-27B

DFlash 2 is a block-diffusion draft model that speeds up Qwen3.8-27B inference up to 3.4x. Here's how it works and what it beats.

DFlash 2speculative decodingQwen3.8-27B

Qwen3.8-4B Distilled: How Emprius Shrank a 2.4T Model to 4B

How Emprius distilled Qwen3.8's 2.4 trillion parameter model into a 4B student using 45,000 reasoning traces, and what survived the compression.

Qwen3.8 distillationreasoning tracessmall language model

Qwen3.8-4B Distilled: Q4 vs Q6 vs Q8 Quantization Compared

Hands-on benchmark of Q4, Q6, and Q8 GGUF quants for the Qwen3.8-4B distilled model, testing speed, VRAM use, and reasoning depth on llama.cpp.

Qwen3.8-4BGGUF quantizationQ4 vs Q8

Run DFlash 2 Speculative Decoding with vLLM and SGLang

How to set up DFlash 2, a lossless draft model for Qwen3.8-27B, using vLLM or SGLang for up to 3.4x faster inference speeds.

DFlash 2 setupvLLM speculative decodingSGLang draft model

How to Run Ornith-1.5-9B Locally with vLLM or SGLang

A practical guide to serving Ornith-1.5-9B, a 9B reasoning model with 262K context, on a single GPU using vLLM or SGLang.

Ornith-1.5-9B localrun on vLLMSGLang deployment