Qwen3.8-27B
Qwen3.8-27B Articles
Browse 5 articles about Qwen3.8-27B.

2-Bit vs FP8 Quantization: What the Escha-W2 Benchmarks Show
Escha-W2 compresses a 27B model to 2-bit and matches FP8 on GPQA, LiveCodeBench, and commonsense tests. Here's what that means for local inference.
2-bit vs FP8quantization benchmarkGPQA Diamond

What Is DFlash 2? Speculative Decoding Explained for Qwen3.8-27B
DFlash 2 is a block-diffusion draft model that speeds up Qwen3.8-27B inference up to 3.4x. Here's how it works and what it beats.
DFlash 2speculative decodingQwen3.8-27B

Qwen3.8-27B OBLITERATED: How This Uncensored Model Actually Works
A breakdown of Qwen3.8-27B-OBLITERATED V2, an abliterated model with a 0% refusal rate that matches or beats stock MMLU scores.
Qwen3.8-27B-OBLITERATEDabliterationuncensored LLM

DFlash 2: Run Qwen3.8-27B at 2x Speed with Speculative Decoding
DFlash 2 speeds up Qwen3.8-27B inference roughly 2x on a single A100 using speculative decoding in SGLang, with no output quality loss.
DFlash 2speculative decodingQwen3.8-27B

Qwen3.8-27B AEON Uncensored: How This Abliteration Actually Works
A community abliteration of Qwen3.8-27B explains its KL-drift methodology, judge-based refusal testing, and how to run the model via vLLM.
Qwen abliterateduncensored LLMAEON Qwen