GLM-5.3 Uncensored EXL3 Quant: What It Takes to Run It Locally
GLM-5.3 Uncensored in EXL3 3.0bpw needs ~273GB and multi-GPU serving. Here's the hardware, TabbyAPI config, and a tool-calling bug to fix first.

What is the GLM-5.3 Uncensored EXL3 release?
It’s a quantized version of a weight-edited GLM-5.3, built for people who want to self-host the model without the vendor’s safety tuning and without paying for 753 billion parameters at full precision. The chain runs: Z.ai’s GLM-5.3 base model, a “crack surgery” weight edit by dealignai that strips alignment behavior without fine-tuning, then an EXL3 3.0bpw quantization by a third party (Infatoshi on Hugging Face) that shrinks it to about 273 GiB. It’s meant to run on local or rented GPU clusters using the ExLlamaV3 backend and TabbyAPI for serving.
TL;DR
- The release is a three-step pipeline: Z.ai’s GLM-5.3 base, a no-fine-tune weight edit from dealignai that removes alignment, and a 3.0bpw EXL3 quant on top.
- At 273 GiB on disk, this needs serious multi-GPU hardware. It was tested on 8x A100 40GB using TabbyAPI’s automatic GPU split.
- Quality loss from quantization looks modest but real: perplexity rises from 3.302 (FP8) to 3.440 (EXL3), and KL divergence stays low on high-confidence tokens.
- On agentic benchmarks (tau2-bench airline and retail), the quant trails the FP8 baseline by small margins that aren’t statistically significant at the tested sample sizes.
- There’s a known tool-calling bug in TabbyAPI’s GLM4.5 parser that mangles string arguments like order IDs and zip codes into integers, and it must be patched before using this model for agents.
- The model ships with an MTP (multi-token prediction) layer that TabbyAPI can use as a speculative decoding draft model, which speeds up generation without a separate draft model.
- Licensing is MIT, inherited from both the upstream GLM-5.3 and the dealignai release.
Other agents start typing. Remy starts asking.
Scoping, trade-offs, edge cases — the real work. Before a line of code.
How was the model built?
The base is zai-org/GLM-5.3, a mixture-of-experts model with 753 billion total parameters, 256 routed experts (8 active per token) plus one shared expert, using multi-head latent attention (MLA) with a sparse DSA indexer across 78 layers plus one multi-token prediction (MTP) layer.
dealignai took that FP8 checkpoint and applied a documented weight edit, not a fine-tune, to produce GLM-5.3-UNCENSORED-FP8. The specifics of the edit live in a file called CRACK_SURGERY.json, which the EXL3 release copies unchanged from the source repo. This is a distinct technique from RLHF-reversal fine-tuning: it edits weights directly rather than training on new data, which is faster and cheaper but means the behavioral changes are whatever that specific surgical edit targeted.
The EXL3 quantization step then compresses that FP8 model for practical local serving. It was converted with ExLlamaV3 (commit d3739fd) using the tool’s default calibration set (250 rows of 2048 tokens each), reading directly from the FP8 checkpoint. The quant isn’t uniform: attention and shared experts sit at 5 bpw, dense MLP layers at 4 bpw, routed experts (the bulk of the parameters) at 3 bpw, and the output head (lm_head) at 6 bpw. The blend works out to an average of 3.04 bits per weight. The MTP layer itself is quantized separately and uncalibrated, with experts at 4 bpw and attention/shared expert at 6 bpw.
What hardware do you need to run it?
Expect to need enough combined VRAM to hold roughly 273 GiB of weights, plus headroom for KV cache and context. The model card’s own test used 8x A100 40GB GPUs (320GB total VRAM) with TabbyAPI’s gpu_split_auto option, which automatically distributes layers across available GPUs rather than requiring manual tensor-parallel configuration.
That’s not a consumer setup. Realistically this targets people with access to a multi-GPU workstation, a rented GPU cluster, or an enterprise inference box. There’s no official CPU-offload or single-GPU path documented for this release, and at 753B total parameters even an aggressive 3-bit quant lands well outside what a single high-end card (even an H100 80GB) can hold on its own.
How good is the quantization compared to the original?
The repo includes a fidelity comparison against the FP8 source, run over 20 rows of 2048 tokens from the wikitext-2 test set using a KL-divergence diff tool. Perplexity goes from 3.302 (FP8) to 3.440 (EXL3 quant), a small but non-trivial increase. KL divergence between the quant and FP8 distributions sits at 0.089 to 0.097 depending on direction, with a median per-token KL of 0.021. More telling: on the 44% of tokens where the FP8 model was already highly confident (top probability at least 0.95), median KL drops to 0.0011, meaning the quant barely disturbs the model’s most confident predictions. Most of the divergence concentrates in harder, more uncertain tokens, which is the expected and generally acceptable failure mode for aggressive quantization.
Does the uncensoring and quantization hurt agentic performance?
Remy doesn't write the code. It manages the agents who do.
Remy runs the project. The specialists do the work. You work with the PM, not the implementers.
There’s some signal that it might, but it’s not conclusive. The model card reports tau2-bench results (airline and retail domains, Pass^1 scoring) comparing this quant against stock GLM-5.3 FP8 served by Z.ai through OpenRouter, both run through the same evaluation harness with GPT-4.1 as the user simulator and judge.
On airline tasks (50 tasks, 2 trials each), the stock model scored 0.710 versus 0.640 for this quant, a gap of about 1.1 standard errors. On retail (114 tasks), stock scored 0.504 versus 0.482, a gap of about 0.4 standard errors. Neither difference clears the bar for statistical significance at these sample sizes. It’s worth being precise about what this comparison actually measures: it mixes two separate effects (the dealign weight edit and the EXL3 quantization) against a baseline that has neither, so you can’t cleanly attribute any gap to compression alone. The honest read is a possible small regression on airline-style tasks, not a confirmed one.
How do you serve it with TabbyAPI?
The tested configuration uses exllamav3 as the backend with GLM’s tool format and reasoning mode enabled:
model:
model_name: GLM-5.3-UNCENSORED-EXL3-3.0bpw
backend: exllamav3
max_seq_len: 65536
cache_size: 98304
gpu_split_auto: true
tool_format: glm4_7
reasoning: true
draft_model:
draft_mode: mtp
The draft_mode: mtp setting tells TabbyAPI to use the model’s own built-in MTP layer as a speculative decoding draft, instead of requiring a separate smaller model. This is a meaningful operational convenience: speculative decoding normally needs you to source and load a compatible draft model, but here the draft capability ships inside the same checkpoint. The tested setup used a 98K-token shared cache against a 65K max sequence length.
What’s the tool-calling bug, and how do you fix it?
This is the detail that matters most if you plan to use this model for agents or function calling. GLM writes tool call arguments as raw text inside tags, like <arg_value>9523456873</arg_value>. TabbyAPI’s glm4_5 parser (as tested, at commit be74bf0) JSON-decodes every one of those values without checking the tool’s schema first. The result: any string parameter that happens to look like a number, order IDs, product IDs, zip codes, gets silently converted to an integer before it reaches your tool.
On the tau2-bench retail evaluation, this bug alone caused roughly 70% of tool calls to fail. The fix is to patch the parser so it checks each parameter’s declared schema type and keeps anything typed as string as raw text rather than coercing it. The benchmark numbers reported above were run with that fix already applied, so anyone deploying this model for tool-calling workloads should apply the same patch before trusting agentic results, and probably before running any production workload that touches order or account IDs.
Frequently Asked Questions
What does “uncensored” mean for this GLM-5.3 release?
It refers to a documented weight edit (not a fine-tune) applied by dealignai to the base GLM-5.3 FP8 checkpoint, intended to remove alignment-driven refusals. The edit’s specifics are recorded in a file called CRACK_SURGERY.json included in the repo.
How much VRAM do I need to run GLM-5.3 Uncensored EXL3?
The quantized model is about 273 GiB on disk. The documented test configuration used 8x A100 40GB GPUs (320GB total) with TabbyAPI’s automatic GPU split. Smaller multi-GPU setups might work with reduced context length, but no single-GPU or consumer-hardware path is documented.
Other agents ship a demo. Remy ships an app.
Real backend. Real database. Real auth. Real plumbing. Remy has it all.
Is the EXL3 quant as good as the original FP8 model?
Close, but not identical. Perplexity rises modestly (3.302 to 3.440) and KL divergence stays very low on tokens the original model was already confident about. Agentic benchmark scores (tau2-bench) show small, statistically inconclusive drops versus the FP8 baseline.
Can I use this model for tool calling or agents right now?
Only after patching TabbyAPI’s tool-call parser. The default glm4_5 parser incorrectly converts string arguments that look numeric (order IDs, zip codes) into integers, which broke about 70% of tool calls in retail-domain testing. Fix the parser to respect the schema’s declared parameter types first.
What license does this model use?
MIT, inherited from both the upstream GLM-5.3 release and the dealignai weight-edited version it’s built on.