IQuest-Q1: Inside the 320B MoE Model Built for Agentic Coding
IQuest-Q1 is a 320B parameter MoE model with 15B active params and 512K context, designed for agentic coding and tool use.

What is IQuest-Q1?
IQuest-Q1 is an open-weight Mixture-of-Experts (MoE) language model built by IQuest for agentic coding, reasoning, and multi-step tool use. It has roughly 320 billion total parameters, but only about 15 billion are activated per token, thanks to a sparse MoE routing scheme with 256 total experts and 8 activated per forward pass. The model supports a context window of 524,288 tokens (512K) and ships with integration paths for both Claude Code and Codex CLI, positioning it as a drop-in alternative for developers who already work inside agentic coding harnesses.
TL;DR
- IQuest-Q1 is a 320B parameter MoE model with only 15B parameters active per token, making inference cheaper than its total size suggests.
- The model uses a hybrid attention pattern (3 sliding-window attention layers to 1 full attention layer) across 88 transformer layers, a design choice aimed at balancing long-context efficiency with global reasoning.
- Native context length is 512K tokens, achieved partly through partial RoPE (32 dimensions) and a 4,096-token sliding window inside the hybrid attention blocks.
- It ships with multi-token prediction (MTP) support, using 2 independent MTP layers during training and a recursive 8x MTP setup at inference for faster decoding via speculative sampling.
- IQuest positions the model against DeepSeek-V4-Flash and DeepSeek-V4-Pro on agentic coding and reasoning benchmarks, including an in-house CLI benchmark called IQuest-CLIBench.
- Deployment is supported through SGLang and vLLM with prebuilt Docker images, and the model integrates directly into Claude Code (2.1.140) and Codex CLI (0.142.0) via custom tool-call and reasoning parsers.
- The model card is explicit about limitations: text-only input, unreliable output on real-world CLI tasks, and an early-stage status that requires human oversight.
Everyone else built a construction worker.
We built the contractor.
One file at a time.
UI, API, database, deploy.
How is IQuest-Q1 architected?
IQuest-Q1 stacks 88 transformer layers with a hidden dimension of 3,072, using a grouped attention setup with 48 query heads and 8 key/value heads, each with a head dimension of 128. The standout design decision is its hybrid attention pattern: three sliding-window attention (SWA) layers for every one full attention (FA) layer. The sliding window is capped at 4,096 tokens, which keeps most layers computationally cheap while periodic full-attention layers let the model still reason across the entire 512K context when needed.
This is paired with partial rotary position embeddings, applied to only 32 dimensions of each head rather than the full head dimension. Partial RoPE is a common trick for extending context length without a full retraining pass on positional encodings, and it likely contributes to how IQuest-Q1 reaches its 512K window without the memory blowup that comes from applying dense attention everywhere.
On the MoE side, the model routes each token to 8 of 256 available experts. That ratio (roughly 3% of experts active) is what keeps the active parameter count down to about 15B despite the 320B total. This is the same basic trade-off other large MoE models make: you pay the memory cost of storing all the experts, but the compute cost of inference tracks closer to the active parameter count.
What is multi-token prediction doing here?
IQuest-Q1 includes a multi-token prediction (MTP) module, a technique that trains the model to predict several future tokens at once rather than just the next one. During training, IQuest uses 2 independent MTP layers. At inference, this becomes a “recursive x8” setup with a 512-token MTP sliding window.
Practically, MTP is used to enable speculative decoding: the draft model predicts multiple tokens ahead, and the main model verifies them in batches, which speeds up generation. The deployment configs for both SGLang and vLLm show this explicitly, using EAGLE-style speculative decoding with a separate MTP draft model path and rejection sampling. For agentic coding workloads, where the model may generate long tool-call sequences or multi-file diffs, faster token throughput has a direct effect on how usable the model feels in an interactive coding session.
How does IQuest-Q1 compare to DeepSeek-V4?
IQuest’s own benchmark charts position the model against DeepSeek-V4-Flash and DeepSeek-V4-Pro (referred to by their 0731 and 0813 releases), along with other coding-oriented models. The evaluation covers agentic coding benchmarks, Humanity’s Last Exam (reported without tool use), and an in-house benchmark called IQuest-CLIBench that specifically measures command-line interface user experience.
Other agents ship a demo. Remy ships an app.
Real backend. Real database. Real auth. Real plumbing. Remy has it all.
The benchmarking methodology is worth noting because it affects how comparable these numbers really are. IQuest says it uses mini-SWE-agent as the harness for DeepSWE v1.1, and Claude Code for other agentic tasks, including Claude Code 2.1.258 specifically for a benchmark called Agents’ Last Exam. Runtime limits were set at six hours for CyberGym and eight hours for Terminal-Bench 2.1. Since IQuest-Q1 currently has no multimodal capability, any multimodal content in agent conversations gets replaced with placeholders during tokenization, which is a meaningful caveat for anyone trying to reproduce results on multimodal-heavy agent benchmarks.
As with any vendor-published benchmark suite, these numbers should be read as a starting point rather than a final verdict. The choice of harness (mini-SWE-agent versus Claude Code) can itself shift results, and IQuest acknowledges this by specifying exact harness versions (Claude Code 2.1.140 or Codex 0.142) for reproducibility.
How do you deploy and run IQuest-Q1?
IQuest-Q1 is distributed as safetensors weights (split across 175 files) on Hugging Face, with a separate MTP module folder for speculative decoding. IQuest recommends two serving backends: SGLang and vLLM, both with prebuilt Docker images tagged for CUDA 13.0.
For SGLang, the launch command uses 8-way tensor parallelism, bfloat16 precision, and the FlashAttention-3 backend, with flags for a custom tool-call parser and reasoning parser (iquest_q1). There’s a variant command that adds recursive MTP via EAGLE speculative decoding, pointing at the model’s mtp subdirectory as the draft model path.
The vLLM path is similar: 8-way tensor parallel serving, auto tool choice enabled, and an optional speculative config block that wires up the EAGLE method with the same MTP draft model and 5 speculative tokens per step.
Both setups assume multi-GPU infrastructure. Given the 320B total parameter count in bfloat16, even with sparse activation, full deployment requires substantial GPU memory across the 8-way tensor parallel setup, which is standard for a model at this scale.
Does it work with Claude Code and Codex CLI?
Yes. IQuest-Q1’s model card includes explicit setup instructions for both. For Claude Code (version 2.1.140 recommended), you point the standard Anthropic environment variables (ANTHROPIC_MODEL, ANTHROPIC_BASE_URL, ANTHROPIC_AUTH_TOKEN) at a gateway serving IQuest-Q1, using a [1m] suffix that signals the million-token client setting without changing the model’s actual 512K limit. Output token limits, autocompact thresholds, and timeout settings are all tuned for the model’s context size.
For Codex CLI (version 0.142.0 recommended), the setup writes a custom config.toml and model_catalog.json pointing at an OpenAI-compatible Responses API endpoint, with context window and auto-compact token limits matched to 512K. Notably, the sample config sets approval_policy = "never" and sandbox_mode = "danger-full-access", which disables the built-in safety guardrails in Codex CLI. Anyone adapting this configuration should treat those settings as something to reconsider for their own environment, not a default to copy blindly.
Both integrations require a serving gateway that understands IQuest’s tool-call and reasoning parser format, since the model uses its own chat template for structured tool calls rather than a generic format.
Is IQuest-Q1 worth using right now?
IQuest is upfront that this is an early-stage release. The model card lists text-only input as a hard limitation (no image, audio, or video), and warns that generated code and explanations “can be incorrect,” recommending that users verify outputs with task-appropriate tests. It also flags that real-world CLI tasks involving iterative debugging often trip the model up: it may overlook constraints, repeat failed approaches, or leave problems unresolved, requiring human oversight throughout.
Built like a system. Not vibe-coded.
Remy manages the project — every layer architected, not stitched together at the last second.
For teams evaluating it, the practical draw is the combination of large total capacity (320B parameters worth of expert knowledge) with a much lighter compute footprint per token (15B active), plus native long-context support and out-of-the-box agentic harness integration. The tradeoff is that, per IQuest’s own admission, it’s not yet a mature, fully reliable coding agent, and benchmark comparisons against DeepSeek-V4 come from IQuest’s own evaluation setup rather than a neutral third party.
Frequently Asked Questions
How many parameters does IQuest-Q1 actually use per token?
It has about 320 billion total parameters but only activates roughly 15 billion per token, since it routes each token through 8 of 256 available experts in its MoE architecture.
What context length does IQuest-Q1 support?
It supports up to 524,288 tokens (512K) natively, enabled in part by partial RoPE embeddings and a hybrid sliding-window/full-attention pattern across its 88 transformer layers.
Can IQuest-Q1 handle images or other non-text input?
No. The current checkpoint is text-only. It has no native support for image, audio, or video inputs, which the model card lists as an explicit limitation.
Which agentic coding tools does IQuest-Q1 integrate with?
It has documented integration paths for Claude Code (version 2.1.140 recommended) and Codex CLI (version 0.142.0 recommended), both requiring a gateway that supports IQuest’s custom tool-call and reasoning parser.
How does IQuest-Q1 compare to DeepSeek-V4?
IQuest benchmarks its model against DeepSeek-V4-Flash and DeepSeek-V4-Pro on agentic coding and reasoning tasks, including an in-house CLI benchmark. These are vendor-reported comparisons using specific harness versions, so results should be treated as a starting point rather than an independent verdict.




