Skip to main content
MindStudio
Pricing
BlogAbout
My Workspace

IQuest-Q1: A 320B MoE Model Built for Agentic Coding

IQuest-Q1 is a 320B MoE model with 15B active params and 512K context, designed for agentic coding with Claude Code and Codex CLI.

Edited by Luis Chavez-Mattos, Director of Product RSS
IQuest-Q1: A 320B MoE Model Built for Agentic Coding

What is IQuest-Q1?

IQuest-Q1 is an open-weight Mixture-of-Experts (MoE) language model from IQuest, built specifically for agentic coding, reasoning, and multi-step tool use. It has about 320 billion total parameters, but only roughly 15 billion are activated per token thanks to its MoE routing (256 total experts, 8 activated per forward pass). The model supports a context window of 524,288 tokens (512K) and ships with first-class integration paths for Claude Code and Codex CLI, two of the most widely used agentic coding harnesses. It was published on Hugging Face by IQuestLab and has already picked up meaningful community attention, with over 700 downloads and 128 likes shortly after release.

TL;DR

  • IQuest-Q1 is a 320B-parameter MoE model that activates only about 15B parameters per token, keeping inference costs closer to a mid-size dense model while retaining a large total capacity.
  • The model uses a hybrid attention pattern (three sliding-window attention layers for every one full-attention layer) plus partial RoPE and multi-token prediction (MTP) layers to speed up inference.
  • It supports a 512K token context window, with Claude Code and Codex CLI both configured to use the full window rather than a truncated default.
  • IQuest-Q1 is positioned as a direct benchmark competitor to DeepSeek-V4-Flash and DeepSeek-V4-Pro on agentic coding and reasoning tasks, including an in-house CLI benchmark called IQuest-CLIBench.
  • Deployment is handled through SGLang or vLLM, with prebuilt Docker images and speculative decoding support via EAGLE-based multi-token prediction.
  • The model card is explicit that IQuest-Q1 is text-only and early-stage, prone to the same kinds of repeated failed attempts and unresolved edge cases seen in other agentic coding models.

How is IQuest-Q1 architected?

IQuest-Q1 is a dense-sparse hybrid in the modern MoE tradition: 88 transformer layers, a hidden dimension of 3,072, and an attention setup with 48 query heads and 8 key/value heads (grouped-query attention) at a head dimension of 128. The routing layer picks 8 experts out of a pool of 256 for each token, which is how the model keeps only about 15B parameters “active” despite a 320B total footprint.

Two architectural details stand out. First, the attention pattern alternates three sliding-window attention (SWA) layers for every one full attention (FA) layer, with a 4,096-token sliding window. This is a common trick for controlling the compute and memory cost of long-context attention: most layers only look at a local window, while periodic full-attention layers let information propagate across the entire sequence. Combined with partial RoPE (applied to just 32 dimensions of the head), this lets the model stretch to its 512K context without the quadratic cost blowing up everywhere.

Second, IQuest-Q1 uses multi-token prediction (MTP) for inference acceleration. During training, it uses two independent MTP layers; at inference, it switches to a recursive scheme (one MTP module applied eight times) with a 512-token sliding window. This is paired with speculative decoding support in both SGLang and vLLM, using an EAGLE-style draft model shipped alongside the main checkpoint. For teams running the model at scale, this combination of MoE sparsity, hybrid attention, and MTP speculative decoding is where most of the real throughput gains will come from, more than from any single headline spec.

How does IQuest-Q1 compare to DeepSeek-V4-Flash and DeepSeek-V4-Pro?

IQuest positions Q1 directly against DeepSeek’s V4-Flash and V4-Pro releases (the 0731 and 0813 builds, per the model card), benchmarking across agentic coding and reasoning tasks. The evaluation suite includes CyberGym and Terminal-Bench 2.1 for sandboxed task execution, Humanity’s Last Exam (reported without tool use) for general reasoning, Agents’ Last Exam for agentic reasoning, and an in-house benchmark called IQuest-CLIBench, which specifically measures the CLI user experience rather than just raw task success.

The methodology details matter here. For agentic coding tasks, IQuest’s team used mini-SWE-agent to evaluate DeepSWE v1.1, and Claude Code as the harness for the other benchmarks (Claude Code 2.1.258 specifically for Agents’ Last Exam). Runtime limits were capped at six hours for CyberGym and eight hours for Terminal-Bench 2.1, which signals these are long-horizon, multi-step tasks rather than single-shot prompts. Since the model has no multimodal input capability, any multimodal content in agent conversations was replaced with placeholders during tokenization, a caveat worth remembering if you’re evaluating the benchmark numbers against models that do handle images.

The practical takeaway: IQuest-Q1 is being benchmarked less like a general chatbot and more like a coding agent runtime component, one meant to be judged by how well it survives long, tool-heavy sessions rather than how well it answers isolated questions.

How do you deploy and run IQuest-Q1?

Other agents ship a demo. Remy ships an app.

UI
React + Tailwind ✓ LIVE
API
REST · typed contracts ✓ LIVE
DATABASE
real SQL, not mocked ✓ LIVE
AUTH
roles · sessions · tokens ✓ LIVE
DEPLOY
git-backed, live URL ✓ LIVE

Real backend. Real database. Real auth. Real plumbing. Remy has it all.

IQuest recommends two serving backends: SGLang and vLLM, both distributed as prebuilt Docker images tagged for CUDA 13.0. A typical SGLang launch uses 8-way tensor parallelism, bfloat16 precision, and the FlashAttention-3 backend, with flags for a custom tool-call parser and reasoning parser specific to IQuest-Q1’s chat template. Recursive MTP can be enabled with the EAGLE speculative algorithm, pointing at a dedicated MTP draft model folder that ships alongside the main weights. vLLM follows the same pattern, exposing a speculative-config JSON block for the same EAGLE-based acceleration plus prefix caching.

For agent integration, the model card gives working configurations for both major coding harnesses. For Claude Code (version 2.1.140 recommended), you set environment variables to route all four Claude model tiers (Sonnet, Opus, Haiku, and the default) to IQuest-Q1[1m], along with context and auto-compaction settings tuned to the 512K window. For Codex CLI (version 0.142.0 recommended), the setup involves writing a config.toml and model_catalog.json into the Codex config directory, registering IQuest as a custom model provider over the OpenAI Responses API shape. Both configurations disable approval prompts and run with full sandbox access by default, which is standard for benchmarking harnesses but worth tightening before any real-world use.

Basic single-turn inference, outside of an agent harness, works through any OpenAI-compatible client once a server is running, using the standard chat.completions.create call pattern with temperature 1.0 and top-p 0.95 as recommended defaults.

Is IQuest-Q1 worth using right now?

For teams already running Claude Code or Codex CLI against hosted frontier models, IQuest-Q1 is interesting mainly as a self-hostable alternative aimed at the same workflows, long agent sessions, large context windows, and heavy tool use, rather than as a general-purpose chatbot replacement. The 15B active parameter count makes it plausible to run at reasonable throughput on multi-GPU setups without needing the full 320B worth of active compute, and the MTP/EAGLE speculative decoding support suggests IQuest has put real engineering effort into inference speed, not just raw benchmark chasing.

That said, the model card is unusually candid about its limitations. It explicitly warns that real-world CLI tasks often involve iterative debugging, and that IQuest-Q1 “may overlook constraints, repeat failed attempts, or leave issues unresolved, necessitating human oversight.” It also flags itself as “early stage” with “substantial limitations in its capabilities and reliability.” That’s a level of honesty worth taking at face value: this is a first release competing against mature, iterated models like DeepSeek-V4, and the benchmark wins (where they exist) should be read in that context. Text-only input, reliance on compatible serving parsers for tool calls, and the general unpredictability of long-horizon agent tasks all mean this is a model for teams willing to do their own evaluation, not a drop-in production swap.

Frequently Asked Questions

What does “15B active parameters” mean for IQuest-Q1?

It means that although the full model has about 320 billion parameters stored across 256 experts, the MoE routing mechanism only activates 8 experts (roughly 15 billion parameters worth of compute) for any given token. This keeps inference cost closer to a much smaller dense model while preserving the larger model’s total knowledge capacity.

Does IQuest-Q1 support images or other modalities?

No. The model card explicitly states it is text-only, with no native image, audio, or video input capability. Any multimodal content encountered during evaluation was replaced with placeholder tokens.

What context length does IQuest-Q1 support?

Other agents start typing. Remy starts asking.

YOU SAID "Build me a sales CRM."
01 DESIGN Should it feel like Linear, or Salesforce?
02 UX How do reps move deals — drag, or dropdown?
03 ARCH Single team, or multi-org with permissions?

Scoping, trade-offs, edge cases — the real work. Before a line of code.

It supports up to 524,288 tokens (512K), and both the Claude Code and Codex CLI integration guides configure the full window rather than truncating it, including specific auto-compaction thresholds tuned to that limit.

How is IQuest-Q1 served in production?

Through SGLang or vLLM, using prebuilt Docker images. Both backends support tensor parallelism across 8 GPUs, bfloat16 precision, and optional EAGLE-based speculative decoding using IQuest-Q1’s multi-token prediction (MTP) module for faster generation.

How does IQuest-Q1 compare to DeepSeek-V4?

IQuest benchmarks Q1 against DeepSeek-V4-Flash and DeepSeek-V4-Pro across agentic coding and reasoning benchmarks, including CyberGym, Terminal-Bench 2.1, Humanity’s Last Exam, Agents’ Last Exam, and an in-house CLI benchmark. Exact comparative scores come from IQuest’s own published benchmark charts, and independent verification is still limited given the model’s recent release.

Editorial standards

Presented by MindStudio

No spam. Unsubscribe anytime.