Skip to main content
MindStudio
Pricing
BlogAbout
My Workspace
IQuest-Q1 benchmarksIQuest-Q1 SWE-Benchagentic coding model

IQuest-Q1 Benchmarks: What the New 320B Coding Model Actually Shows

IQuest-Q1 is a 320B MoE agentic coding model with a 524K context window. Here's what its benchmarks and limitations tell us so far.

Edited by Luis Chavez-Mattos, Director of Product RSS
IQuest-Q1 Benchmarks: What the New 320B Coding Model Actually Shows

What is IQuest-Q1?

IQuest-Q1 is an open-weight Mixture-of-Experts model built by IQuest for agentic coding, reasoning, and multi-step tool use. It has roughly 320 billion total parameters, but only about 15 billion are activated per token, which is what makes a model this large practical to run at all. It uses 88 transformer layers, 256 total experts with 8 activated per token, and a hybrid attention pattern that mixes sliding-window attention with full attention in a 3:1 ratio. Context length tops out at 524,288 tokens, which IQuest markets as a “1m”-class context setting in some client configurations.

The model is positioned squarely against other large reasoning and coding models, with its own model card comparing it to DeepSeek-V4-Flash and DeepSeek-V4-Pro. It ships with support for Claude Code and Codex CLI as agent harnesses, and IQuest provides Docker images for both SGLang and vLLM serving.

TL;DR

  • IQuest-Q1 is a 320B-parameter Mixture-of-Experts model with only 15B parameters activated per token, keeping inference costs closer to a mid-size dense model despite its total size.
  • The model supports a 524,288-token context window, among the longer context lengths available in an open-weight coding model right now.
  • IQuest reports results on agentic coding and reasoning benchmarks including SWE-Bench-style tasks, Terminal-Bench 2.1, CyberGym, and an in-house CLI benchmark called IQuest-CLIBench, evaluated through Claude Code and Codex harnesses rather than simple prompt-response testing.
  • Benchmark runs use generous time budgets, up to eight hours for Terminal-Bench 2.1 and six hours for CyberGym, which matters when comparing scores against models tested under tighter limits.
  • The model card is explicit that IQuest-Q1 is text-only and still early-stage, with known problems in long, iterative CLI debugging sessions.
  • Deployment is GPU-heavy, with reference configs assuming 8-way tensor parallelism, which rules out casual single-GPU use.
  • No independent benchmark listings for IQuest-Q1 currently appear on Hugging Face’s model search, so the only published numbers right now come directly from IQuest.

Remy doesn't write the code. It manages the agents who do.

R
Remy
Product Manager Agent
Leading
Design
Engineer
QA
Deploy

Remy runs the project. The specialists do the work. You work with the PM, not the implementers.

How was IQuest-Q1 benchmarked?

IQuest didn’t just run a one-shot coding eval. The model card describes a harness-based evaluation approach, meaning the model is tested inside real agent tooling rather than through isolated prompt completions. For most agentic coding benchmarks, IQuest used Claude Code as the harness. For DeepSWE v1.1 comparisons, they used mini-SWE-agent instead. For a benchmark called “Agents’ Last Exam,” they used a more recent Claude Code build (2.1.258).

This distinction matters. A model’s raw ability to write correct code is different from its ability to use tools, read file diffs, run tests, and iterate inside a coding agent loop. Benchmarking through Claude Code and Codex CLI specifically targets that second skill, agentic workflow competence, rather than just code generation in isolation.

Recommended sampling settings for reproducing IQuest’s numbers are temperature 1.0, top-p 0.95, and top-k 20, with Claude Code version 2.1.140 or Codex 0.142 as the harness. Anyone trying to replicate published scores needs to match these settings and harness versions closely, since agentic benchmarks are sensitive to tool-calling behavior and context handling that can shift between harness versions.

What benchmarks does IQuest-Q1 report?

The model card points to a performance chart covering multiple benchmark categories rather than a single leaderboard number. Based on the benchmark notes, the suite includes:

  • Agentic coding benchmarks evaluated via Claude Code or mini-SWE-agent (the kind of task family SWE-Bench and similar suites fall into)
  • Terminal-Bench 2.1, a test of command-line agent competence, run with an eight-hour time limit per task
  • CyberGym, run with a six-hour time limit
  • Humanity’s Last Exam, reported without tool access
  • IQuest-CLIBench, an in-house benchmark IQuest built specifically to measure CLI user experience

Comparisons are drawn against DeepSeek-V4-Flash and DeepSeek-V4-Pro (the 0731 and 0813 releases respectively). That’s a reasonable set of comparison points given those are also large, agent-capable models, but it also means the published numbers come from IQuest’s own testing pipeline, not a neutral third-party leaderboard. As of now, there’s no independent benchmark listing for IQuest-Q1 visible on Hugging Face’s public model search, so there’s no outside corroboration yet.

Why do the long time limits matter?

The eight-hour cap for Terminal-Bench 2.1 and six-hour cap for CyberGym are worth paying attention to because they’re unusually generous. Many published agentic benchmarks cap task time much more tightly, which forces models to solve problems efficiently. A long time budget lets a model retry, backtrack, and explore more before giving up, which can inflate success rates compared to a model tested under a stricter clock.

This isn’t necessarily a flaw in IQuest’s methodology, since real-world coding agents often do get long runtimes in practice. But anyone comparing IQuest-Q1’s scores against another model’s published Terminal-Bench numbers should check whether both were run under the same time constraints. Scores from different time budgets aren’t directly comparable.

What are IQuest-Q1’s known limitations?

IQuest is unusually direct about the model’s weaknesses in its own documentation. The listed limitations include:

  • No multimodal input. The model only accepts text. Image, audio, and video inputs aren’t supported in this checkpoint, and during evaluation, any multimodal content in agent conversations gets replaced with placeholders rather than actually processed.
  • Unreliable output. The card states plainly that generated code and explanations can be wrong, and recommends reviewing changes and verifying them with tests rather than trusting output directly.
  • Tool integration overhead. Tool calls depend on an IQuest-specific chat template format, which means serving the model correctly requires compatible parsers (iquest_q1 tool-call and reasoning parsers) and a properly configured agent harness. This isn’t a drop-in OpenAI-format model.
  • Struggles with real-world CLI iteration. The card specifically calls out that IQuest-Q1 “may overlook constraints, repeat failed attempts, or leave issues unresolved” during long, iterative debugging tasks, and recommends human oversight.
  • Early-stage status. IQuest describes the model as still early in development with “substantial limitations in its capabilities and reliability.”

That last point is notable for a benchmark-focused model release. Most vendors lead with strengths; IQuest’s documentation leads with a fairly blunt admission that this is a work in progress.

Is IQuest-Q1 worth running for agentic coding work?

For teams already running large open-weight models on multi-GPU infrastructure, IQuest-Q1 is worth evaluating directly against whatever is currently in production, especially if long-context agent workflows (500K+ tokens) are a priority. The MoE design with only 15B active parameters keeps per-token compute lower than the 320B total parameter count suggests, though serving still requires substantial tensor-parallel GPU setups (reference configs use 8-way tensor parallelism).

For smaller teams or anyone without multi-GPU infrastructure, the practical bar is higher. The model needs specific tool-call and reasoning parsers, specific harness versions (Claude Code 2.1.140, Codex 0.142) for benchmark-matching behavior, and careful prompt/context management given the known issues with long CLI debugging sessions. The model card’s own warnings about repeated failed attempts and unresolved issues in real CLI tasks suggest this isn’t a model to point at production codebases unsupervised yet.

Frequently Asked Questions

What benchmarks has IQuest-Q1 been tested on?

IQuest reports results on agentic coding tasks (evaluated through Claude Code and mini-SWE-agent), Terminal-Bench 2.1, CyberGym, Humanity’s Last Exam (without tools), and an in-house CLI benchmark called IQuest-CLIBench. Comparisons are made against DeepSeek-V4-Flash and DeepSeek-V4-Pro.

How many parameters does IQuest-Q1 have?

It has about 320 billion total parameters in a Mixture-of-Experts architecture, with roughly 15 billion activated per token across 8 of 256 total experts.

What context length does IQuest-Q1 support?

Up to 524,288 tokens. Some client configurations label this as a “1m” context setting, but the model card clarifies that this labeling doesn’t change the actual 512K context limit.

Can IQuest-Q1 handle images or other media?

No. The model card confirms it’s text-only, with no native image, audio, or video input support in this checkpoint.

What hardware is needed to run IQuest-Q1?

Reference deployment configs from IQuest use 8-way tensor parallelism with SGLang or vLLM, indicating this is a multi-GPU workload rather than something suited to a single consumer GPU.

Editorial standards

Presented by MindStudio

No spam. Unsubscribe anytime.