Gemini 4 Argon's Million-Token Output, Explained
Gemini 4 Argon can output a million tokens in one response. Here's why that ceiling matters for test-time compute and long agent runs.

What is Gemini 4 Argon’s million-token output limit?
Gemini 4 Argon is Google’s new frontier model, and its headline feature is a single-response output cap of up to one million tokens. Most current frontier models, including GPT and Claude variants, cap a single response around 128,000 tokens. Argon’s previous Gemini generation topped out even lower, around 64,000 tokens per response. By pushing the ceiling roughly eight times higher than competing models, Argon lets a model generate far more in one continuous run before anything needs to be cut off, summarized, or restarted.
TL;DR
- Gemini 4 Argon raises the single-response output ceiling to one million tokens, compared to the roughly 128,000-token cap common across other frontier models today.
- The bigger output window reduces the need for harness-level stitching, where a system has to stop a model mid-task, compress its progress, and restart it, a process that loses detail every time it happens.
- Longer output room also means longer uninterrupted chain-of-thought, since models with smaller caps have historically had to rush or truncate their reasoning before hitting the limit.
- Google is positioning Argon for enterprise knowledge work, long-horizon coding, and cybersecurity defense, citing internal use in migrating C and C++ code (including an OS kernel) to Rust.
- Independent benchmarking from Artificial Analysis puts Argon’s intelligence score near GPT-level frontier models and notably ahead of Google’s prior Pro model, while also showing a much lower hallucination rate in their testing.
- At launch, Argon is priced to match Claude Sonnet 5.5, moving to Opus 5.5 pricing after an introductory discount period, and it is currently limited to trusted testers rather than general availability.
- The real efficiency story isn’t just the million-token ceiling, it’s how many tokens the model actually needs to finish a task, since cost and latency scale with tokens used, not tokens available.
Plans first. Then code.
Remy writes the spec, manages the build, and ships the app.
Why does a bigger output window matter for test-time compute?
Test-time compute, also called test-time scaling, is the idea that a model can get smarter on a hard problem not by being retrained, but by being given more room to think during inference. Reasoning models already do this by generating long internal chains of thought before producing a final answer. The catch has been the output cap: if a model is limited to 128,000 tokens per response, its reasoning trace and its final answer have to fit inside that budget together.
Argon’s million-token ceiling changes the shape of that budget. A model can spend far more tokens reasoning through a genuinely hard problem, interleave tool calls or function calls within that reasoning, and still deliver a complete answer in a single pass. This matters because the alternative, hitting a cap partway through a task, forces a system to stop, summarize what’s been done, and restart. Each restart throws away detail that the model had in view, and summarization is lossy by nature. A larger ceiling means fewer seams, and fewer seams means less accumulated error over a long task.
There’s a well-known line of research on this tradeoff: Google DeepMind has published findings showing that a smaller model given more time and tokens to think at inference can outperform a model many times its size, at least on problems where the smaller model already has a reasonable shot at the answer. Argon’s output ceiling is effectively a bet that giving a single frontier model dramatically more room to “think out loud” in one shot pays off on exactly the kind of multi-step, long-horizon tasks that enterprises and coding agents care about.
How does this change long-running agent workflows?
Agent workflows, especially coding agents working on large codebases, are where output caps bite hardest. A task like rewriting a module, migrating a codebase to a new language, or running many rounds of iterative experimentation produces a lot of intermediate output: draft code, test results, revised code, more test results. If a model’s response window runs out in the middle of that, an agent harness has to compact the conversation and resume, often losing variable names, prior reasoning, or edge cases it had already worked out.
A million-token response window means an agent can carry a much larger amount of working context through to the end of a task without that kind of compaction. Google has pointed to internal examples at this scale: teams of Argon-based agents migrating C and C++ code, including an entire operating system kernel, to Rust, and an open-source video decoder that was rewritten in Rust to run meaningfully faster with identical output. These aren’t toy benchmarks, they’re examples run at the scale of software used by large numbers of people, which is the point Google is making: long-horizon software engineering is exactly the domain this output ceiling was built for.
Built like a system. Not vibe-coded.
Remy manages the project — every layer architected, not stitched together at the last second.
It’s worth noting a caveat here. For agents that are “tool-heavy,” meaning they make many separate calls out to external tools or APIs rather than reasoning in one continuous block, a bigger single-response ceiling matters less, because the task is already split across many calls regardless of output cap.
Is more output automatically better?
Not necessarily, and this is the part worth paying attention to. A million-token ceiling is a maximum, not a target. The number that actually determines cost and speed in production is how many tokens a model uses to finish a given task, not how many tokens it’s allowed to use. Some recent models have been criticized for burning through very large numbers of tokens to reach answers that a more efficient model could reach in far fewer. Independent benchmarking from Artificial Analysis frames this as a “cost per task” metric rather than a pure intelligence score, and it’s a more useful number for anyone actually paying for inference.
By that measure, Argon reportedly lands at a lower cost per task than some competing frontier models, though evaluation reports attribute most of that advantage to Argon’s per-token pricing rather than to it using meaningfully fewer tokens than rivals. In other words, Argon isn’t necessarily more frugal with tokens, it’s priced more favorably per token at launch. That’s a meaningful distinction for anyone budgeting inference costs: a bigger output ceiling gives a model room to be verbose when a task calls for it, but it doesn’t guarantee efficiency on simpler tasks.
Independent testing has also reported a notably lower hallucination rate for Argon compared to some rival frontier models, suggesting Google has tuned the model to say “I don’t know” rather than produce a confident wrong answer, something that matters more in production use than raw benchmark accuracy alone.
What are the practical tradeoffs of running a model this way?
Generating up to a million tokens in one response takes time. Decoding speed hasn’t been officially published, but if a model generates in the range of 100 tokens per second, a full million-token response could take on the order of hours, not seconds. That’s a real operational consideration: long single-shot runs need to be weighed against splitting work into shorter, parallelized calls, especially for latency-sensitive applications.
There’s also a cost dimension. Argon’s launch pricing has been positioned to match Claude Sonnet 5.5, with cached input priced at a steep discount, before moving up to Opus 5.5-level pricing once the introductory period ends. A longer response costs more in absolute terms even at a good per-token rate, so the million-token ceiling is best thought of as headroom for the hardest tasks, not a default mode for everyday queries.
Frequently Asked Questions
What is Gemini 4 Argon?
Argon is Google’s newest frontier model, positioned for enterprise knowledge work, long-horizon software engineering, and cybersecurity defense. It is currently being rolled out to trusted testers rather than the general public.
How big is Argon’s output limit compared to other models?
Argon can generate up to one million tokens in a single response. Most other current frontier models cap a single response around 128,000 tokens, and Google’s own prior-generation models were capped closer to 64,000 tokens.
Does a bigger output ceiling make a model smarter?
One coffee. One working app.
You bring the idea. Remy manages the project.
Not directly. It gives a model more room to reason and generate without being cut off and restarted, which can improve performance on long, complex tasks. But the actual intelligence and efficiency of a model still depend on how well it reasons and how many tokens it needs to use, not just how many tokens it’s allowed to use.
Is Gemini 4 Argon available now?
As of its announcement, Argon is in limited testing with select users and companies rather than generally available. Broader access is expected, though no firm date has been confirmed.
Why does token efficiency matter if the output cap is so high?
Because cost and latency in production scale with tokens actually used, not tokens available. A model that solves a task in fewer tokens is cheaper and faster to run, even if it has access to a much larger ceiling it rarely needs to use.