Gemini 4 Argon Pricing: Is It Actually Cheaper Than Rival AI Models?
Gemini 4 Argon's launch pricing looks cheap next to Opus and GPT, but the discount expires and doubles rates. Here's the real cost picture.

How much does Gemini 4 Argon cost?
Google launched Gemini 4 Argon with introductory API pricing of $2 per million input tokens and $10 per million output tokens, plus a 95% discount on cached input. That rate is temporary. Google’s own footnote says the price doubles to $4 input and $20 output once the introductory period ends, though no end date has been published. Judge Argon’s pricing by the standard rate, not the launch number, if you’re planning a production budget.
TL;DR
- Launch pricing undercuts rivals at $2/$10 per million tokens, but Google has confirmed that rate doubles to $4/$20 after an unspecified introductory window.
- Artificial Analysis measured roughly $1.99 per intelligence-index task at the promotional rate, rising to about $3.98 once the discount lapses, compared to roughly $3.26 for Astra at max and about $0.72 for GPT 6.1 Solo at max on the same workload.
- The Vals index still ranks Argon as the cheapest per test even using its higher standard price, reporting about $15.68 per test versus $21.34 for Sonnet 5.5 and $32.14 for Opus 5.5.
- Argon generates far more output tokens per task (around 62,000 versus roughly 27,000 for Astra on one evaluation), which means the lower per-token price is partly offset by the model simply writing more.
- Coding costs don’t follow the same pattern. On a cybersecurity benchmark, Argon came in cheaper per rollout than Opus, but coding accuracy results were mixed across different software engineering tests.
- No single price-to-performance number tells the whole story. Cost comparisons shift depending on whether you weight finance and business tasks (where Argon does well) or coding tasks (where results are inconsistent).
Remy doesn't write the code. It manages the agents who do.
Remy runs the project. The specialists do the work. You work with the PM, not the implementers.
What is Gemini 4 Argon?
Gemini 4 Argon is Google’s newest frontier model, positioned against top-tier competitors like Anthropic’s Opus 5.5 and Sonnet 5.5, OpenAI’s GPT 6 series, and xAI’s Astra, rather than against lighter “Flash” or small-model tiers. Google released it alongside a wide benchmark sheet covering knowledge work, agentic coding, science and math reasoning, long-context handling, and computer use. Independent trackers including Artificial Analysis, the Vals index, Zapier’s automation bench, Arena, and Serge’s spreadsheet benchmark have published their own evaluations, giving a check on Google’s self-reported numbers.
Initial access rolled out through Google’s Fairwind program for select cybersecurity partners, with broader availability promised to paid API customers and Google AI Ultra subscribers. No firm public release date has been confirmed, so a subscription today doesn’t guarantee hands-on access yet.
Is Gemini 4 Argon actually cheap?
It depends which price you look at and which workload you’re pricing. At the introductory rate, Argon looks genuinely inexpensive relative to other frontier models. Artificial Analysis put its cost at about $1.99 per intelligence-index task, cheaper than Astra at max (around $3.26) on the same measure. But GPT 6.1 Solo at max reportedly costs about $0.72 per task while scoring only one index point below Argon, which makes Solo the more efficient option on that specific comparison even though Argon posts a higher raw intelligence score.
Once Argon’s standard pricing kicks in, doubling to $4 input and $20 output, Artificial Analysis estimated the per-task cost rises to roughly $3.98, putting it much closer to Astra’s price rather than clearly undercutting it. That’s an important detail for anyone budgeting around the launch numbers rather than the long-term rate.
The Vals index tells a more favorable story for Argon on pure dollar cost. Even using Argon’s higher standard token price (not the introductory discount), Vals calculated about $15.68 per test, against $21.34 for Sonnet 5.5 and $32.14 for Opus 5.5. That’s a real argument for value, but Vals weights finance, legal, tax, and coding benchmarks by their share of the US economy, and finance accounts for a little over half that weighting. If your actual workload skews toward coding rather than financial analysis, that aggregate number may not reflect what you’d pay for your specific tasks.
Why does Argon’s output volume matter for pricing?
One detail that complicates any simple price comparison: Argon tends to generate a lot more output per task than competing models. Artificial Analysis reported around 62,000 output tokens per task for Argon versus roughly 27,000 for Astra on the same evaluation. Google has also advertised a 1 million token output limit, meaning the model can keep working on a single response far longer than typical frontier models, using a technique Artificial Analysis called “long decode continuation” that resumes long responses through follow-up calls rather than timing out.
That changes how you should read any per-million-token price. A lower per-token rate multiplied against a much higher token count can land close to, or even above, a pricier model that answers more concisely. The practical question isn’t “what does Argon charge per million tokens,” it’s “what does it cost to get a finished, correct answer,” including all the extra tokens it spends thinking through the problem.
Does the price hold up for coding work?
This is where the pricing picture gets murkier. On a cybersecurity benchmark (CWE bench), Argon’s cost came in at roughly $63 per rollout compared to $79 for Opus, a real advantage. But that calculation used different agent harnesses for each model and a cached input price that doesn’t match Google’s introductory discount, so it shouldn’t be treated as a universal cost ratio across coding tasks generally.
More importantly, Argon’s coding accuracy itself is inconsistent across benchmarks. It beat Opus and Astra on one software engineering test (Deep SWE version 1.1) but lost to both on another (Frontier SWE version 2), and also trailed Opus on a terminal-based benchmark. The Vals index separately put Argon second behind Sonnet 5.5 on both app-building and code migration tests, though the margin on app building was smaller than the benchmark’s own reported margin of error. Paying less per token doesn’t help much if the model needs more retries to land a working patch. Anyone choosing a model primarily for coding work should weigh accuracy on the specific coding benchmarks that match their use case before leaning on the overall price comparison.
Is Gemini 4 Argon worth it on price alone?
Not by itself. The launch pricing is attractive and the Vals index backs up a genuine cost advantage for business-style workloads like finance, legal, and tax tasks. But three factors complicate any blanket claim that Argon is “the cheap frontier model”: the introductory rate is confirmed to double, the model tends to generate significantly more output tokens than rivals, and coding performance, where a lot of professional AI spending happens, is mixed rather than uniformly ahead of Opus or Sonnet.
The more useful framing is to price the actual job, not the token rate. That means factoring in output volume, retry rates, and which specific benchmark category matches your real workload, whether that’s financial analysis, automation workflows, or repository-level coding. A model that’s cheaper per million tokens but writes twice as much, or needs a second pass to fix its own mistakes, can end up costing about the same as a pricier model that gets it right the first time.
Frequently Asked Questions
What is Gemini 4 Argon’s standard API price?
Google’s introductory rate is $2 per million input tokens and $10 per million output tokens, with a 95% discount on cached input. The company has confirmed this doubles to $4 and $20 respectively once the introductory period ends, though no specific end date has been announced.
Is Gemini 4 Argon cheaper than Claude Opus 5.5?
On the Vals index, which measures cost per completed test using standard pricing, Argon came out cheaper (about $15.68 versus $32.14 for Opus 5.5). But Opus scored higher on Artificial Analysis’s overall intelligence index, so the comparison depends on whether you’re optimizing for cost or for raw capability on a given task.
Does Gemini 4 Argon’s low hallucination rate affect its value?
One coffee. One working app.
You bring the idea. Remy manages the project.
Argon scored a 15% hallucination rate on Artificial Analysis’s omniscience benchmark, the lowest among models scoring at least 45 on that intelligence index. That metric measures how often incorrect answers are confidently wrong versus honestly flagged as uncertain, not overall accuracy. Its knowledge accuracy (50%) was actually lower than Astra’s (63%), so the tradeoff is fewer confident wrong answers, not more correct ones.
Why does Gemini 4 Argon generate so much more output than other models?
Google gave Argon a 1 million token output limit, letting it continue reasoning through long, multi-step problems in a single response rather than cutting off. Artificial Analysis measured about 62,000 output tokens per task for Argon versus roughly 27,000 for Astra, which means part of its lower per-token price is offset by generating more tokens overall.
Should I switch my coding workflow to Gemini 4 Argon for the lower price?
Not based on pricing alone. Argon beat rivals on some coding benchmarks (Deep SWE version 1.1) but lost on others (Frontier SWE version 2, terminal bench), and the Vals index placed it second behind Sonnet 5.5 on app building and code migration. Test it on a difficult task your current model struggles with and compare total cost, including retries, before making a switch.