Gemini 4 Argon Benchmarks: How It Really Stacks Up Against Rivals
Independent benchmarks show Gemini 4 Argon leads on business tasks and hallucination control, but coding and pricing tell a mixed story.

What do independent benchmarks actually show about Gemini 4 Argon?
Gemini 4 Argon lands near the top of several third-party evaluations, but no single number tells the whole story. On artificial analysis’s intelligence index, Argon scores 53 at high reasoning, roughly tied with GPT 6.1 Astra at max reasoning and a point behind Claude Opus 5.5 (58) and Sonnet 5.5 (56). It does better on task-specific suites: it tops Val’s index at 68.9%, leads Zapier’s automation benchmark at 51.29%, and posts the lowest hallucination rate among high-scoring models at 15%. Coding results are split, with wins on some suites and losses on others.
TL;DR
- Argon’s overall intelligence index score of 53 on artificial analysis is statistically close to GPT 6.1 Astra and slightly behind Claude Opus 5.5 and Sonnet 5.5, so it’s competitive but not a clear leader on general reasoning.
- It takes first place on Val’s index (68.9% vs. roughly 67% for both Sonnet 5.5 and Opus 5.5), a suite weighted toward finance, legal, tax, and coding tasks based on their share of the US economy.
- On Zapier’s automation benchmark, which only counts a task as passed if every backend check succeeds, Argon scores 51.29% against Sonnet 5.5’s 44.75%, though roughly half of tasks still fail some check.
- Coding results are inconsistent across suites: Argon beats Opus and Astra on Deep SWE v1.1 (77.9%) but loses on Frontier SWE v2 and Google’s own terminal bench comparison, and Val’s index puts it second behind Sonnet 5.5 on app building and code migration.
- Argon’s hallucination rate of 15% on AA Omniscience is the lowest among models scoring 45+ on the intelligence index, but that figure only counts wrong answers among incorrect responses, not overall accuracy, which actually trails Astra (50% vs. 63% knowledge accuracy).
- Headline pricing of $2 per million input tokens and $10 per million output tokens is introductory and doubles to $4 and $20 once the promotional period ends, which changes the cost comparison against Astra and GPT 6.1 Soul significantly.
- Early access is limited to Google’s Fairwind program for select partners, with broader API and Google AI Ultra access promised but no public release date confirmed yet.
One coffee. One working app.
You bring the idea. Remy manages the project.
How does Argon compare to GPT-6 and Claude on overall intelligence?
On artificial analysis’s intelligence index, a composite reasoning measure, Argon scores 53 at high reasoning settings. GPT 6.1 Astra at max reasoning scores the same when rounded, putting the two in a dead heat on this particular metric. Claude’s Opus 5.5 scores 58 and Sonnet 5.5 scores 56, both measured at max settings with default fallback routing for refusals, meaning Anthropic holds a modest lead on this specific index.
The gap between Argon and the Claude models is small enough that a one or two point difference likely won’t translate into a noticeable difference on most individual tasks. These are index points, not completion percentages, so a score of 53 doesn’t mean the model finishes 53% of anything in particular. The index is useful for spotting broad trends across releases, not for predicting performance on a specific job.
Where does Argon actually win, and where does it lose?
The clearest wins show up in business-oriented and agentic evaluations rather than raw reasoning benchmarks.
Val’s index, which blends finance, coding, legal, and tax benchmarks weighted by their share of the US economy, puts Argon first at 68.9%, ahead of both Sonnet 5.5 and Opus 5.5 at roughly 67% each. Val also calculates this using Argon’s full standard (non-promotional) token price and still finds it cheaper per test than the Claude models. That’s a genuine value signal, though finance makes up more than half the index’s weighting, so the overall ranking may not reflect workloads dominated by code.
Zapier’s automation benchmark checks whether an agent actually completes multi-step business workflows across simulated apps, verifying the resulting records and changes rather than trusting the model’s own claim of success. Argon scores 51.29% at high reasoning versus 44.75% for Sonnet 5.5. That’s a real lead, but it also means close to half of tasks still fail some part of the check, which argues for supervision rather than full autonomy.
Spreadsheet work shows a similar pattern. On Serge’s GDPval.xlsx benchmark, Argon scores 38.3% against 30.3% for Opus 5.5 and 29.1% for Sonnet 5.5. This is a rubric-based score rather than a pass/fail completion rate, so it reflects partial credit across many criteria.
Coding is where the results pull in different directions. Google’s own comparison shows Argon ahead on Deep SWE v1.1 (77.9% vs. 74.2% for Opus and 74.1% for Astra) but behind on Frontier SWE v2 (55% vs. 65.5% for Astra and 62.3% for Opus) and Google’s terminal bench comparison (57.4% vs. 66.4% for Opus). Val’s independent testing places Argon second on both app building and code migration benchmarks, behind Sonnet 5.5, though the app-building gap is smaller than the benchmark’s own margin of error. There’s also a methodology wrinkle: Google computed Argon’s Deep SWE score using its own agent harness, while the Astra and Claude numbers came from separate leaderboards and system cards, so it’s not a fully controlled comparison.
What’s going on with the hallucination numbers?
- ✕a coding agent
- ✕no-code
- ✕vibe coding
- ✕a faster Cursor
The one that tells the coding agents what to build.
Argon’s 15% hallucination rate on artificial analysis’s AA Omniscience benchmark is the lowest among models scoring at least 45 on the intelligence index, well below Astra’s 51% at max reasoning. That sounds like a clean win, but the metric is narrower than it appears. It only measures, among the responses that weren’t fully correct, how often the model gave a confidently wrong answer instead of a partial answer or an honest “I don’t know.”
Correct answers aren’t counted in that denominator at all, which is why Argon’s overall knowledge accuracy on the same benchmark (50%) actually trails Astra’s (63%). In other words, Argon answers fewer questions correctly than Astra, but when it doesn’t know something, it’s more likely to say so rather than invent an answer. For research and fact-heavy work, that’s a meaningful behavioral difference: a model that admits uncertainty wastes less of your time than one that confidently fabricates a citation or API. Whether that caution carries over into coding and browsing tasks, where tools are available to verify claims, is still an open question.
Is the Argon pricing actually a good deal?
It depends heavily on timing. Google’s advertised rate of $2 per million input tokens and $10 per million output tokens (with a 95% discount on cached input) is explicitly introductory. A footnote confirms the standard rate doubles to $4 and $20 once the promotional period ends, with no fixed end date announced.
At the promotional price, artificial analysis calculates about $1.99 per intelligence index task for Argon, cheaper than Astra’s $3.26. At the standard rate, that cost nearly doubles to $3.98, while GPT 6.1 Soul at max comes in at just $0.72 for a score only one index point lower. Val’s index, using the standard (non-promotional) price, already shows Argon costing less per test than both Sonnet 5.5 and Opus 5.5 ($15.68 vs. $21.34 and $32.14 respectively), which suggests the value case holds up even without the launch discount, at least for that particular workload. Argon also tends to generate far more output tokens per task (around 62,000) than Astra (around 27,000), so its lower per-token price is partly offset by generating more text. The practical takeaway: calculate cost using the standard rate and your own task mix rather than trusting the launch pricing as a permanent baseline.
Is Gemini 4 Argon worth switching to?
For business and document-heavy workflows, finance analysis, spreadsheet work, automation agents, and tasks where admitting uncertainty matters, the independent results make a credible case. Multiple separate benchmarks (Val’s index, Zapier, Serge’s spreadsheet test) point the same direction.
For coding, the picture is mixed enough that wholesale migration isn’t justified by the launch numbers alone. Argon wins some software engineering benchmarks and loses others, sometimes within Google’s own comparison table. The sensible approach is testing it against a real migration task or a problem your current model struggled with, then comparing both the output quality and the total cost at standard (non-promotional) pricing.
Frequently Asked Questions
Is Gemini 4 Argon better than Claude Opus 5.5?
It depends on the task. Opus 5.5 scores higher on artificial analysis’s general intelligence index (58 vs. 53) and leads on Val’s app-building and code migration benchmarks. Argon leads on Val’s overall index, Zapier’s automation benchmark, and hallucination avoidance.
How does Gemini 4 Argon compare to GPT-6?
Argon and GPT 6.1 Astra score almost identically on artificial analysis’s intelligence index (both around 53). Argon shows a lower hallucination rate but also lower raw knowledge accuracy than Astra, and the two trade wins across different coding benchmarks.
Does Gemini 4 Argon actually reduce hallucinations?
It reduces confidently wrong answers among its incorrect responses (15% vs. Astra’s 51%), but its overall knowledge accuracy (50%) is lower than Astra’s (63%). It answers fewer questions correctly but is more likely to admit uncertainty instead of fabricating an answer.
Is Gemini 4 Argon pricing cheaper than competitors?
At launch, yes, roughly $2 per million input tokens and $10 per million output tokens, undercutting Astra on cost-per-task. But that rate doubles after an unspecified introductory period, and GPT 6.1 Soul remains far cheaper per task even at Argon’s promotional price.
Can I access Gemini 4 Argon right now?
Initial access is limited to Google’s Fairwind program for selected partners. Google says broader access will come through paid API customers and Google AI Ultra subscribers, but no public release date has been confirmed.
