Skip to main content
MindStudio
Pricing
BlogAbout
My Workspace
Gemini 3.7 Flash benchmarksGemini 3.7 Flash vs 3.6Gemini Flash coding benchmarks

Gemini 3.7 Flash Benchmarks: How Much Better Is It Than 3.6?

Gemini 3.7 Flash beats 3.6 Flash by wide margins on coding and agentic benchmarks just three weeks after launch. Here's the full breakdown.

Edited by Luis Chavez-Mattos, Director of Product RSS
Gemini 3.7 Flash Benchmarks: How Much Better Is It Than 3.6?

What are the actual benchmark differences between Gemini 3.7 Flash and 3.6 Flash?

Gemini 3.7 Flash posts double-digit gains over Gemini 3.6 Flash across every major coding and agentic benchmark, despite arriving just three weeks after its predecessor. On DeepSweep, a long-horizon software engineering benchmark, it scores 65.3% versus 49% for 3.6 Flash. On Frontier Code, a production-quality coding benchmark, it hits 43.6% against 34.4%. Web Dev Arena Elo climbs from roughly 1538 to 1588. On Automation Bench, which measures agentic task completion, the score nearly doubles from 17% to about 30%.

TL;DR

  • DeepSweep scores jumped 16 points, from 49% on Gemini 3.6 Flash to 65.3% on Gemini 3.7 Flash, a large move for a long-horizon software engineering benchmark in a single point release.
  • Frontier Code improved by roughly 9 points, going from 34.4% to 43.6%, suggesting the model produces more production-ready code rather than just passing simple test cases.
  • Web Dev Arena Elo rose from about 1538 to 1588, a smaller but still meaningful gain on a benchmark that ranks models by head-to-head web development quality.
  • Automation Bench nearly doubled, climbing from 17% to around 30%, which is the clearest signal that agentic tool-calling and multi-step planning improved the most.
  • The model ships with a 1 million token context window and reportedly runs at around 250 tokens per second, keeping Flash-tier speed while closing ground on frontier-level coding performance.
  • Pricing dropped sharply too, with an introductory API rate of $0.75 per million input tokens and $3.75 per million output tokens, roughly half of what 3.6 Flash cost at launch, plus a temporary Open Router discount pushing it even lower.
  • Access is free or nearly free right now through Google Antigravity and AI Studio, and heavily discounted on Open Router, which makes the benchmark jump more relevant since builders can actually test it without cost pressure.

Remy is new. The platform isn't.

Remy
Product Manager Agent
THE PLATFORM
200+ models 1,000+ integrations Managed DB Auth Payments Deploy
BUILT BY MINDSTUDIO
Shipping agent infrastructure since 2021

Remy is the latest expression of years of platform work. Not a hastily wrapped LLM.

Why did Gemini 3.7 Flash improve so much in just three weeks?

Google released Gemini 3.7 Flash on August 13th, about three weeks after Gemini 3.6 Flash. That kind of turnaround usually signals a minor patch: bug fixes, small instruction-following tweaks, nothing structural. The benchmark data doesn’t match that pattern. A 16-point gain on DeepSweep and a near-doubling on Automation Bench point to real changes in how the model plans, calls tools, and handles multi-step tasks rather than a superficial tune-up.

Google positioned the release as its “most intelligent workhorse model for coding and agents,” which lines up with where the biggest gains show up: agentic benchmarks moved more than general knowledge or reasoning benchmarks. That’s consistent with a targeted training push toward tool use and long-horizon task completion rather than a broad capability upgrade.

How does Gemini 3.7 Flash perform in real coding and agentic tasks?

Benchmarks are one thing, but the practical test is whether a Flash-tier model can actually build something without breaking halfway through. In hands-on use inside Google’s Antigravity coding environment, Gemini 3.7 Flash has handled full app builds in a single pass, producing clean layouts and working functionality without requiring constant correction. Follow-up edits, like changing a design element or adding a feature, don’t tend to destabilize the rest of the project, which was a common complaint with earlier Flash-class models.

The agentic behavior is where the benchmark story translates most directly into usability. The model plans multi-step tasks, calls tools in the right order, and tracks long chains of actions without losing state. Combined with a reported speed of around 250 tokens per second and a 1 million token context window, it can hold large codebases or long conversations in memory while still responding quickly.

How does the pricing compare to the previous Flash model?

Google is running an introductory API price for Gemini 3.7 Flash through the end of December: $0.75 per million input tokens and $3.75 per million output tokens. That’s about half of what Gemini 3.6 Flash cost at launch, which already made Flash models attractive for high-volume use.

On top of that, Open Router is running an additional limited-time discount through August 27th, bringing the effective price down to roughly $0.38 per million input tokens and $1.88 per million output tokens. That works out to about 75% off the standard rate for a model beating its predecessor by wide margins on coding and agentic benchmarks. For anyone running agentic pipelines through tools like Cline, Roo Code, or Open Code, that pricing changes the cost calculus for running the model continuously rather than sparingly.

Where can you use Gemini 3.7 Flash for free right now?

There are three main access points, and all of them currently cost nothing or close to it:

  • Google Antigravity: Gemini 3.7 Flash is available for free inside Google’s Antigravity agentic coding environment. Log in, select the model, and start building.
  • AI Studio: The model is free to use for chatting and prompt testing, and it’s also available in AI Studio Build for generating and deploying small apps.
  • Open Router: API access at the discounted rate mentioned above, valid until August 27th, after which pricing reverts to the standard introductory rate.
Cursor
ChatGPT
Figma
Linear
GitHub
Vercel
Supabase
goremy.ai

Seven tools to build an app. Or just Remy.

Editor, preview, AI agents, deploy — all in one tab. Nothing to install.

Free tiers on Antigravity and AI Studio are meant for testing and experimentation. Google can adjust quotas at any time, so relying on the free tier for production workloads carries risk. For personal projects, prototyping, and evaluation, the free access is more than sufficient to judge whether the model fits a given workflow.

Is Gemini 3.7 Flash worth using over a larger frontier model?

For the hardest reasoning or coding problems, a frontier-class Pro model will likely still outperform a Flash model on raw capability. Gemini 3.7 Flash isn’t positioned to replace that tier. What it changes is the economics of everyday work: day-to-day coding tasks, agentic workflows, and general-purpose queries where a fast, cheap model that gets most of the way to frontier quality is more useful than a slower, expensive model that’s marginally better.

The benchmark jumps matter most in this context. A model that nearly doubles its agentic task completion rate while running at Flash speed and Flash pricing removes the tradeoff that used to force a choice between speed and reliability. That’s a bigger practical shift for people building automation pipelines or coding assistants than another few points on a frontier leaderboard.

Frequently Asked Questions

What is Gemini 3.7 Flash?

It’s Google’s latest Flash-tier model, released August 13th as the successor to Gemini 3.6 Flash. It’s built for coding and agentic workloads, with a 1 million token context window and fast inference speeds.

How much better is Gemini 3.7 Flash than 3.6 Flash on benchmarks?

It scores 65.3% versus 49% on DeepSweep, 43.6% versus 34.4% on Frontier Code, about 1588 versus 1538 Elo on Web Dev Arena, and roughly 30% versus 17% on Automation Bench.

Is Gemini 3.7 Flash free to use?

Yes, it’s currently free in Google Antigravity and in AI Studio (including AI Studio Build). API access through providers like Open Router carries a cost, though a temporary discount was in effect through August 27th.

How fast is Gemini 3.7 Flash?

It has been reported to run at around 250 tokens per second, which combined with its 1 million token context window makes it suitable for large, fast-moving coding and agentic tasks.

Does Gemini 3.7 Flash replace frontier Pro models?

No. It’s not designed to beat top-tier frontier models on the hardest tasks. Its value is in delivering close to frontier-level performance on everyday coding and agentic work at a fraction of the cost and speed penalty.

Editorial standards

Presented by MindStudio

No spam. Unsubscribe anytime.