Skip to main content
MindStudio
Pricing
BlogAbout
My Workspace
Gemini 3.8 Flash pricingGemini Flash costAI token pricing

Gemini 3.8 Flash Pricing: How Much Does It Cost to Use?

Gemini 3.8 Flash costs 75 cents per million input tokens and $3.75 per million output tokens as an introductory rate that expires this year.

Edited by Luis Chavez-Mattos, Director of Product RSS
Gemini 3.8 Flash Pricing: How Much Does It Cost to Use?

How much does Gemini 3.8 Flash cost per million tokens?

Gemini 3.8 Flash currently costs 75 cents per million input tokens and $3.75 per million output tokens. That’s an introductory rate, and Google has flagged in its own pricing documentation that these numbers change at the end of the year, with a higher rate listed in smaller print underneath the headline price. Even after that increase, the model still lands well below what Anthropic and OpenAI charge for comparable-tier models, which is the whole point of Google’s flash lineup: get strong agentic and reasoning performance at a price point built for production-scale usage, not just demos.

TL;DR

  • Introductory pricing for Gemini 3.8 Flash is 75 cents per million input tokens and $3.75 per million output tokens, but Google has stated this is temporary and set to rise around December 31 or January 1.
  • Claude Opus 5 costs roughly $5.25 by comparison, meaning Flash’s current rate is a fraction of frontier-tier pricing, even after the scheduled increase.
  • GPT-5.6 Sol runs around $4.20 and GPT-5.6 Terra around $2.12, putting Gemini 3.8 Flash well under even OpenAI’s mid-tier offering.
  • Cost per task matters more than sticker price, since one creator’s testing showed Flash-family models can use up to 30% more output tokens per task than the previous generation, partially offsetting the savings.
  • On the DeepSWE 1.1 benchmark, Gemini 3.8 Flash scores competitively with Opus 5 (around 73.7% versus roughly even numbers for Opus), while GPT-5.6 Sol trails slightly behind at 72.7%, making the cost-to-performance ratio the real story.
  • A specialized Cyber variant, available only to trusted partners through a limited access program, reportedly matches Sonnet-level cyber benchmark performance at under $4 versus roughly $10 for the comparison model.
  • The model isn’t a frontier replacement everywhere: it lags notably behind Opus 5 on harder agentic coding benchmarks like Terminal Bench, so pricing needs to be weighed against task type, not treated as a blanket discount.
Cursor
ChatGPT
Figma
Linear
GitHub
Vercel
Supabase
goremy.ai

Seven tools to build an app. Or just Remy.

Editor, preview, AI agents, deploy — all in one tab. Nothing to install.

Why is Google pricing Flash this aggressively?

Google has released three flash-tier models in about six weeks, an unusually fast cadence that signals a deliberate push to own the cost-efficient end of the market rather than chase frontier benchmark headlines with every release. Analysts covering the release point to the “Pareto frontier” framing used by outside benchmark trackers: Gemini 3.8 Flash sits in the sweet spot where you get a strong chunk of frontier-model capability without paying frontier-model prices.

This matters because most real-world AI usage isn’t a single high-stakes prompt to a top-tier model. It’s thousands or millions of repeated calls in production pipelines, coding agents, or customer-facing tools, where the difference between $0.75 and $5.25 per million input tokens compounds fast. Google appears to be betting that developers building at scale care more about consistent, cheap, “good enough” performance than about winning every individual benchmark.

How does Gemini 3.8 Flash pricing compare to Opus 5 and GPT-5.6?

The gap is large. Reported figures put Claude Opus 5 at around $5.25, GPT-5.6 Sol at around $4.20, and GPT-5.6 Terra at around $2.12, all measured on a comparable blended basis. Against that field, Gemini 3.8 Flash’s 75 cent input and $3.75 output pricing is dramatically cheaper, even when compared to Terra, which is itself positioned as OpenAI’s more budget-friendly tier.

The catch is the fine print: those Flash prices are introductory and Google has said they’ll rise before the new year. Some coverage has noted the actual post-introductory price is displayed in small text beneath the bolded promotional number, which is worth a second look if you’re budgeting a project around current rates. Even at the higher rate, Flash is still expected to undercut Terra, so the pricing advantage doesn’t disappear, it just shrinks somewhat.

Does cheaper pricing actually mean cheaper tasks?

Not automatically. Per-token price is only half the equation. What actually determines your bill is cost per completed task, which depends on how many tokens a model burns solving a given problem. One analysis using outside benchmark tracking noted that Gemini 3.8 Flash, compared to the prior flash generation, showed as much as a 30% increase in output tokens per task. A cheaper model that needs more tokens to reach the same answer can end up costing about the same as a pricier model that’s more efficient.

This is why benchmark charts that plot accuracy against average cost per task (rather than just listed per-token price) are more useful for real budgeting. On the DeepSWE 1.1 benchmark specifically, Gemini 3.8 Flash was reported sitting favorably on this cost-versus-performance curve, scoring competitively with Opus 5 while costing a fraction as much, which is the strongest evidence so far that the pricing advantage holds up even after accounting for token usage.

Is Gemini 3.8 Flash worth it for coding and agentic tasks?

Other agents ship a demo. Remy ships an app.

UI
React + Tailwind ✓ LIVE
API
REST · typed contracts ✓ LIVE
DATABASE
real SQL, not mocked ✓ LIVE
AUTH
roles · sessions · tokens ✓ LIVE
DEPLOY
git-backed, live URL ✓ LIVE

Real backend. Real database. Real auth. Real plumbing. Remy has it all.

It depends heavily on the specific task. On DeepSWE 1.1, a long-horizon software engineering benchmark, Gemini 3.8 Flash performed close to Opus 5, which is a notable result for a model priced this far below frontier tiers. But on Terminal Bench, a newer and harder agentic coding benchmark, Opus 5 was reported to score more than twice as high as Flash on one version of the test. On another real-world knowledge work benchmark (GDPval, covering tasks like PDF extraction, data analysis, and presentation building), Gemini 3.8 Flash came in noticeably behind both Opus 5 and GPT-5.6 Sol.

The upshot from creators testing the model directly: it’s not a frontier model replacement, and nobody covering the release claimed it was. It’s positioned as a workhorse for well-scoped, high-volume tasks, agentic coding loops, legal document review (it reportedly leads on the Harvey legal benchmark), financial analysis workflows, and everyday multimodal work, where paying frontier prices for every call doesn’t make economic sense. For open-ended creative or design-heavy work, testers found other models (including GPT-5.6 Sol) still ahead on raw output quality.

What is Gemini 3.8 Flash Cyber and how is it priced?

Alongside the general release, Google introduced Gemini 3.8 Flash Cyber, a variant tuned for cybersecurity tasks like vulnerability discovery, and made available only to a limited set of trusted partners through what’s described as a “Fair Wind” style access program rather than general availability. On the CyberGym benchmark, it reportedly performs at a level close to a more expensive comparison model, at a price under $4 versus about $10 for that competitor, according to Google’s own materials. Because it isn’t publicly available, its pricing is more a signal of Google’s broader cost strategy than something most developers can act on directly today.

Frequently Asked Questions

What is the current price of Gemini 3.8 Flash?

As of release, it’s priced at 75 cents per million input tokens and $3.75 per million output tokens. This is described as an introductory rate.

When does Gemini 3.8 Flash pricing change?

Coverage points to a change around December 31 or January 1, after which the rate is expected to rise, though Google has not been fully transparent about how much, listing the future price in smaller text below the current promotional rate.

Is Gemini 3.8 Flash cheaper than GPT-5.6?

Yes, on a per-million-token basis it’s cheaper than both GPT-5.6 Sol and GPT-5.6 Terra, based on reported pricing figures of roughly $4.20 and $2.12 respectively for those OpenAI models.

How does Gemini 3.8 Flash compare to Claude Opus 5 on price and performance?

It’s a fraction of the cost of Opus 5, which runs around $5.25, and on some benchmarks like DeepSWE 1.1 it scores close to Opus 5. On others, like Terminal Bench, Opus 5 performs substantially better, so the value depends on the task.

Does a lower price always mean lower total cost?

No. Some testing found the newer flash model uses more output tokens per task than its predecessor, up to roughly 30% more in some cases, which can offset part of the per-token savings depending on the workload.

Editorial standards

Presented by MindStudio

No spam. Unsubscribe anytime.