Gemini 3.8 Flash Tested: Cheap, Fast, and Harness-Dependent
Gemini 3.8 Flash hits Opus-5 scores on Deep SWE at a fraction of the cost, but real output quality swings hard by harness.

What is Gemini 3.8 Flash and why does it matter?
Gemini 3.8 Flash is Google’s newest lightweight Gemini model, released just three weeks after Gemini 3.7 Flash and the third flash-tier release in about six weeks. It matters because it lands near the top of the cost-to-performance curve on real coding benchmarks while staying priced like a budget model, and because two independent hands-on tests show its actual output quality depends heavily on which coding harness you run it in.
TL;DR
- Deep SWE parity puts Gemini 3.8 Flash at 73.7% on Deep SWE v1.1, edging out GPT 5.6 Soul (72.7%) and landing effectively even with Claude Opus 5, despite being a flash-tier model.
- Pricing stays aggressive, at 75 cents per million input tokens and $3.75 per million output tokens as an introductory rate through the end of the year, well under Opus 5 ($5-25), GPT 5.6 Soul ($4-20), and GPT 5.6 Terra ($2-12).
- Harness choice changes the output dramatically: the same detailed prompt produces a noticeably weaker result in the plain Gemini app than in Google’s Antigravity harness, which adds multi-turn implementation and verification steps.
- Benchmark results are uneven, with Gemini 3.8 Flash winning Humanity’s Last Exam (55.9) and the Harvey legal benchmark (61.4%, the best score on the planet according to one tester), but only scoring “okay” on GDPval (1545) and 19.1% on the harder Terminal-Bench 4.0.
- Token usage rose by up to 30% per task compared to the previous Gemini flash model according to Artificial Intelligence Index data, which matters for anyone tracking real production costs rather than sticker price.
- Google also quietly shipped Gemini 3.8 Flash Cyber, a cybersecurity-focused variant limited to trusted partners, which scores close to more expensive frontier models on CyberGym-style benchmarks at a fraction of the cost.
- Demo tests against GLM 5.3 and GPT 5.6 Soul put Gemini 3.8 Flash in the middle of the pack for 3D biome generation and product landing pages, but produced standout results on a 3D topographic map of Mount Everest and a playable Doom-style demo.
How does Gemini 3.8 Flash perform on benchmarks?
The headline number is Deep SWE v1.1, a long-horizon software engineering benchmark that both testers in this roundup treat as one of the more trustworthy proxies for how a model actually feels to use day to day. Gemini 3.8 Flash scores 73.7% there, essentially tied with Claude Opus 5 and ahead of GPT 5.6 Soul’s 72.7%. For a flash-tier model, that’s a strong result, and it’s the number Google leans on hardest in its own materials.
Outside that one benchmark, results scatter. On GDPval, a knowledge-work benchmark built by OpenAI covering tasks like PDF extraction and presentation building, Gemini 3.8 Flash scores 1545, well behind Opus 5 (1824) and GPT 5.6 Soul (1710). On Terminal-Bench 2.1, an agentic terminal-coding benchmark, it hits a leading 89.4%, but that benchmark is described as largely saturated. On the newer, harder Terminal-Bench 4.0, its score drops to 19.1%, ahead of Sonnet 5 but well behind Opus 5 and GPT 5.6 Soul. On OSWorld, a computer-use benchmark, it scores a respectable 59% against Opus 5’s 75%.
Then there are the categories where it wins outright. It tops Humanity’s Last Exam at 55.9, and it posts the best score on the Harvey legal agent benchmark at 61.4%, ahead of even its own predecessor, Gemini 3.7 Flash. The takeaway from both source videos is consistent: this is not a model that’s uniformly strong or weak, it’s a model that’s very good at specific things (legal reasoning, long-horizon coding, general knowledge exams) and mediocre at others (real-world knowledge work, the hardest agentic terminal tasks).
Why does the harness matter more than the model here?
One of the more useful observations from hands-on testing is that the exact same prompt, run through the exact same model, produces meaningfully different output depending on the surrounding tool. Testing a detailed coding prompt directly in the Gemini app produced a comparatively thin result. Running the identical prompt inside Antigravity, Google’s own agentic coding harness, produced a far more elaborate output: a medieval castle simulation with internal building mechanics, ambient sound effects, a cannonball “cart mode” for knocking down structures, a flying dragon, adjustable time of day with realistic shadow rendering, and changeable weather.
The explanation given is that Antigravity supports multi-turn verification, where the model implements a piece of functionality, checks its own work, and then revises it, rather than generating a single-shot response. That loop appears to matter more for output quality than raw model capability in isolation. The practical implication for anyone evaluating a new model release: testing it only through a chat interface will understate what it can actually do in an agentic coding environment, and comparisons across models are only fair if the harness is held constant.
The clearest demonstration of this cited in testing was an ISS tracker built with a live API call refreshed every five seconds, rendering the station’s real-time position on a 3D Earth split between day and night lighting based on actual time zones, with optional weather layers. That kind of result, combining accurate data-fetching, 3D rendering, and creative polish, was singled out as one of the best single outputs seen from any model tested in that harness.
How does it compare to GLM 5.3 and GPT 5.6 Soul in practice?
A separate round of side-by-side testing ran Gemini 3.8 Flash against GLM 5.3, Fable 5.1, and GPT 5.6 Soul across identical prompts.
For generating seven low-poly 3D biomes, Gemini 3.8 Flash landed in a rough tie with GLM 5.3, behind GPT 5.6 Soul (the favorite in that comparison) and Fable 5.1. It showed small rendering glitches, such as fish clustering oddly under a single boat and water geometry poking through an ice biome.
For simple product landing pages (Apple, an Nvidia DGX Spark, a rubber duck company, a Galaxy Z Fold, a Tesla Model Y), results were mixed. The DGX Spark page came out clean with working interactive stats. The Galaxy Z Fold page had a functional drag-to-fold interaction but looked nothing like an actual phone. The Apple and rubber duck pages were ranked among the weaker outputs across all four models tested, with the Apple checkout page in particular flagged as the worst individual result in that batch.
For a PowerPoint deck on data centers, the model correctly picked up brand colors and typography from a provided reference and produced structurally sound, information-correct slides, though testers noted a Claude model would likely be the better pick for a design-heavy deck.
The two standout wins in this round were unrelated to any competitor comparison: a 3D topographic map of Mount Everest with adjustable cutaway angles, vertical exaggeration, solar azimuth sliders, and labeled camp locations, and a playable, if simple, Doom-style first-person shooter generated from a single prompt.
Is Gemini 3.8 Flash worth using for coding work?
For cost-sensitive, well-scoped coding tasks, yes. The combination of near-Opus-5 Deep SWE performance and a fraction of the token price makes it a reasonable default for production agentic workloads where you don’t need frontier-level performance on every request. Google’s own positioning, reinforced by the Pareto-frontier framing from the Artificial Intelligence Index, is explicit that this is a workhorse model, not a frontier flagship.
The caveats are real. Token consumption rose by as much as 30% per task compared to the prior flash generation, which eats into the raw per-token price advantage once you measure cost per completed task rather than cost per token. Performance on harder agentic benchmarks like Terminal-Bench 4.0 lags well behind Opus 5 and GPT 5.6 Soul. And output quality is highly sensitive to harness: expect underwhelming results in a plain chat interface and noticeably better results in a proper agentic coding environment like Antigravity.
Frequently Asked Questions
What is the difference between Gemini 3.8 Flash and Gemini 3.8 Flash Cyber?
Gemini 3.8 Flash is the general-purpose model. Gemini 3.8 Flash Cyber is a variant tuned for cybersecurity tasks with fewer built-in cyber guardrails, available only to trusted partners through Google’s Fair Wind program rather than to the general public.
How much does Gemini 3.8 Flash cost?
Other agents ship a demo. Remy ships an app.
Real backend. Real database. Real auth. Real plumbing. Remy has it all.
Introductory pricing is 75 cents per million input tokens and $3.75 per million output tokens, in effect through the end of the year before the price rises, according to figures shown in Google’s own materials. That’s still substantially cheaper than Claude Opus 5, GPT 5.6 Soul, or GPT 5.6 Terra even after the increase.
Does Gemini 3.8 Flash beat Claude Opus 5?
It essentially matches Opus 5 on the Deep SWE v1.1 benchmark and wins on Humanity’s Last Exam and the Harvey legal benchmark, but it trails Opus 5 by a wide margin on harder benchmarks like Terminal-Bench 4.0 and on real-world knowledge work (GDPval).
Why does the same prompt give different results in different tools?
Output quality depends on the harness wrapped around the model. Simple chat interfaces like the Gemini app return single-shot responses, while agentic harnesses like Antigravity let the model implement, verify, and revise its own work across multiple turns, which produced visibly richer results in side-by-side testing.
Is Gemini 3.8 Flash good for non-coding tasks like presentations?
It can produce structurally correct, on-brand PowerPoint decks and passable landing pages, but testers noted it lacks strong visual design instincts compared to some Claude models, making it more suitable for content-correct output than design-heavy deliverables.