Gemini 2.5 Flash
Gemini 2.5 Flash is a multimodal vision model from Google with a 1,000,000-token context window and configurable thinking.
Fast multimodal reasoning with large context
Gemini 2.5 Flash is a vision-capable language model developed by Google and released in June 2025. It supports a context window of up to one million tokens and can process both text and images as input. The model includes a configurable thinking budget, allowing developers to control how much internal reasoning the model applies before generating a response.
Gemini 2.5 Flash is designed for use cases that require a balance of speed, cost efficiency, and multimodal understanding at scale. Its large context window makes it well-suited for tasks involving long documents, extended conversations, or large volumes of mixed-format content. The model also supports tool use, making it applicable to agent-based workflows and structured task automation.
What Gemini 2.5 Flash supports
Large Context Window
Processes up to 1,000,000 tokens in a single request, enabling analysis of long documents, codebases, or extended conversation histories.
Image Understanding
Accepts image inputs alongside text, supporting tasks like visual question answering, document parsing, and image-based reasoning.
Configurable Thinking
Exposes a thinking budget input that lets developers set how much internal reasoning the model performs before producing output, with a numeric limit parameter.
Tool Use
Supports structured tool inputs, enabling the model to call external functions or APIs as part of agent-based and multi-step workflows.
Priority Mode
Includes a priority mode selector that allows developers to tune the model's behavior toward speed or quality depending on the task.
Fast Inference
Tagged as fast, making it suitable for latency-sensitive applications that also require multimodal or long-context capabilities.
Cost-Effective Scaling
Positioned as a cost-effective option within the Gemini 2.5 family, suitable for high-volume production workloads.
Ready to build with Gemini 2.5 Flash?
Get Started FreeBenchmark scores
Scores represent accuracy — the percentage of questions answered correctly on each test.
| Benchmark | What it tests | Score |
|---|---|---|
| MMLU-Pro | Expert knowledge across 14 academic disciplines | 80.9% |
| GPQA Diamond | PhD-level science questions (biology, physics, chemistry) | 68.3% |
| MATH-500 | Undergraduate and competition-level math problems | 93.2% |
| AIME 2024 | American math olympiad problems | 50.0% |
| LiveCodeBench | Real-world coding tasks from recent competitions | 49.5% |
| HLE | Questions that challenge frontier models across many domains | 5.1% |
| SciCode | Scientific research coding and numerical methods | 29.1% |
Common questions about Gemini 2.5 Flash
What is the context window size for Gemini 2.5 Flash?
Gemini 2.5 Flash supports a context window of 1,000,000 tokens, allowing it to process very long documents or extended conversations in a single request.
What is the maximum response size?
The model can generate responses of up to 64,000 tokens per request.
Does Gemini 2.5 Flash support image inputs?
Yes, Gemini 2.5 Flash is a vision model and accepts image inputs alongside text prompts.
What is the thinking budget and how does it work?
The thinking budget is a configurable input that controls how much internal reasoning the model performs before generating a response. Developers can set a numeric limit to balance response quality against latency and cost.
What is the pricing for Gemini 2.5 Flash on MindStudio?
Specific pricing was not included in the available metadata. Check the MindStudio platform or Google's official API pricing page for current rates.
What is the knowledge cutoff date for Gemini 2.5 Flash?
The knowledge cutoff date is not specified in the available metadata. Google's official documentation for Gemini 2.5 Flash is the best source for this information.
What people think about Gemini 2.5 Flash
Community discussions around Gemini 2.5 Flash reflect general interest in its image editing capabilities and the rollout of related model variants, with posts about the Flash Image model receiving hundreds of upvotes for demonstrating high-level image edits. Users have also noted Google's informal naming of a Flash image preview variant as "Nano Banana," which generated notable engagement.
A recurring concern in the community is pricing, with one widely discussed thread noting that the cost of thinking output tokens doubled from $0.15 to $0.30 after the model reached general availability. This pricing change prompted significant discussion among developers evaluating the model for cost-sensitive production use cases.
Google is now officially calling "Gemini 2.5 Flash image preview", "Nano Banana"
Google doubled the price of Gemini 2.5 Flash thinking output after GA from 0.15 to 0.30 what
Google's new Gemini 2.5 Flash Image model can do some very impressive high-level image edits
BREAKING: OpenAI releases "GPT-Image-1.5" (ChatGPT Images) & It instantly takes the #1 Spot on LMArena, beating Google's Nano Banana Pro.
Parameters & options
Must be less than Max Response Size
Explore similar models
Start building with Gemini 2.5 Flash
No API keys required. Create AI-powered workflows with Gemini 2.5 Flash in minutes — free.