Skip to main content
MindStudio
Pricing
BlogAbout
My Workspace
Text Generation Model

Gemini 2.5 Flash

Gemini 2.5 Flash is a multimodal vision model from Google with a 1,000,000-token context window and configurable thinking.

PublisherGoogle
TypeText
Context Window1,000,000 tokens
ReleasedJune 2025
Input$0.30/MTok
Output$2.50/MTok
FASTCOST EFFECTIVELARGE CONTEXTMULTI-MODAL

Fast multimodal reasoning with large context

Gemini 2.5 Flash is a vision-capable language model developed by Google and released in June 2025. It supports a context window of up to one million tokens and can process both text and images as input. The model includes a configurable thinking budget, allowing developers to control how much internal reasoning the model applies before generating a response.

Gemini 2.5 Flash is designed for use cases that require a balance of speed, cost efficiency, and multimodal understanding at scale. Its large context window makes it well-suited for tasks involving long documents, extended conversations, or large volumes of mixed-format content. The model also supports tool use, making it applicable to agent-based workflows and structured task automation.

What Gemini 2.5 Flash supports

Large Context Window

Processes up to 1,000,000 tokens in a single request, enabling analysis of long documents, codebases, or extended conversation histories.

Image Understanding

Accepts image inputs alongside text, supporting tasks like visual question answering, document parsing, and image-based reasoning.

Configurable Thinking

Exposes a thinking budget input that lets developers set how much internal reasoning the model performs before producing output, with a numeric limit parameter.

Tool Use

Supports structured tool inputs, enabling the model to call external functions or APIs as part of agent-based and multi-step workflows.

Priority Mode

Includes a priority mode selector that allows developers to tune the model's behavior toward speed or quality depending on the task.

Fast Inference

Tagged as fast, making it suitable for latency-sensitive applications that also require multimodal or long-context capabilities.

Cost-Effective Scaling

Positioned as a cost-effective option within the Gemini 2.5 family, suitable for high-volume production workloads.

Ready to build with Gemini 2.5 Flash?

Get Started Free

Benchmark scores

Scores represent accuracy — the percentage of questions answered correctly on each test.

BenchmarkWhat it testsScore
MMLU-ProExpert knowledge across 14 academic disciplines80.9%
GPQA DiamondPhD-level science questions (biology, physics, chemistry)68.3%
MATH-500Undergraduate and competition-level math problems93.2%
AIME 2024American math olympiad problems50.0%
LiveCodeBenchReal-world coding tasks from recent competitions49.5%
HLEQuestions that challenge frontier models across many domains5.1%
SciCodeScientific research coding and numerical methods29.1%

Common questions about Gemini 2.5 Flash

What is the context window size for Gemini 2.5 Flash?

Gemini 2.5 Flash supports a context window of 1,000,000 tokens, allowing it to process very long documents or extended conversations in a single request.

What is the maximum response size?

The model can generate responses of up to 64,000 tokens per request.

Does Gemini 2.5 Flash support image inputs?

Yes, Gemini 2.5 Flash is a vision model and accepts image inputs alongside text prompts.

What is the thinking budget and how does it work?

The thinking budget is a configurable input that controls how much internal reasoning the model performs before generating a response. Developers can set a numeric limit to balance response quality against latency and cost.

What is the pricing for Gemini 2.5 Flash on MindStudio?

Specific pricing was not included in the available metadata. Check the MindStudio platform or Google's official API pricing page for current rates.

What is the knowledge cutoff date for Gemini 2.5 Flash?

The knowledge cutoff date is not specified in the available metadata. Google's official documentation for Gemini 2.5 Flash is the best source for this information.

What people think about Gemini 2.5 Flash

Community discussions around Gemini 2.5 Flash reflect general interest in its image editing capabilities and the rollout of related model variants, with posts about the Flash Image model receiving hundreds of upvotes for demonstrating high-level image edits. Users have also noted Google's informal naming of a Flash image preview variant as "Nano Banana," which generated notable engagement.

A recurring concern in the community is pricing, with one widely discussed thread noting that the cost of thinking output tokens doubled from $0.15 to $0.30 after the model reached general availability. This pricing change prompted significant discussion among developers evaluating the model for cost-sensitive production use cases.

View more discussions →

Parameters & options

Max Temperature1
Max Response Size64,000 tokens
Thinking BudgetSelect
Default: auto
OffManualAuto
Thinking Budget LimitNumber

Must be less than Max Response Size

Range: 1–24576
ToolsTools
Google SearchGround responses in current events and facts from the web to reduce hallucinations.
Google MapsBuild location-aware assistants that can find places, get directions, and provide rich local context.
Code ExecutionAllow the model to write and run Python code to solve math problems or process data accurately.
URL ContextDirect the model to read and analyze content from specific web pages or documents.
Priority ModeSelect
Default: false
StandardPriority (1.8x cost, higher reliability)

Start building with Gemini 2.5 Flash

No API keys required. Create AI-powered workflows with Gemini 2.5 Flash in minutes — free.