Skip to main content
MindStudio
Pricing
BlogAbout
My Workspace
Text Generation Model

Gemini 2.5 Flash Lite

Gemini 2.5 Flash Lite is a fast text and image understanding model from Google with a 1,000,000-token context window.

PublisherGoogle
TypeText
Context Window1,000,000 tokens
ReleasedJuly 2025
Input$0.10/MTok
Output$0.40/MTok
FAST

Fast multimodal text generation with thinking budget

Gemini 2.5 Flash Lite is a text generation model developed by Google, released in July 2025. It belongs to the Gemini 2.5 family and is positioned as a lightweight, speed-optimized variant of the Flash line, designed for high-throughput workloads where latency and cost efficiency matter. The model supports a 1,000,000-token context window and accepts image inputs alongside text, making it suitable for multimodal tasks at scale.

A distinguishing feature of Gemini 2.5 Flash Lite is its configurable thinking budget, which allows developers to control how much internal reasoning the model applies before generating a response. This makes it adaptable across use cases ranging from quick classification and summarization tasks to more deliberate, structured reasoning when needed. The model is well-suited for applications that process large documents, handle high request volumes, or require a balance between response speed and reasoning depth.

What Gemini 2.5 Flash Lite supports

Large Context Window

Processes up to 1,000,000 tokens in a single request, enabling analysis of very long documents, codebases, or conversation histories without truncation.

Image Understanding

Accepts image inputs alongside text prompts, supporting tasks like visual question answering, image description, and document parsing.

Configurable Thinking Budget

Exposes a Thinking Budget control that lets developers set how much internal reasoning the model performs, with a numeric limit to cap token usage.

Priority Mode

Offers a Priority Mode selector that allows developers to tune the trade-off between response speed and output quality at the request level.

Fast Inference

Tagged as FAST, the model is optimized for low-latency responses, making it suitable for real-time applications and high-throughput pipelines.

Long-Form Text Generation

Supports a maximum response size of 65,535 tokens, allowing generation of detailed reports, long-form content, and extended structured outputs.

Ready to build with Gemini 2.5 Flash Lite?

Get Started Free

Benchmark scores

Scores represent accuracy — the percentage of questions answered correctly on each test.

BenchmarkWhat it testsScore
MMLU-ProExpert knowledge across 14 academic disciplines72.4%
GPQA DiamondPhD-level science questions (biology, physics, chemistry)47.4%
MATH-500Undergraduate and competition-level math problems92.6%
AIME 2024American math olympiad problems50.0%
LiveCodeBenchReal-world coding tasks from recent competitions40.0%
HLEQuestions that challenge frontier models across many domains3.7%
SciCodeScientific research coding and numerical methods17.7%

Common questions about Gemini 2.5 Flash Lite

What is the context window size for Gemini 2.5 Flash Lite?

Gemini 2.5 Flash Lite supports a context window of 1,000,000 tokens, allowing very large documents or long conversation histories to be processed in a single request.

Does Gemini 2.5 Flash Lite support image inputs?

Yes, the model accepts image inputs in addition to text, enabling multimodal tasks such as visual question answering and image-based document analysis.

What is the Thinking Budget and how does it work?

The Thinking Budget is a configurable parameter that controls how much internal reasoning the model performs before generating a response. Developers can set a numeric limit to cap the tokens used for thinking, balancing reasoning depth against speed and cost.

What is the maximum response length?

The model can generate responses of up to 65,535 tokens in a single output, suitable for long-form content generation and detailed structured responses.

When was Gemini 2.5 Flash Lite released?

Gemini 2.5 Flash Lite was released in July 2025 and became available on MindStudio in September 2025.

What people think about Gemini 2.5 Flash Lite

Community discussion around Gemini 2.5 Flash Lite is generally positive, with users highlighting the addition of Thinking, Live Audio, and Grounding as notable features for Google's most affordable model in the 2.5 family. The thread received 139 upvotes with minimal controversy, suggesting broad approval of the feature expansion.

Commenters focused primarily on the model's cost-efficiency and the practical value of optional reasoning at a low price point, with limited discussion of limitations. The small comment count (5) indicates the announcement was well-received but did not generate significant debate.

View more discussions →

Parameters & options

Max Temperature2
Max Response Size65,535 tokens
Thinking BudgetSelect
Default: auto
OffManualAuto
Thinking Budget LimitNumber

Must be less than Max Response Size

Range: 1–24576
Priority ModeSelect
Default: false
StandardPriority (1.8x cost, higher reliability)

Start building with Gemini 2.5 Flash Lite

No API keys required. Create AI-powered workflows with Gemini 2.5 Flash Lite in minutes — free.