Gemini 2.5 Flash Lite
Gemini 2.5 Flash Lite is a fast text and image understanding model from Google with a 1,000,000-token context window.
Fast multimodal text generation with thinking budget
Gemini 2.5 Flash Lite is a text generation model developed by Google, released in July 2025. It belongs to the Gemini 2.5 family and is positioned as a lightweight, speed-optimized variant of the Flash line, designed for high-throughput workloads where latency and cost efficiency matter. The model supports a 1,000,000-token context window and accepts image inputs alongside text, making it suitable for multimodal tasks at scale.
A distinguishing feature of Gemini 2.5 Flash Lite is its configurable thinking budget, which allows developers to control how much internal reasoning the model applies before generating a response. This makes it adaptable across use cases ranging from quick classification and summarization tasks to more deliberate, structured reasoning when needed. The model is well-suited for applications that process large documents, handle high request volumes, or require a balance between response speed and reasoning depth.
What Gemini 2.5 Flash Lite supports
Large Context Window
Processes up to 1,000,000 tokens in a single request, enabling analysis of very long documents, codebases, or conversation histories without truncation.
Image Understanding
Accepts image inputs alongside text prompts, supporting tasks like visual question answering, image description, and document parsing.
Configurable Thinking Budget
Exposes a Thinking Budget control that lets developers set how much internal reasoning the model performs, with a numeric limit to cap token usage.
Priority Mode
Offers a Priority Mode selector that allows developers to tune the trade-off between response speed and output quality at the request level.
Fast Inference
Tagged as FAST, the model is optimized for low-latency responses, making it suitable for real-time applications and high-throughput pipelines.
Long-Form Text Generation
Supports a maximum response size of 65,535 tokens, allowing generation of detailed reports, long-form content, and extended structured outputs.
Ready to build with Gemini 2.5 Flash Lite?
Get Started FreeBenchmark scores
Scores represent accuracy — the percentage of questions answered correctly on each test.
| Benchmark | What it tests | Score |
|---|---|---|
| MMLU-Pro | Expert knowledge across 14 academic disciplines | 72.4% |
| GPQA Diamond | PhD-level science questions (biology, physics, chemistry) | 47.4% |
| MATH-500 | Undergraduate and competition-level math problems | 92.6% |
| AIME 2024 | American math olympiad problems | 50.0% |
| LiveCodeBench | Real-world coding tasks from recent competitions | 40.0% |
| HLE | Questions that challenge frontier models across many domains | 3.7% |
| SciCode | Scientific research coding and numerical methods | 17.7% |
Common questions about Gemini 2.5 Flash Lite
What is the context window size for Gemini 2.5 Flash Lite?
Gemini 2.5 Flash Lite supports a context window of 1,000,000 tokens, allowing very large documents or long conversation histories to be processed in a single request.
Does Gemini 2.5 Flash Lite support image inputs?
Yes, the model accepts image inputs in addition to text, enabling multimodal tasks such as visual question answering and image-based document analysis.
What is the Thinking Budget and how does it work?
The Thinking Budget is a configurable parameter that controls how much internal reasoning the model performs before generating a response. Developers can set a numeric limit to cap the tokens used for thinking, balancing reasoning depth against speed and cost.
What is the maximum response length?
The model can generate responses of up to 65,535 tokens in a single output, suitable for long-form content generation and detailed structured responses.
When was Gemini 2.5 Flash Lite released?
Gemini 2.5 Flash Lite was released in July 2025 and became available on MindStudio in September 2025.
What people think about Gemini 2.5 Flash Lite
Community discussion around Gemini 2.5 Flash Lite is generally positive, with users highlighting the addition of Thinking, Live Audio, and Grounding as notable features for Google's most affordable model in the 2.5 family. The thread received 139 upvotes with minimal controversy, suggesting broad approval of the feature expansion.
Commenters focused primarily on the model's cost-efficiency and the practical value of optional reasoning at a low price point, with limited discussion of limitations. The small comment count (5) indicates the announcement was well-received but did not generate significant debate.
Parameters & options
Must be less than Max Response Size
Explore similar models
Start building with Gemini 2.5 Flash Lite
No API keys required. Create AI-powered workflows with Gemini 2.5 Flash Lite in minutes — free.