Gemini 3.1 Flash Lite
Gemini 3.1 Flash Lite is a text generation model from Google with a 1,048,576-token context window and real-time latency.
Fast, large-context text generation at low cost
Gemini 3.1 Flash Lite is a text generation model developed by Google, released in February 2026. It belongs to the Flash Lite tier of the Gemini 3.1 family, a line designed for high-throughput, cost-efficient inference. The model supports a context window of 1,048,576 tokens and can process image inputs alongside text, making it suitable for multimodal tasks. It also exposes a configurable thinking level, allowing developers to adjust the depth of reasoning the model applies at inference time.
Gemini 3.1 Flash Lite is positioned for use cases that require low latency and large context handling without incurring high per-token costs. Its 65,536-token maximum response size accommodates long-form outputs such as document summarization, extended dialogue, and batch content generation. The model supports tool use, enabling integration with external APIs and function-calling workflows. These characteristics make it well-suited for production applications where speed, scale, and cost efficiency are primary constraints.
What Gemini 3.1 Flash Lite supports
Large Context Window
Processes up to 1,048,576 tokens in a single request, enabling analysis of long documents, codebases, or extended conversation histories without truncation.
Real-Time Latency
Optimized for low-latency inference, making it suitable for interactive applications and user-facing products that require fast response times.
Low Cost Inference
Designed for cost-efficient operation at scale, targeting use cases where high request volume makes per-token pricing a primary concern.
Image Input Support
Accepts image inputs alongside text prompts, allowing the model to reason about visual content within the same request.
Tool Use
Supports function calling and external tool integration, enabling the model to invoke APIs or structured workflows as part of its response.
Configurable Thinking Level
Exposes a selectable thinking level input that lets developers control how much internal reasoning the model applies before generating a response.
Long-Form Output
Supports a maximum response size of 65,536 tokens, accommodating extended outputs such as full document drafts, detailed reports, or lengthy code files.
Ready to build with Gemini 3.1 Flash Lite?
Get Started FreeCommon questions about Gemini 3.1 Flash Lite
What is the context window size for Gemini 3.1 Flash Lite?
Gemini 3.1 Flash Lite has a context window of 1,048,576 tokens, allowing it to process very long documents or extended conversations in a single request.
What is the maximum response length this model can produce?
The model supports a maximum response size of 65,536 tokens per request, which is suitable for generating long-form content such as reports, summaries, or extended code.
Does Gemini 3.1 Flash Lite support image inputs?
Yes, the model supports image inputs alongside text, enabling multimodal prompts that combine visual and textual content.
What is the thinking level input and how does it work?
The thinking level is a configurable input that lets you adjust the depth of reasoning the model applies before generating a response. Selecting a higher level may improve reasoning quality at the cost of additional latency.
When was Gemini 3.1 Flash Lite released?
Gemini 3.1 Flash Lite was released in February 2026 by Google.
Does Gemini 3.1 Flash Lite support tool or function calling?
Yes, the model supports tool use, allowing it to call external functions or APIs as part of its response generation workflow.
Parameters & options
Explore similar models
Start building with Gemini 3.1 Flash Lite
No API keys required. Create AI-powered workflows with Gemini 3.1 Flash Lite in minutes — free.