Gemini 3.5 Flash Lite
Gemini 3.5 Flash Lite is a text generation model from Google designed for low-cost, real-time tasks with a 1M token context window.
Lightweight text generation with large context
Gemini 3.5 Flash Lite is a text generation model developed by Google, released in July 2026 as part of the Gemini model family. It is designed for use cases that require low latency and low cost, while still supporting a context window of 1,048,576 tokens and a maximum response size of 65,536 tokens. The model accepts image inputs alongside text, making it suitable for multimodal tasks that need to run at scale or in real time.
This model includes a configurable Thinking Level input, which allows developers to adjust how much internal reasoning the model applies before generating a response, offering a tradeoff between speed and depth. It also supports tool use, enabling integration with external functions and APIs within agentic workflows. Gemini 3.5 Flash Lite is best suited for high-volume, latency-sensitive applications such as classification, summarization, extraction, and lightweight assistants where cost efficiency is a priority.
What Gemini 3.5 Flash Lite supports
Large Context Window
Supports up to 1,048,576 tokens of context in a single request, allowing long documents, transcripts, or multi-turn conversations to be processed without truncation.
Real-Time Latency
Optimized for low-latency inference, making it suitable for interactive applications and pipelines where response speed is a requirement.
Low Cost Operation
Positioned as a cost-efficient model within the Gemini family, intended for high-volume workloads where per-request cost needs to be minimized.
Image Input Support
Accepts image inputs alongside text prompts, enabling multimodal tasks such as visual question answering and image-based extraction.
Configurable Thinking Level
Exposes a Thinking Level selector that lets developers control the depth of internal reasoning before output, trading latency for response quality as needed.
Tool Use
Supports tool and function calling inputs, allowing the model to be integrated into agentic workflows that invoke external APIs or structured actions.
Ready to build with Gemini 3.5 Flash Lite?
Get Started FreeCommon questions about Gemini 3.5 Flash Lite
What is the context window size for Gemini 3.5 Flash Lite?
Gemini 3.5 Flash Lite supports a context window of 1,048,576 tokens, which allows very long documents or extended conversation histories to be included in a single request.
What is the maximum response length this model can produce?
The model has a maximum response size of 65,536 tokens per generation.
Does Gemini 3.5 Flash Lite support image inputs?
Yes, the model supports image inputs alongside text, making it capable of handling multimodal prompts.
What is the Thinking Level input and how does it work?
Thinking Level is a configurable input that controls how much internal reasoning the model performs before generating a response. Adjusting it lets you trade off between faster, lighter responses and more deliberate outputs.
When was Gemini 3.5 Flash Lite released?
Gemini 3.5 Flash Lite was released in July 2026 by Google.
What kinds of tasks is this model best suited for?
Based on its tags and design, the model is best suited for high-volume, latency-sensitive tasks such as classification, summarization, data extraction, and lightweight conversational assistants where cost efficiency matters.
Parameters & options
Explore similar models
Start building with Gemini 3.5 Flash Lite
No API keys required. Create AI-powered workflows with Gemini 3.5 Flash Lite in minutes — free.