Skip to main content
MindStudio
Pricing
BlogAbout
My Workspace
Text Generation Model

Gemini 3.5 Flash Lite

Gemini 3.5 Flash Lite is a text generation model from Google designed for low-cost, real-time tasks with a 1M token context window.

PublisherGoogle
TypeText
Context Window1,048,576 tokens
ReleasedJuly 2026
Input$0.30/MTok
Output$2.50/MTok
LARGE CONTEXTREAL-TIME LATENCYLOW COST

Lightweight text generation with large context

Gemini 3.5 Flash Lite is a text generation model developed by Google, released in July 2026 as part of the Gemini model family. It is designed for use cases that require low latency and low cost, while still supporting a context window of 1,048,576 tokens and a maximum response size of 65,536 tokens. The model accepts image inputs alongside text, making it suitable for multimodal tasks that need to run at scale or in real time.

This model includes a configurable Thinking Level input, which allows developers to adjust how much internal reasoning the model applies before generating a response, offering a tradeoff between speed and depth. It also supports tool use, enabling integration with external functions and APIs within agentic workflows. Gemini 3.5 Flash Lite is best suited for high-volume, latency-sensitive applications such as classification, summarization, extraction, and lightweight assistants where cost efficiency is a priority.

What Gemini 3.5 Flash Lite supports

Large Context Window

Supports up to 1,048,576 tokens of context in a single request, allowing long documents, transcripts, or multi-turn conversations to be processed without truncation.

Real-Time Latency

Optimized for low-latency inference, making it suitable for interactive applications and pipelines where response speed is a requirement.

Low Cost Operation

Positioned as a cost-efficient model within the Gemini family, intended for high-volume workloads where per-request cost needs to be minimized.

Image Input Support

Accepts image inputs alongside text prompts, enabling multimodal tasks such as visual question answering and image-based extraction.

Configurable Thinking Level

Exposes a Thinking Level selector that lets developers control the depth of internal reasoning before output, trading latency for response quality as needed.

Tool Use

Supports tool and function calling inputs, allowing the model to be integrated into agentic workflows that invoke external APIs or structured actions.

Ready to build with Gemini 3.5 Flash Lite?

Get Started Free

Common questions about Gemini 3.5 Flash Lite

What is the context window size for Gemini 3.5 Flash Lite?

Gemini 3.5 Flash Lite supports a context window of 1,048,576 tokens, which allows very long documents or extended conversation histories to be included in a single request.

What is the maximum response length this model can produce?

The model has a maximum response size of 65,536 tokens per generation.

Does Gemini 3.5 Flash Lite support image inputs?

Yes, the model supports image inputs alongside text, making it capable of handling multimodal prompts.

What is the Thinking Level input and how does it work?

Thinking Level is a configurable input that controls how much internal reasoning the model performs before generating a response. Adjusting it lets you trade off between faster, lighter responses and more deliberate outputs.

When was Gemini 3.5 Flash Lite released?

Gemini 3.5 Flash Lite was released in July 2026 by Google.

What kinds of tasks is this model best suited for?

Based on its tags and design, the model is best suited for high-volume, latency-sensitive tasks such as classification, summarization, data extraction, and lightweight conversational assistants where cost efficiency matters.

Parameters & options

Max Temperature2
Max Response Size65,536 tokens
Thinking LevelSelect
Default: low
MinimalLowMediumHigh
ToolsTools
Google SearchGround responses in current events and facts from the web to reduce hallucinations.
Code ExecutionAllow the model to write and run Python code to solve math problems or process data accurately.
URL ContextDirect the model to read and analyze content from specific web pages or documents.

Start building with Gemini 3.5 Flash Lite

No API keys required. Create AI-powered workflows with Gemini 3.5 Flash Lite in minutes — free.