Skip to main content
MindStudio
Pricing
BlogAbout
My Workspace
Text Generation Model

Gemini 3 Flash

Gemini 3 Flash is a text and image understanding model from Google with a 1,048,576-token context window.

PublisherGoogle
TypeText
Context Window1,048,576 tokens
ReleasedDecember 2025
Input$0.50/MTok
Output$3.00/MTok
LARGE CONTEXTREAL-TIME LATENCYLATESTTOOLS

Large-context flash model with thinking support

Gemini 3 Flash is a multimodal text generation model developed by Google, released in December 2025 under the identifier gemini-3-flash-preview. It accepts both text and image inputs and supports a context window of 1,048,576 tokens, making it suitable for tasks that require processing large volumes of content in a single request. The model also exposes a configurable thinking budget, allowing developers to control how much internal reasoning the model applies before producing a response.

Gemini 3 Flash is designed for use cases where low latency matters alongside large-context handling, such as document analysis, multi-turn conversation, and tool-augmented workflows. It supports external tool calling natively and includes a priority mode setting that lets developers tune the balance between speed and thoroughness. The model is available directly through MindStudio without requiring separate API key configuration.

What Gemini 3 Flash supports

Large Context Window

Processes up to 1,048,576 tokens in a single request, enabling analysis of long documents, codebases, or extended conversation histories without truncation.

Real-Time Latency

Optimized for low-latency responses, making it suitable for interactive applications and real-time user-facing workflows.

Tool Calling

Supports native tool use, allowing the model to invoke external functions or APIs as part of a response and enabling agent-style task completion.

Image Understanding

Accepts image inputs alongside text, supporting tasks such as visual question answering, image description, and document parsing from screenshots.

Configurable Thinking

Exposes a thinking budget input that lets developers set how much internal reasoning the model performs, with an optional numeric limit to cap compute usage.

Priority Mode

Includes a priority mode selector that allows developers to tune the trade-off between response speed and output thoroughness at inference time.

Ready to build with Gemini 3 Flash?

Get Started Free

Benchmark scores

Scores represent accuracy — the percentage of questions answered correctly on each test.

BenchmarkWhat it testsScore
MMLU-ProExpert knowledge across 14 academic disciplines88.2%
GPQA DiamondPhD-level science questions (biology, physics, chemistry)81.2%
LiveCodeBenchReal-world coding tasks from recent competitions79.7%
HLEQuestions that challenge frontier models across many domains14.1%
SciCodeScientific research coding and numerical methods49.9%
SWE-bench VerifiedReal GitHub issues requiring multi-file code fixes78.0%

Common questions about Gemini 3 Flash

What is the context window size for Gemini 3 Flash?

Gemini 3 Flash supports a context window of 1,048,576 tokens, which allows very long documents, transcripts, or conversation histories to be processed in a single call.

What is the maximum response size?

The model can return up to 65,535 tokens in a single response.

Does Gemini 3 Flash support image inputs?

Yes. The model accepts image inputs in addition to text, enabling multimodal tasks such as visual question answering and image-based document analysis.

What is the thinking budget, and how does it work?

The thinking budget is a configurable input that controls how much internal reasoning the model applies before generating a response. A numeric limit can also be set to cap the amount of compute used for this reasoning step.

When was Gemini 3 Flash released?

Gemini 3 Flash was released in December 2025 and became available on MindStudio on December 17, 2025.

Is pricing information available for Gemini 3 Flash?

Published pricing details are not listed in the current metadata. Check the MindStudio platform or Google's official documentation for the latest pricing information.

What people think about Gemini 3 Flash

Community reception on Reddit has been largely positive, with users highlighting the model's benchmark results including a reported 99.7% score on AIME and a rank of #3 on LMArena at the time of release. The low cost of approximately $0.50 per 1 million tokens relative to its reported reasoning performance has been a frequently cited point of interest.

Discussions have also focused on specific capabilities such as agentic vision features introduced in a subsequent update, and independent benchmark results including a reported high "Omniscience" score. Some threads reference deleted posts from researchers at Google DeepMind, suggesting community interest in behind-the-scenes development context.

View more discussions →

Parameters & options

Max Temperature2
Max Response Size65,535 tokens
Thinking BudgetSelect
Default: auto
OffManualAuto
Thinking Budget LimitNumber

Must be less than Max Response Size

Range: 1–24576
ToolsTools
Google SearchGround responses in current events and facts from the web to reduce hallucinations.
Google MapsBuild location-aware assistants that can find places, get directions, and provide rich local context.
Code ExecutionAllow the model to write and run Python code to solve math problems or process data accurately.
URL ContextDirect the model to read and analyze content from specific web pages or documents.
Priority ModeSelect
Default: false
StandardPriority (1.8x cost, higher reliability)

Start building with Gemini 3 Flash

No API keys required. Create AI-powered workflows with Gemini 3 Flash in minutes — free.