Skip to main content
MindStudio
Pricing
BlogAbout
My Workspace
Text to Speech Model

Gemini 3.1 Flash TTS

Gemini 3.1 Flash TTS is a text-to-speech model from Google that converts text into audio with style control.

PublisherGoogle
TypeText to Speech
Context Window16,384 tokens
ReleasedApril 2026
Input$1.00/MTok
Output$20.00/MTok
LATEST

Text to speech synthesis from Google Gemini

Gemini 3.1 Flash TTS is a text-to-speech model developed by Google and released in April 2026 under the Gemini model family. It accepts text input and produces synthesized speech, with a context window of 16,384 tokens. The model supports voice selection and style instructions as inputs, giving developers control over how the generated audio sounds.

This model is suited for applications that require programmatic speech generation, such as voice interfaces, accessibility tools, content narration, and automated audio production. The style instruction input allows users to specify delivery characteristics like tone, pacing, or affect directly through a prompt. As a first-party Google model available on MindStudio, it can be integrated into AI workflows without managing separate API credentials.

What Gemini 3.1 Flash TTS supports

Text to Speech

Converts written text into synthesized audio output using a 16,384-token context window, enabling long-form narration in a single request.

Voice Selection

Allows users to choose from available voice options via a select input, controlling the speaker identity of the generated audio.

Style Instructions

Accepts a natural-language prompt to guide delivery style, such as tone, pacing, or emotional affect, without requiring code-level configuration.

Long Context Input

Supports up to 16,384 tokens of input text, accommodating lengthy documents, scripts, or articles in a single synthesis request.

Ready to build with Gemini 3.1 Flash TTS?

Get Started Free

Common questions about Gemini 3.1 Flash TTS

What is the context window for Gemini 3.1 Flash TTS?

Gemini 3.1 Flash TTS has a context window of 16,384 tokens, which also matches its maximum response size.

What inputs does Gemini 3.1 Flash TTS accept?

The model accepts two inputs: a voice selector, which lets you choose the speaker, and a style instruction prompt, which lets you describe how the speech should be delivered.

What is the pricing for Gemini 3.1 Flash TTS?

Pricing information for Gemini 3.1 Flash TTS has not been published in the available metadata. Check MindStudio or Google's official pricing pages for current rates.

When was Gemini 3.1 Flash TTS released?

Gemini 3.1 Flash TTS was released in April 2026 and is currently available with a status of 'available' on MindStudio.

Does Gemini 3.1 Flash TTS support image or video input?

No. Gemini 3.1 Flash TTS is a text-to-speech model and does not support image or video inputs. Its inputs are limited to a voice selection and a text-based style instruction.

Parameters & options

Max Response Size16,384 tokens
VoiceSelect

Prebuilt voice preset to use.

Default: Kore
Zephyr (bright)Puck (upbeat)Charon (informative)Kore (firm)Fenrir (excitable)Leda (youthful)Orus (firm)Aoede (breezy)Callirrhoe (easy-going)Autonoe (bright)Enceladus (breathy)Iapetus (clear)Umbriel (easy-going)Algieba (smooth)Despina (smooth)Erinome (clear)Algenib (gravelly)Rasalgethi (informative)Laomedeia (upbeat)Achernar (soft)Alnilam (firm)Schedar (even)Gacrux (mature)Pulcherrima (forward)Achird (friendly)Zubenelgenubi (casual)Vindemiatrix (gentle)Sadachbia (lively)Sadaltager (knowledgeable)Sulafat (warm)
Style InstructionPrompt

Optional natural-language direction for delivery (e.g. "Say cheerfully:", "Whisper softly:", "Narrate dramatically:"). Prepended to the input before synthesis. Leave blank for a neutral read. You can also embed expressive audio tags directly in your input text like [happy], [whisper], [laughing].

Start building with Gemini 3.1 Flash TTS

No API keys required. Create AI-powered workflows with Gemini 3.1 Flash TTS in minutes — free.