Gemini 3.1 Flash TTS
Gemini 3.1 Flash TTS is a text-to-speech model from Google that converts text into audio with style control.
Text to speech synthesis from Google Gemini
Gemini 3.1 Flash TTS is a text-to-speech model developed by Google and released in April 2026 under the Gemini model family. It accepts text input and produces synthesized speech, with a context window of 16,384 tokens. The model supports voice selection and style instructions as inputs, giving developers control over how the generated audio sounds.
This model is suited for applications that require programmatic speech generation, such as voice interfaces, accessibility tools, content narration, and automated audio production. The style instruction input allows users to specify delivery characteristics like tone, pacing, or affect directly through a prompt. As a first-party Google model available on MindStudio, it can be integrated into AI workflows without managing separate API credentials.
What Gemini 3.1 Flash TTS supports
Text to Speech
Converts written text into synthesized audio output using a 16,384-token context window, enabling long-form narration in a single request.
Voice Selection
Allows users to choose from available voice options via a select input, controlling the speaker identity of the generated audio.
Style Instructions
Accepts a natural-language prompt to guide delivery style, such as tone, pacing, or emotional affect, without requiring code-level configuration.
Long Context Input
Supports up to 16,384 tokens of input text, accommodating lengthy documents, scripts, or articles in a single synthesis request.
Ready to build with Gemini 3.1 Flash TTS?
Get Started FreeCommon questions about Gemini 3.1 Flash TTS
What is the context window for Gemini 3.1 Flash TTS?
Gemini 3.1 Flash TTS has a context window of 16,384 tokens, which also matches its maximum response size.
What inputs does Gemini 3.1 Flash TTS accept?
The model accepts two inputs: a voice selector, which lets you choose the speaker, and a style instruction prompt, which lets you describe how the speech should be delivered.
What is the pricing for Gemini 3.1 Flash TTS?
Pricing information for Gemini 3.1 Flash TTS has not been published in the available metadata. Check MindStudio or Google's official pricing pages for current rates.
When was Gemini 3.1 Flash TTS released?
Gemini 3.1 Flash TTS was released in April 2026 and is currently available with a status of 'available' on MindStudio.
Does Gemini 3.1 Flash TTS support image or video input?
No. Gemini 3.1 Flash TTS is a text-to-speech model and does not support image or video inputs. Its inputs are limited to a voice selection and a text-based style instruction.
Documentation & links
Parameters & options
Prebuilt voice preset to use.
Optional natural-language direction for delivery (e.g. "Say cheerfully:", "Whisper softly:", "Narrate dramatically:"). Prepended to the input before synthesis. Leave blank for a neutral read. You can also embed expressive audio tags directly in your input text like [happy], [whisper], [laughing].
Explore similar models
Start building with Gemini 3.1 Flash TTS
No API keys required. Create AI-powered workflows with Gemini 3.1 Flash TTS in minutes — free.