Skip to main content
MindStudio
Pricing
BlogAbout
My Workspace
Text to Speech Model

GPT-4o-mini TTS

GPT-4o-mini TTS is a text-to-speech model from OpenAI that converts text to audio with adjustable voice, speed, and format.

PublisherOpenAI
TypeText to Speech
Context Window2,000 tokens
ReleasedMarch 2025
Input$0.60/MTok
Output$2.40/MTok
LATEST

Lightweight text-to-speech with voice control

GPT-4o-mini TTS is a text-to-speech model developed by OpenAI and released in March 2025. It accepts a text prompt along with configurable inputs — including voice selection, playback speed, and output format — and returns synthesized audio. The model supports a context window of 2,000 tokens, making it suited for shorter passages, narration, and conversational audio generation.

This model is the smaller, more efficient variant in OpenAI's TTS lineup, designed for use cases where lower latency or cost efficiency is a priority. It accepts natural-language instructions to guide tone and delivery style, giving developers some control over how the speech is rendered beyond just voice selection. GPT-4o-mini TTS is well suited for applications like voice interfaces, audio previews, accessibility tooling, and automated narration pipelines.

What GPT-4o-mini TTS supports

Text-to-Speech Synthesis

Converts a text prompt into spoken audio output. Supports a context window of up to 2,000 tokens per request.

Voice Selection

Allows selection from multiple available voices via a dropdown input. The chosen voice determines the speaker identity of the generated audio.

Delivery Instructions

Accepts a natural-language instructions field to guide tone, pacing, or style of the synthesized speech beyond default voice behavior.

Speed Control

Accepts a numeric input to adjust playback speed of the generated audio, enabling faster or slower delivery as needed.

Output Format Selection

Supports multiple audio output formats selectable at request time, allowing integration with different downstream audio pipelines.

Ready to build with GPT-4o-mini TTS?

Get Started Free

Common questions about GPT-4o-mini TTS

What is the context window for GPT-4o-mini TTS?

GPT-4o-mini TTS has a context window of 2,000 tokens, which limits the length of text that can be converted to audio in a single request.

What inputs does GPT-4o-mini TTS accept?

The model accepts four inputs: a voice selection, a text prompt, a natural-language instructions field for delivery guidance, a numeric speed value, and an output format selection.

What audio output formats does GPT-4o-mini TTS support?

The model supports multiple output formats selectable via the response_format input. OpenAI's TTS API commonly supports formats including mp3, opus, aac, flac, wav, and pcm, though the exact options available may depend on the API version.

How does GPT-4o-mini TTS differ from GPT-4o TTS?

GPT-4o-mini TTS is the smaller variant in OpenAI's TTS lineup. It shares the same input structure but is designed for use cases where lower latency or reduced cost is a priority compared to the full GPT-4o TTS model.

Is pricing information available for GPT-4o-mini TTS?

Pricing details are not included in the available metadata. For current pricing, refer to OpenAI's official pricing page at openai.com/pricing.

Does GPT-4o-mini TTS have a knowledge cutoff?

As a text-to-speech model, GPT-4o-mini TTS does not have a knowledge cutoff in the traditional sense — it synthesizes audio from text input rather than answering questions from a training dataset.

What people think about GPT-4o-mini TTS

Community discussion directly about GPT-4o-mini TTS is limited in the found threads, though one Reddit post notes that OpenAI quietly released updated versions of their TTS and related audio models in the API in December 2025, suggesting ongoing development. General sentiment around OpenAI's audio model updates tends to be positive, with developers noting the availability of new versioned releases.

The threads found are largely focused on other TTS models or broad AI release roundups rather than GPT-4o-mini TTS specifically, so detailed community feedback on its limitations or use cases is not well represented in this sample. Developers interested in community discussion may find more targeted feedback in OpenAI developer forums or the OpenAI API changelog.

View more discussions →

Parameters & options

VoiceSelect

Voice to use in TTS

Default: alloy
AlloyAshCoralEchoFableOnyxNovaSageShimmer
InstructionsPrompt

Control the voice of your generated audio with additional instructions.

SpeedNumber

The speed of the generated audio. (Default is 1.0)

Default: 1Range: 0.25–4 (step 0.05)
Output FormatSelect
Default: mp3
MP3WAV

Start building with GPT-4o-mini TTS

No API keys required. Create AI-powered workflows with GPT-4o-mini TTS in minutes — free.