Microsoft MAI-Voice-2.1 and MAI-Transcribe-2: Voice AI Models Tested
Hands-on look at Microsoft's MAI-Voice-2.1 TTS and MAI-Transcribe-2 speech-to-text, covering languages, latency, accuracy and refusals.

What are MAI-Voice-2.1 and MAI-Transcribe-2?
MAI-Voice-2.1 and MAI-Transcribe-2 are Microsoft AI’s latest pair of voice models, built to make AI voice agents sound and respond like a real person on a call. MAI-Transcribe-2 is the speech-to-text side, converting spoken audio into text in real time across 60 languages with partial transcripts appearing in just over 100 milliseconds. MAI-Voice-2.1 is the text-to-speech side, an expressive model covering 23 languages that keeps a single voice’s identity and native accent consistent across all of them. Together they’re designed to form a fast hear-think-speak loop for conversational agents.
TL;DR
- MAI-Transcribe-2 streams speech to text across 60 languages, delivers first partial transcripts in about 100 milliseconds, and posts a word error rate of roughly 2.5% on the Artificial Analysis accuracy leaderboard.
- MAI-Voice-2.1 is Microsoft’s most expressive TTS model, supporting 23 languages while preserving one voice’s identity and native accent in each language.
- A flash variant of Voice 2.1 trades some expressiveness for speed, hitting around 150 milliseconds end-to-end latency at $15 per million characters, which the reviewer flagged as pricey.
- In hands-on testing across languages including Spanish, Arabic, Brazilian Portuguese, Hindi, Urdu, Czech, Chinese, Danish, Italian, Dutch, French and Turkish, transcription speed and accuracy looked consistently strong.
- The voice model handled conversational interruptions smoothly during a live demo, cutting in and responding naturally rather than waiting for a full pause.
- Refusal behavior is aggressive: the demo assistant (“Harper”) declined even mildly odd or off-topic requests, repeatedly steering back to a narrow set of safe topics.
- Microsoft has not released open weights for either model, so testing is limited to the official playground rather than local deployment.
How does MAI-Transcribe-2 perform across languages?
MAI-Transcribe-2 is a streaming speech-to-text model covering 60 languages, with two output modes: a verbatim transcript and a “clean” version that smooths out disfluencies. In testing through Microsoft’s playground, the model was run against sample audio in Russian, Spanish, Arabic, Brazilian Portuguese, Hindi, and Urdu.
Most languages transcribed quickly and accurately, with Spanish, Arabic and Brazilian Portuguese in particular standing out for speed. The one exception was Urdu, which took noticeably longer to process than the others. Overall, the result lines up with Microsoft’s claim of topping the Artificial Analysis accuracy leaderboard with about a 2.5% word error rate, a low figure for a model handling dozens of languages rather than just one or two.
The headline technical claim is latency: first partial transcripts appear in just over 100 milliseconds, which matters for live captioning, voice agents, and call center tooling where a laggy transcript breaks the sense of a real conversation.
How expressive is MAI-Voice-2.1 in practice?
MAI-Voice-2.1 is pitched as Microsoft’s most expressive text-to-speech model, and it supports 23 languages while keeping a single voice’s identity and accent consistent no matter which language it’s speaking. In the playground, voices can be tuned by tone, with presets like narrator, educational, customer call center and agent.
Testing moved through Czech, Chinese, Danish, Italian, Indonesian, Dutch and French, including a deliberately tricky Czech tongue twister built around consonant clusters, a common pronunciation stress test for TTS systems. The output held up well across most languages tested, with pronunciation sounding natural rather than robotic, even on harder phonetic cases.
Alongside the full Voice 2.1 model, Microsoft offers a flash version aimed at speed-sensitive use cases. It runs at around 150 milliseconds end-to-end latency, fast enough for real-time back-and-forth conversation, but it comes at $15 per million characters, a price point that stands out as high compared to typical TTS pricing.
How well does it handle real conversation, including interruptions?
One of the more telling tests wasn’t about raw accuracy but about conversational mechanics: what happens when a user interrupts the AI mid-sentence. In a live demo built around a voice agent named “Harper,” the model handled interruptions cleanly, stopping what it was saying and picking up the new input without the awkward lag or garbled overlap that trips up many voice bots.
That matters more than it sounds. A model that nails pronunciation but can’t handle a user talking over it still feels mechanical on a call. Combining fast transcription (around 100ms to first partial) with a quick-responding voice model is what lets an agent feel like it’s listening in real time rather than processing in a rigid turn-based loop. Microsoft’s pitch is that this combination frees up more time inside the response loop for the underlying model to reason or call tools while still sounding natural, since the voice and transcription layers aren’t eating up the whole latency budget.
How aggressive is the refusal behavior?
Remy doesn't build the plumbing. It inherits it.
Other agents wire up auth, databases, models, and integrations from scratch every time you ask them to build something.
Remy ships with all of it from MindStudio — so every cycle goes into the app you actually want.
The same demo that showed off smooth interruption handling also surfaced a sharper edge: Harper refused almost any request that strayed from a narrow set of safe topics. Mildly odd, non-harmful, or just unconventional prompts were met with a flat “I’m sorry, I can’t help with that,” followed by a redirect toward gardening or similarly safe small talk.
This isn’t a transcription or voice-quality issue, it’s a guardrail design choice, but it’s relevant to anyone evaluating these models for a real product. An assistant that refuses too readily risks frustrating users on completely benign requests, not just genuinely problematic ones. For builders, this means the refusal threshold and system prompt design will matter as much as the underlying voice quality when deploying an agent built on these models.
Is MAI-Voice-2.1 / MAI-Transcribe-2 worth testing for your project?
For teams building voice agents, the appeal is straightforward: low latency on both ends (around 100ms for transcription, around 150ms for flash TTS), broad language coverage (60 languages for transcription, 23 for speech), and a transcription accuracy figure that leads at least one public leaderboard. That’s a strong combination on paper for call centers, multilingual support bots, or any product where conversational speed is the differentiator.
The catches are pricing and openness. The flash TTS tier at $15 per million characters isn’t cheap at scale, and there’s no open-weight release to test or self-host, unlike some competing voice models that ship weights for local or fine-tuned deployment. Anyone evaluating these models today is doing so through Microsoft’s hosted playground and APIs rather than their own infrastructure. For teams already in the Microsoft AI ecosystem, that’s a minor friction. For teams wanting to own their deployment, it’s a real limitation worth weighing against the latency and accuracy gains.
Frequently Asked Questions
What languages does MAI-Transcribe-2 support?
MAI-Transcribe-2 supports streaming transcription across 60 languages, with both verbatim and cleaned-up transcript output modes.
What languages does MAI-Voice-2.1 support?
MAI-Voice-2.1 covers 23 languages and keeps a single voice’s identity and native accent consistent when speaking in any of them.
How fast are these models?
MAI-Transcribe-2 returns first partial transcripts in just over 100 milliseconds. The flash version of MAI-Voice-2.1 generates audio with around 150 milliseconds of end-to-end latency.
How accurate is MAI-Transcribe-2?
Microsoft reports a word error rate of about 2.5% on the Artificial Analysis accuracy leaderboard, which it currently tops.
Are MAI-Voice-2.1 and MAI-Transcribe-2 open source?
No. Both are closed models accessible through Microsoft’s playground and APIs, with no open-weight release available for local deployment.
Why does the demo assistant refuse so many requests?
The demo voice agent applies strict guardrails, declining even harmless off-topic requests and redirecting conversation to a limited set of safe subjects. This reflects a deliberate safety design choice rather than a limitation of the voice or transcription models themselves.