MAI Playground, Microsoft AI’s public test bench for its in-house models, gained three new audio models on 1 October 2026. MAI-Transcribe-2-Streaming turns live speech into text as people talk, and two new voice models, MAI-Voice-2.1 and the faster MAI-Voice-2.1-Flash, turn text back into speech in 23 languages. Microsoft also added a demo called Chatter, a voice assistant built from the new models, so anyone can hear the pieces working together.

TestingCatalog spotted the additions the same day, quoting Microsoft’s own summary: “MAI-Transcribe and MAI-Voice models: accurate, fast, low cost, and chart-topping audio understanding and generation for building the best conversational voice agents.” The models are also available to developers through Microsoft Foundry and partners.

This article explains what is new in MAI Playground, what each model does and costs, how the streaming transcription model ranks against rivals, how quickly Microsoft’s voice line has moved this year, and what developers in the UK should check before building on it. For the batch model that came first, see our earlier report on MAI-Transcribe-2.

What Microsoft Added to MAI Playground

mai playground microsoft mai voice transcribe models b stenotype machine feeding out paper tape

Microsoft announced the three models in a post titled “Our first streaming transcription model debuts at no. 1 on Artificial Analysis”, published at 16:00 UTC on 1 October.

Three models in one release

MAI-Transcribe-2-Streaming is Microsoft’s first real-time speech recognition model. MAI-Voice-2.1 is its most capable text-to-speech model, and MAI-Voice-2.1-Flash is the low-latency version built for volume. Microsoft frames them as the two ends of a voice agent: one model listens, the other speaks.

Chatter, the new demo

Chatter is a new voice assistant demo inside MAI Playground, labelled Beta and “Powered by voice & transcribe”. Microsoft says it built Chatter “to show these models working together in a live agent.” It replaces the idea of testing each model separately with a single spoken conversation.

What the MAI Playground menu shows now

When we checked on 2 October, the MAI Playground model picker listed MAI-Transcribe-2, MAI-Voice-2.1, MAI-Image-2.6 and MAI-Thinking-1, with Chatter under AI Experiences. Microsoft’s own model pages for MAI-Transcribe-2 and MAI-Voice-2.1 were updated on 1 October with “Try in playground” buttons.

ModelJobPriceLanguages
MAI-Transcribe-2-StreamingReal-time speech to text$0.54 per audio hour (introductory, to end of 2026)60
MAI-Voice-2.1Expressive text to speech$22 per 1M characters23
MAI-Voice-2.1-FlashLow-latency text to speech$15 per 1M characters23
Chatter (demo)Voice assistant using the models aboveFree in MAI PlaygroundNot stated

MAI-Transcribe-2-Streaming: Real-Time Speech to Text

mai playground microsoft mai voice transcribe models c songbird singing on a flowering branch

The streaming model is the headline. Batch transcription waits for a whole file; streaming transcription has to produce words while someone is still talking, which is much harder to do accurately.

First words in about 100 milliseconds

Microsoft says the model produces its first guesses, known as partials, “in just over 100ms of receiving audio”, then revises them as more context arrives and commits a stable transcript. That lets a voice agent start reasoning or calling tools before the speaker finishes. Microsoft’s internal tests show words appearing “2x faster” than its closest competitor for dictation and subtitles.

Sixty languages with automatic detection

The streaming model covers 60 languages, matching the batch MAI-Transcribe-2, and supports automatic, continuous language detection. That matters for call centres and meetings where people switch languages mid-conversation.

How it ranks on Artificial Analysis

Microsoft cites the Artificial Analysis streaming speech-to-text leaderboard as of 28 September 2026. There, MAI-Transcribe-2-Streaming had the lowest final word error rate of the streaming models shown, 2.50%, and a first-partial error rate of 2.80%. It returned a final transcript 0.130 seconds after speech ended.

Final word error rate on Artificial Analysis streaming leaderboard, 28 Sep 2026, lower is better (bar length = WER ÷ 5.25%)

MAI-Transcribe-2-Streaming: 2.50%
Grok Voice Transcribe 2.0 (Streaming): 2.73%
Muse Voice Transcribe: 3.06%
ElevenLabs Scribe v2 Realtime: 3.59%
GPT Live Transcribe: 3.92%
Gemini 3.5 Transcribe Live: 4.00%
Azure STT Real-time Transcription: 5.25%

Microsoft beats its own older service

One detail stands out in Microsoft’s chart: its existing Azure real-time speech service sits at 5.25%, more than double the new model’s 2.50%. Teams already using Azure Speech for live captions have a clear reason to test the new model.

Speed is not uniform

Low error is only half of a streaming result. On the same chart, some rivals return a final transcript faster: Soniox v5 Real-Time at 0.054 seconds and Cartesia Ink-2 at 0.067 seconds, against Microsoft’s 0.130. Microsoft’s claim is that it sits on the best balance of accuracy and speed, not that it is the fastest.

Price compared with the batch model

The streaming model costs an introductory $0.54 per hour of audio until the end of the year. The batch MAI-Transcribe-2 launched on 3 September at $0.10 per hour. Real-time transcription therefore costs 5.4 times as much per hour ($0.54 ÷ $0.10), which is normal for streaming but worth budgeting for if you only need recordings transcribed after the fact.

MAI-Voice-2.1 and MAI-Voice-2.1-Flash

mai playground microsoft mai voice transcribe models d long eared bat listening in flight

The two voice models share the same languages and voices but are tuned for different jobs, and both now appear in MAI Playground.

Twenty-three languages, one voice

Microsoft says MAI-Voice-2.1 supports 23 languages and lets “one single voice” speak all of them “with a truly native accent”, so a brand can keep the same voice everywhere. The model page lists English, Italian, French, German, Hindi, Spanish, Portuguese, Korean, Chinese (Simplified), Turkish, Russian, Thai, Dutch, Romanian, Hungarian, Czech, Danish, Finnish, Indonesian, Polish, Swedish, Norwegian and Vietnamese. Japanese and Arabic are not on the list.

Counting the locales

The blog says 23 languages and 26 locales. The model page’s list has 28 entries, because English appears in four variants (US, Australia, UK and India) and Spanish and Portuguese in two each. Microsoft has not explained the difference, so check the exact locale you need in MAI Playground before you commit.

Prices have not moved

MAI-Voice-2.1 costs $22 per million characters and Flash costs $15. When Microsoft launched MAI-Voice-2-Flash in July, it said Flash was 32% cheaper than MAI-Voice-2 at $15, which implies MAI-Voice-2 was about $22 ($15 ÷ 0.68). So the 2.1 upgrade adds eight languages at the same price.

Latency figures

Microsoft’s model page lists model inference latency of about 550 milliseconds for MAI-Voice-2.1 and about 45 milliseconds for Flash. The blog gives Flash “an end-to-end latency” of 150 milliseconds and says it delivers “55% faster model inference” and is “~60% cheaper than comparable models.” Inference time and end-to-end time measure different things, so expect real calls to land nearer the larger figure.

FeatureMAI-Voice-2.1MAI-Voice-2.1-Flash
Model inference latencyAbout 550 msAbout 45 ms
Price per 1M characters$22$15
Languages2323
Emotion control and voice promptingYesYes
Best for (Microsoft)Audiobooks, content, voice-overCall centres, assistants, IVR

What listeners thought

Microsoft reports two listener tests. In a blind side-by-side comparison of 5,032 judgements, MAI-Voice-2.1 was preferred 58.9% of the time against 41.1% for MAI-Voice-2. In a test with 4,000 listeners, 50.3% rated the MAI voices as equally or more human-like than human recordings. Both are Microsoft’s own studies, and the second combines results for both 2.1 models.

Cloning with consent guardrails

Both voice models can clone a voice “using just a few seconds of reference audio” in every supported language, and Microsoft says they have “built-in consent guardrails that prevent misuse.” For the legal side of voice cloning, see our report on a Japanese court ruling on voice rights.

How Fast Microsoft's Voice Line Has Moved

mai playground microsoft mai voice transcribe models e camera flash gun firing its bulb

The additions to MAI Playground cap a year in which Microsoft has shipped a new audio model roughly every month. Microsoft AI, led by Mustafa Suleyman, has steadily reduced its reliance on partner models for speech.

DateReleaseKey fact
2 April 2026MAI-Transcribe-125 languages
4 June 2026MAI-Voice-2English-only to 15 languages
24 July 2026MAI-Voice-2-Flash public preview$15 per 1M characters, 2x faster
3 September 2026MAI-Transcribe-260 languages, $0.10 per hour
1 October 2026MAI-Transcribe-2-Streaming, MAI-Voice-2.1, MAI-Voice-2.1-FlashReal-time, 23 voice languages, added to MAI Playground

Languages supported, by release (bar length = languages ÷ 60)

MAI-Voice-2 (June): 15
MAI-Voice-2.1 (October): 23
MAI-Transcribe-1 (April): 25
MAI-Transcribe-2 and Streaming (Sept/Oct): 60

On these figures, voice coverage grew by about half in four months (from 15 to 23 languages), while transcription coverage more than doubled in five (from 25 to 60).

Why the pace matters

Each release has gone into Microsoft’s own products before or alongside developers. In July Microsoft said a MAI transcription model was replacing the previous model in Dragon Copilot, its clinical documentation tool used by 170,000 medical providers. Our report on Microsoft’s draft code of conduct covers the principles it says govern these releases.

Why Microsoft Is Building Its Own Voice Models

mai playground microsoft mai voice transcribe models f wind up chattering teeth on two shoes

The steady stream of releases into MAI Playground is part of a bigger strategy. Microsoft has long relied on partner models, above all OpenAI’s, inside Copilot. Its MAI line gives it in-house alternatives that it controls, prices and ships on its own schedule.

Products first, developers second

Each MAI audio model has gone into Microsoft’s own products as well as to developers. Microsoft said earlier MAI-Transcribe versions were used in Copilot Voice and Teams, and that MAI-Voice-2 was being built into VS Code and the Dynamics 365 Contact Center. When a model performs in those products, it then appears in MAI Playground and Foundry for everyone else.

Price as a weapon

Microsoft is competing hard on price. MAI-Transcribe-2 launched at $0.10 per audio hour, and the streaming model at an introductory $0.54. Both introductory prices run only until the end of 2026, so budget for a possible increase next year rather than assuming today’s rate will last.

What it means for Azure Speech customers

Azure’s existing real-time speech service scored 5.25% on the same streaming leaderboard, against 2.50% for the new model. Microsoft has not said whether the older service will be retired, but customers using it for captions or call transcription now have a clearly better option from the same vendor. Testing it in MAI Playground first costs nothing.

A small lab moving quickly

Microsoft describes its MAI team as “a lean, fast-moving lab” with compute that “is ramping quickly and extensively.” Five audio releases in six months suggest that pace will continue, so treat any comparison you run in MAI Playground today as a snapshot rather than a final verdict.

Where to Use the Models Beyond MAI Playground

MAI Playground is for listening and testing. For production, Microsoft lists several routes.

Microsoft Foundry and Azure Voice Live

All three models are available through Microsoft Foundry, and Microsoft points voice-agent builders to Azure Voice Live, its speech-to-speech service. The MAI-Voice-2.1 page also links to Azure Speech and to Copilot Audio Expressions.

OpenRouter, Vercel and LiveKit

MAI-Voice-2.1 and Flash are available through OpenRouter. Microsoft also lists Vercel, and says LiveKit support is coming soon. That makes it easier to swap MAI voices into an existing agent stack without moving everything to Azure.

Building a voice agent loop

Microsoft’s own advice is to pair MAI-Transcribe-2-Streaming with MAI-Voice-2.1-Flash. “A voice agent is a loop. It has to hear, understand, decide, and speak,” its post says, and every component “either buys you time in that window… or spends it.” Saving time on listening and speaking leaves more time for the agent to reason and use tools.

Use cases Microsoft names

Microsoft lists customer service agents that act before the caller finishes, multilingual assistants that reply in the language they are addressed in, and interactive learning with distinct speakers for tutoring and role-play. Each depends on natural language processing between the two audio models, typically a large model that decides what to say.

How to Test the New Models in MAI Playground

MAI Playground is free to use, so the quickest way to judge the new models is to try them on your own material. Here is a simple plan.

Start with Chatter

Open Chatter in MAI Playground and hold a normal conversation. Notice how quickly it replies, whether it waits for you to finish, and what happens when you interrupt it mid-sentence. That tells you more about the streaming transcription and Flash voice working together than any benchmark chart.

Test transcription on your own audio

Record two or three minutes that look like your real use: a phone call, a meeting, a support conversation. Include regional accents, technical terms and some background noise. Run it through MAI Playground and count the errors by hand. If your current provider can process the same clip, compare the two results line by line.

Compare the two voices

Give MAI-Voice-2.1 and MAI-Voice-2.1-Flash the same script and listen to both. Try a sentence with a product name, a long number and a question, then ask the same voice to repeat it in a second language. The difference in pacing and expression will show whether you need the full model or whether Flash is good enough.

Know what MAI Playground cannot tell you

MAI Playground shows quality, not operations. It will not tell you about service levels, data retention, regional hosting or costs at volume. Those live in Microsoft Foundry and Azure documentation and in your contract, so treat a good MAI Playground session as the start of an evaluation, not the end of one.

How the New Models Compare

MAI Playground makes it easy to hear Microsoft’s models, but buyers will want to compare them with rivals they may already use.

Streaming transcription rivals

On the Artificial Analysis chart Microsoft published, the nearest rivals on accuracy were xAI’s Grok Voice Transcribe 2.0 and Meta’s Muse Voice Transcribe, followed by Cartesia, ElevenLabs Scribe v2 Realtime, OpenAI’s GPT Live Transcribe and Google’s Gemini 3.5 Transcribe Live. The gap between first place and ElevenLabs in seventh is 1.09 percentage points (3.59% minus 2.50%), so test on your own audio. Our report on Gemini 3.5 Transcribe covers Google’s batch model.

Voice rivals

ElevenLabs launched Eleven v4 and v4 Turbo on 28 September and topped the TTS Arena leaderboard, as we reported in our story on ElevenLabs Eleven v4. Microsoft’s pitch against that is price and language breadth per voice, rather than a leaderboard win.

Streaming modelFinal WERFirst partial WERTime to final
MAI-Transcribe-2-Streaming2.50%2.80%0.130 s
Grok Voice Transcribe 2.02.73%3.36%0.490 s
Muse Voice Transcribe3.06%3.57%0.163 s
ElevenLabs Scribe v2 Realtime3.59%3.59%0.141 s
Gemini 3.5 Transcribe Live4.00%5.77%0.395 s

Figures are from the Artificial Analysis data Microsoft embedded in its 1 October post, dated 28 September 2026.

What It Means for UK Developers

The new models are easy to try in MAI Playground, but production use raises the usual questions.

Voice data is personal data

Recordings of people’s voices are personal data under UK GDPR, and a voice used to identify someone can be special category biometric data. If you transcribe calls, tell callers, set retention periods and check where processing happens in your Azure or Foundry region.

Cloning needs clear consent

Cloning a staff member’s or actor’s voice needs written consent that covers how long and where the voice will be used. Microsoft’s guardrails help, but the responsibility stays with you. For the security side of call-centre voice, our IT security team can help you plan against cloned-voice fraud.

Test with your own accents

Benchmarks use standard test sets. Before switching a UK contact centre, run your own recordings through MAI Playground, including regional accents and background noise, and compare the error rate with your current provider. Our AI and machine learning team can help design that comparison.

MAI Playground FAQ

What is MAI Playground?

MAI Playground is Microsoft AI’s free web space for trying Microsoft’s in-house models, including transcription, voice, image and reasoning models, plus demos such as Chatter.

What was added to MAI Playground on 1 October 2026?

MAI-Transcribe-2-Streaming, MAI-Voice-2.1 and MAI-Voice-2.1-Flash, together with Chatter, a voice assistant demo built on the new models.

How much do the new models cost?

MAI-Transcribe-2-Streaming costs an introductory $0.54 per audio hour until the end of 2026. MAI-Voice-2.1 costs $22 and MAI-Voice-2.1-Flash $15 per million characters.

How many languages do they support?

The streaming transcription model supports 60 languages. Both voice models support 23 languages, with one voice able to speak all of them.

Is MAI-Transcribe-2-Streaming the most accurate?

On Artificial Analysis’s streaming leaderboard dated 28 September 2026, as published by Microsoft, it had the lowest final word error rate of the models shown, at 2.50%.

Where can developers use the models?

Through Microsoft Foundry, Azure Speech and Voice Live, OpenRouter (voice models), Vercel and, soon, LiveKit, as well as for testing in MAI Playground.

References and Further Reading