MAI-Transcribe-2 has just landed on AIxploria, ten and a half days after Microsoft AI released it. The directory published the card at 04:12 UTC on 14 September 2026 with a 4.4-star rating, a “Free” pricing tag, a “Verified Tool” label, a badge reading “#25 in Transcriber” and a full editorial review. We checked it the way we check every new listing in our series on AI tools that reach the directory.
We compared the MAI-Transcribe-2 card with Microsoft AI’s launch post, its model page, the Azure Speech documentation, the REST API reference and the Azure pricing page. We tested the badge against AIxploria’s own Transcriber category and read the live Artificial Analysis leaderboard that Microsoft cites. Sixteen of the 30 items we could check match. The problems are sharper than usual: the #25 slot belongs to another card, the Free tag contradicts the card’s own FAQ, and the review gets the diarization defaults and a file limit wrong.
Google’s rival reached the same directory 18 days earlier, which we covered in Gemini 3.5 Transcribe Is Now Available on AIxploria. This article follows our Google Pics review, published earlier today, and it uses the same method: vendor pages first, then the badge, then the claims, then the lists.
Table of contents
- What MAI-Transcribe-2 Is, in Microsoft’s Own Words
- What the AIxploria Card for MAI-Transcribe-2 Says
- Testing the MAI-Transcribe-2 Card’s #25 Badge
- Fact-Checking the MAI-Transcribe-2 Review, Claim by Claim
- The Free Tag on a $0.10-an-Hour Model
- Is MAI-Transcribe-2 the Fastest, Most Accurate and Cheapest?
- What the MAI-Transcribe-2 Card Leaves Out for Businesses
- The Alternatives and Featured Blocks Around the Card
- Where the MAI-Transcribe-2 Card Has Propagated
- How MAI-Transcribe-2 Compares With Earlier Cards
- What MAI-Transcribe-2 Means for Your Business
- How to Check a Directory Card Before Acting on It
- Frequently Asked Questions
- References and Further Reading
What MAI-Transcribe-2 Is, in Microsoft's Own Words
Microsoft AI, the company’s in-house model lab, announced MAI-Transcribe-2 at 14:00 UTC on 3 September 2026 under the headline “MAI-Transcribe-2 is the fastest, most accurate and cheapest speech recognition model in the world”. It is a batch speech-to-text model: you send a recorded audio file and receive text. Every check in this article runs against Microsoft’s own pages and the benchmark data Microsoft cites, not against press coverage.
A preview model inside Azure Speech
The Azure documentation calls the product “MAI-Transcribe in Azure Speech (preview)”. It says the feature is in public preview, “provided without a service-level agreement, and is not recommended for production workloads”. Developers call MAI-Transcribe-2 through the Fast Transcription API by setting enhancedMode.enabled to true and naming the model. The same model can transcribe input audio in the Voice Live API, and Microsoft also offers it in the MAI Playground and through OpenRouter.
The feature list Microsoft published
Microsoft’s launch post lists nine capabilities for MAI-Transcribe-2:
- Faster inference, “up to 10× faster processing than leading competitors”, especially on long audio
- Speaker diarization that attributes words to the right person
- Word-level timestamps for alignment, search and editing
- Keyword biasing for domain terms, abbreviations and names
- A “verbatim” style that keeps filler words and a “clean” style that removes them
- Code switching for blended pairs such as Hinglish and Spanglish
- Automatic language identification
- Robust performance in noisy conditions
- Accuracy across 60 languages
Three versions in 154 days
MAI-Transcribe-2 is the third model in the family. Microsoft released MAI-Transcribe-1 on 2 April 2026 and MAI-Transcribe-1.5 on 2 June. The documentation now marks the first version “Deprecated on Aug 20, 2026”, 140 days after its launch.
| Version | Released | Languages | AA-WER on Artificial Analysis | Diarization and word timestamps | Price per audio hour |
|---|---|---|---|---|---|
| MAI-Transcribe-1 | 2 April 2026 | 25 | 2.61% | No | $0.36 |
| MAI-Transcribe-1.5 | 2 June 2026 | 43 | 2.38% | No | $0.36 |
| MAI-Transcribe-2 | 3 September 2026 | 60 | 2.04% | Yes | $0.10 until 31 December 2026 |
Seventeen new languages
The language table in Microsoft’s documentation has one column for MAI-Transcribe-1.5 and one for MAI-Transcribe-2. All 43 languages of the older model carry over, and 17 appear only in the new column: Afrikaans, Armenian, Azerbaijani, Bosnian, Cantonese, Filipino, Galician, Hebrew, Icelandic, Kazakh, Latvian, Macedonian, Malay, Nepali, Persian, Swahili and Urdu. That makes exactly 60.
Speech is only one strand of Microsoft AI’s voice work. We reported last month on MAI Realtime, a bidirectional voice model that surfaced in the MAI Playground before Microsoft confirmed anything about it.
What the AIxploria Card for MAI-Transcribe-2 Says
The card opens with a two-sentence blurb: “A transcription model that converts audio to text, even in challenging acoustic conditions, with diarization and support for over 60 languages. Add your industry-specific terms to improve recognition of specialized vocabulary.” Underneath sits a full review: a summary, five pros and three cons, sections on features, rivals and pricing, a version table, a four-question FAQ and a verdict.
The card at a glance
| Field on the card | Value read on 14 September 2026 |
|---|---|
| Card address | aixploria.com/en/mai-transcribe-2-microsoft/ (AIxploria post 51364) |
| Published | 04:12:07 UTC, edited at 04:16:24 UTC |
| Rating | 4.4 out of 5 from 5 votes |
| Upvotes and views | 54 upvotes and 6,002 views at 11:34 UTC |
| Pricing tag and label | Free, Verified Tool |
| Tags | Latest AI, Transcriber |
| Rank badge | #25 in Transcriber |
| Useful links | Official Website, Playground |
| Alternatives | 24 in the page title, 8 shown on the page |
Five votes behind 4.4 stars
The structured data behind the rating reports a rating count of five. Five votes on a card that is seven hours old tell you nothing about MAI-Transcribe-2. The control cards we opened in the same category are no better: Whisper Large V3 Turbo shows a perfect 5 from a single vote, and ElevenLabs Scribe V2 shows 4.6 from five.
Ten and a half days behind Microsoft
Microsoft’s launch post went live at 14:00 UTC on 3 September, and the card went live at 04:12 UTC on 14 September, 10 days and 14 hours later. Google’s rival moved through the same directory much faster. Gemini 3.5 Transcribe was announced on 26 August and had its card by 03:33 UTC the next morning, at most 27 and a half hours later.
Testing the MAI-Transcribe-2 Card's #25 Badge
The badge above the review reads “#25 in Transcriber” and links to AIxploria’s Transcriber category, which covers text, audio, image and video transcription. In earlier articles, badges have named a category that did not contain the tool, pointed at a predecessor and placed a successor below its older model. So we read the category pages and compared cards at known positions with their own badges.
Six controls, six exact matches
Pages one to three of the category loaded normally and held 36 tools. We opened six cards at known positions, and every badge matched its slot exactly.
| Tool | Position on the category pages | Badge on its own card |
|---|---|---|
| GLM-OCR | 3 | #3 in Transcriber |
| ElevenLabs Scribe V2 | 4 | #4 in Transcriber |
| Gemini 3.5 Transcribe | 8 | #8 in Transcriber |
| FineVoice Speech to Text | 11 | #11 in Transcriber |
| Whisper OpenAI | 21 | #21 in Transcriber |
| Whisper Large V3 Turbo | 25 | #25 in Transcriber |
| MAI-Transcribe-2 | Not in the first 36 | #25 in Transcriber |
The badges are a working ordinal, which is exactly what makes the result on this card hard to explain away.
Two cards, one slot
Position 25, the first card on page three, is Whisper Large V3 Turbo, a card AIxploria published in October 2024. Its own badge reads “#25 in Transcriber”. The MAI-Transcribe-2 card carries the identical badge, yet the new model appears nowhere in the first 36 places. Either the new card’s badge was computed against a list that has not been refreshed, or two tools share one rank. A reader who follows the badge link finds OpenAI’s open-weights model in the slot, not Microsoft’s.
The Whisper link points at the same card
The review compares MAI-Transcribe-2 with “Whisper-Large-V3”, the model Microsoft names in its benchmark, and links that name to AIxploria’s Whisper Large V3 Turbo card. They are different models: Turbo is OpenAI’s faster variant of Large V3 with a much smaller decoder. AIxploria has no separate Large V3 card, and its whisper-large-v3 address redirects to the Turbo card. The one comparison link in the review lands on the tool that holds the badge slot.
About 100 tools, and the limit of this check
The category’s page title reads “Page 2 of 9”, and each page holds 12 cards, so Transcriber lists between 97 and 108 tools. Page four returned a 5,650-byte Cloudflare challenge, the same block every card in this series has hit. We cannot say where MAI-Transcribe-2 really sits. The verified claim is narrower: it is absent from the first 36 places, and its badge names a slot another card holds.
Fact-Checking the MAI-Transcribe-2 Review, Claim by Claim
We split the review, the blurb and the card’s metadata into 30 checkable items. We compared each one with Microsoft’s pages, the REST reference, the Artificial Analysis data and, for one claim about Whisper, OpenAI’s own code. Sixteen items match. Two are rounded generously, two are incomplete, two go further than their source, two appear nowhere in Microsoft’s pages, four are contradicted by a primary source, and two are badge or link errors.
Where the MAI-Transcribe-2 review matches Microsoft
The core description is accurate. The card correctly gives the 3 September public-preview release, the $0.10-per-hour launch price through the end of 2026, first place on FLEURS with a 5.2% average word error rate across 60 languages, keyword biasing, the verbatim and clean styles, automatic language identification, code switching and the WAV, MP3 and FLAC formats. Its version table and its rival prices are right too. Note that “topped FLEURS” is Microsoft’s own evaluation, not an independent test.
“Included by default” is the wrong way round
The pros list says “Speaker diarization and word-level timestamps included by default”. Microsoft’s parameter table says the opposite. The diarization.enabled setting defaults to false and timestamps defaults to “none”, so a plain request to MAI-Transcribe-2 returns neither. The card’s own FAQ later says both are “switched on through API parameters”, which is correct and contradicts the pros list above it.
The 15-minute ceiling the card leaves out
The review’s headline example is “a four-person steering call” that comes back labelled by speaker. Its FAQ adds only that diarization “supports shorter recordings than plain transcription does”. Microsoft’s documentation is far more specific. In preview, requests with diarization switched on “fail for recordings of about 15 minutes and longer”, returning HTTP 408, 500 or 503 errors. An hour-long call is four times that length. Microsoft’s workaround is to transcribe without diarization, keep the word-level timestamps and run a separate speaker diarization step.
A file limit Microsoft does not publish
The cons list says files are “capped at 300 MB or two hours”. The MAI-Transcribe documentation sends readers to the REST reference for limits, and that reference, for API version 2025-10-15, says the audio “must be shorter than 2 hours in audio duration and smaller than 250 MB in size”. The two-hour figure is right, and the 300 MB figure appears on neither page. The general Fast Transcription guide gives different figures again, under five hours and 500 MB, so test the real limit before building a pipeline around it.
Two claims about Whisper
One FAQ answer says MAI-Transcribe-2 “ships with diarization and timestamps built in, features Whisper does not provide on its own”. The diarization half is fair. The timestamp half is not: OpenAI’s open-source Whisper returns start and end times for every segment, and its transcribe function includes a word_timestamps option for word-level timing. Another answer says MAI-Transcribe-1 “was deprecated five months after release”. The documented gap is 140 days, closer to four and a half months.
Speed and price claims that outrun the data
The review says Artificial Analysis “clocked it at up to ten times the speed of rivals on long recordings”. Microsoft attributes 10x, 7x and 5x ratios to Artificial Analysis evaluations, but the medians Artificial Analysis publishes today give smaller gaps, as the speed section below shows. The verdict’s “roughly a third of rival rates” is accurate against Gemini 3.5 Transcribe, $0.10 against $0.30, but not against ElevenLabs Scribe v2, where $0.10 is 45% of $0.22.
Two claims Microsoft’s pages do not contain
The review says the model is “paired with its MAI-Voice-2 Flash speech model on the output side”. Microsoft’s MAI-Transcribe-2 post, model page and documentation never mention MAI-Voice-2 Flash; the pairing idea dates from April, when Microsoft suggested combining MAI-Transcribe-1 with MAI-Voice-1. The review also says Microsoft is “step by step reducing its reliance on OpenAI’s models”. That may be a fair reading of the strategy, but no Microsoft page we read says it.
The full MAI-Transcribe-2 ledger
| # | Item on the card | What the sources say | Verdict |
|---|---|---|---|
| 1 | Microsoft AI’s in-house speech-to-text model | “Built in-house by the Microsoft AI team” | Matches |
| 2 | Converts audio across 60 languages | 60 languages in the documentation | Matches |
| 3 | Blurb: “over 60 languages” | Exactly 60 | Stronger than the source |
| 4 | Released 3 September 2026 in public preview | Launch post dated 3 September; documentation says public preview | Matches |
| 5 | $0.10 per audio hour until the end of 2026 | “A limited-time offer until the end of the year” | Matches |
| 6 | First on FLEURS, 5.2% average WER over 60 languages | Same figures, from Microsoft’s own evaluation | Matches |
| 7 | Ahead of Whisper-Large-V3, GPT-Transcribe, Scribe v2 and Gemini 3.5 Transcribe | The same four models named | Matches |
| 8 | Diarization and word timestamps included by default | Both default to off | Contradicted by the documentation |
| 9 | Up to 10x faster on long recordings | “Up to 10× faster processing”, especially long-form audio | Matches |
| 10 | Artificial Analysis clocked up to 10x the speed of rivals | Published medians give 7.7x, 5.1x and 3.3x | Stronger than the source |
| 11 | Keyword biasing for jargon and names | A phrase list parameter | Matches |
| 12 | Verbatim and clean styles | Both documented, verbatim by default | Matches |
| 13 | Automatic language detection and code switching | Both documented, including Hinglish and Spanglish | Matches |
| 14 | Paired with MAI-Voice-2 Flash | Not mentioned on the three launch pages | Not in Microsoft’s pages |
| 15 | Public preview with no service-level agreement | Same wording | Matches |
| 16 | No ready-made consumer app | Earlier versions already run in Copilot, Teams and Dragon Copilot | Incomplete |
| 17 | Files capped at 300 MB or two hours | Under 2 hours and under 250 MB | Contradicted by the documentation |
| 18 | WAV, MP3 and FLAC files | The same three formats | Matches |
| 19 | Diarization supports shorter recordings | Fails at about 15 minutes, with a two-step workaround | Incomplete |
| 20 | Version table: 25, 43 and 60 languages in April, June and September | Matches Microsoft’s three launch posts | Matches |
| 21 | A 72% cut from the $0.36 launch price | $0.36 to $0.10 is a 72.2% cut | Matches |
| 22 | Scribe v2 near $0.22, GPT-Transcribe $0.27, Gemini 3.5 Transcribe $0.30 an hour | The same prices on Artificial Analysis | Matches |
| 23 | Access through Foundry, MAI Playground and OpenRouter | All three live; OpenRouter lists $0.10 an hour | Matches |
| 24 | Whisper has no timestamps on its own | Whisper returns segment times and offers word timestamps | Contradicted by Whisper’s code |
| 25 | MAI-Transcribe-1 deprecated five months after release | 140 days | Rounded generously |
| 26 | Labelled transcripts for roughly a third of rival rates | 33% of Gemini’s price, 45% of Scribe v2’s | Rounded generously |
| 27 | Reducing reliance on OpenAI’s models | Not stated by Microsoft | Not in Microsoft’s pages |
| 28 | “Free” pricing tag | The card’s own FAQ answers “No” | Contradicted by the card itself |
| 29 | “#25 in Transcriber” badge | Slot 25 is Whisper Large V3 Turbo | Badge or link error |
| 30 | “Whisper-Large-V3” link | Opens the Whisper Large V3 Turbo card | Badge or link error |
Grouping the 30 verdicts shows where the MAI-Transcribe-2 card is strong and where it is weak.
The Free Tag on a $0.10-an-Hour Model
Free in the metadata, “No” in the FAQ
The card’s pricing tag reads “Free”, and its structured data publishes an offer in the category “Free” at a price of 0. The first question in the card’s own FAQ is “Is MAI-Transcribe-2 free?”, and the answer begins “No, it costs $0.10 per hour of audio”. Microsoft agrees with the FAQ. The MAI Playground lets you try the model without code and an Azure account costs nothing to create, but API calls are billed per audio hour.
What $0.10 an hour buys
At the launch rate, 1,000 hours of recorded audio cost $100. The same volume costs $360 at the $0.36 rate Microsoft still lists for MAI-Transcribe-1.5, and between $220 and $300 at the Artificial Analysis prices for the three rivals the card names. For a contact centre recording 10,000 hours a year, the gap between the old and new Microsoft rates is $2,600.
The price after 31 December
The $0.10 rate is a promotion. Microsoft’s launch post calls it “a limited-time offer until the end of the year”, and the Azure Speech pricing page notes that “MAI-Transcribe-2 offered at a discounted price through December 31st 2026”. Neither page publishes the rate that follows, and Azure’s public retail price API returned no meter with “Transcribe” in its name when we queried it on 14 September. The promotion lasts 119 days from launch, 108 of them after the card appeared.
Is MAI-Transcribe-2 the Fastest, Most Accurate and Cheapest?
Microsoft’s headline makes three superlatives. Its own body text is more careful: the model “ranks first on the FLEURS benchmark” and “ranks second on the Artificial Analysis Word-Error-Rate leaderboard”. We read the live Artificial Analysis non-streaming speech-to-text leaderboard on 14 September to test all three claims.
Accuracy: first on Microsoft’s test, second on the independent one
Artificial Analysis lists MAI-Transcribe-2 at an AA-WER of 2.0%, second behind Fun-Realtime-ASR-preview at 1.7% and ahead of ElevenLabs Scribe v2 at 2.2%. Microsoft’s first place comes from FLEURS, a public multilingual benchmark that Microsoft ran itself. Both results are real. “Most accurate in the world” holds on one leaderboard and not on the other.
Speed: fourth on the median
Artificial Analysis measures speed as a speed factor, the seconds of input audio transcribed per second. Its own summary names Deepgram’s Nova-3 the fastest, at 605.1 times real time. The median for MAI-Transcribe-2 is 285.8 times, fourth among the 33 default provider entries in the page’s data, and faster than every model Microsoft compares it with.
| Model on Artificial Analysis | AA-WER | Median speed factor | One hour of audio takes | Price per audio hour |
|---|---|---|---|---|
| MAI-Transcribe-2 | 2.04% | 285.8x | 12.6 seconds | $0.10 |
| ElevenLabs Scribe v2 | 2.18% | 56.2x | 64.1 seconds | $0.22 |
| MAI-Transcribe-1.5 | 2.38% | 194.1x | 18.5 seconds | $0.36 |
| Gemini 3.5 Transcribe | 2.60% | 87.8x | 41.0 seconds | $0.30 |
| GPT Transcribe | 3.31% | 37.0x | 97.4 seconds | $0.27 |
| Deepgram Nova-3 | 5.18% | 605.1x | 5.9 seconds | $0.26 |
Ten times faster, or 7.7?
Microsoft writes that “based on evals run by Artificial Analysis, the model is 10x faster than OpenAI’s GPT-Transcribe, 7x faster than ElevenLabs’ Scribe v2, and 5x faster than Gemini 3.5 Transcribe”. Dividing the published medians gives 7.7x, 5.1x and 3.3x. Microsoft’s ratios may come from long files or an earlier run, and its post does not say which. The model page’s “1hr audio → 10 sec” sits closer to the fastest 5% of runs, 9.7 seconds at the 95th-percentile speed, than to the 12.6-second median.
The same gap between blog and data applies to the previous version. Microsoft’s June post said MAI-Transcribe-1.5 handles an hour of audio “in under 15 seconds”, its model page now says 20 seconds, and the Artificial Analysis median works out at 18.5 seconds.
Price: cheap, not the cheapest
At $1.67 per 1,000 minutes, MAI-Transcribe-2 is cheaper than every other entry with an AA-WER below 3%. It is not the cheapest speech-to-text model Artificial Analysis lists. The site’s summary names StepAudio 2.5 ASR at $0.3667 per 1,000 minutes, and MAI-Transcribe-2 ranks eighth on price among the same 33 entries. The honest superlative is “cheapest of the most accurate models”, which is still a strong claim.
What the MAI-Transcribe-2 Card Leaves Out for Businesses
The card writes for a developer deciding whether to try the model. A business deciding whether to depend on it needs a different set of facts, and most of them sit in Microsoft’s documentation rather than on the card.
Microsoft says not to use the preview in production
The card lists “no service-level agreement” as a con and stops there. Microsoft’s documentation goes further: preview features are “not recommended for production workloads”, and “certain features might not be supported or might have constrained capabilities”. For a regulated workload, such as call recording in financial services or clinical dictation, that is a procurement answer rather than a footnote.
MAI-Transcribe-1 lasted 140 days
Microsoft deprecated MAI-Transcribe-1 on 20 August 2026, 140 days after launch, and MAI-Transcribe-1.5 arrived only 93 days before MAI-Transcribe-2. A team that built on the first version in April has already migrated once. Anyone building on MAI-Transcribe-2 should pin the model name, keep a small set of reference recordings with known-good transcripts, and budget time to re-test each release.
Long meetings need a two-step pipeline
The 15-minute diarization ceiling changes the architecture for meetings, interviews and contact-centre calls. Microsoft’s own workaround is to transcribe the whole file without diarization, request word-level timestamps, and assign speakers with a separate diarization step. That adds a second component to build, test and possibly buy. The REST limit of two hours per file also means splitting longer recordings, such as all-day hearings or conference sessions.
The family already runs inside Microsoft products
The card’s verdict says there is “no consumer app here, just an API to wire in”. That is true of the MAI-Transcribe-2 API, but the family is already inside Microsoft’s own software. Microsoft said MAI-Transcribe-1 was rolling out in Copilot’s Voice mode and Teams, and that MAI-Transcribe-1.5 was being integrated into Copilot, Teams, GitHub and Dynamics 365 Contact Centre. In July it added that MAI-Transcribe-1.5 supports Dragon Copilot, “a solution used by 170,000 medical providers”.
Data handling is not on the card
The card says nothing about where audio is processed or how long it is kept, and neither do the three Microsoft pages we used for this review. Call recordings and clinical dictation are personal data in almost every case. Those answers belong in a data protection review, alongside Microsoft’s Azure terms, before the first real recording is sent.
The Alternatives and Featured Blocks Around the Card
Eight visible alternatives out of 24
The page title advertises “24 Alternatives”, and eight are shown under the review: FineVoice Speech to Text, ElevenLabs Scribe V2, Video To Blog, Audio Transcriber AI, DeVoice, Noiz Agent, Peech and GLM-OCR. Those eight occupy positions 3, 4, 6, 7, 9, 11, 12 and 13 in the Transcriber category. The strip is simply the top of the category with five tools skipped: FireFlies, Otter AI, Free Subtitles Generator, Gemini 3.5 Transcribe and Notta AI.
An image tool listed as a speech alternative
GLM-OCR’s own card describes it as a way to “extract text from any image” with optical character recognition. It belongs in a category that covers image transcription, but it cannot transcribe a phone call, and it is offered as the eighth alternative to a speech model.
Three of the four named rivals are missing
The review compares MAI-Transcribe-2 with Whisper-Large-V3, GPT-Transcribe, Scribe v2 and Gemini 3.5 Transcribe. Only Scribe V2 appears in the alternatives strip. Gemini 3.5 Transcribe sits at position 8, inside the range the strip draws from, and is left out anyway.
The same four featured tools as this morning
A featured block beside the review promotes Anyvids, ThumbnailCreator.com, Meshy AI and Shoomble: a video aggregator, a thumbnail maker, a 3D-model generator and an AI-character chat site. It is the same four tools, in the same order, that appeared beside the Google Pics card this morning. None transcribes audio. AIxploria’s pages link to its advertising page, but the card does not say whether any featured slot was bought, and we make no claim that one was.
Where the MAI-Transcribe-2 Card Has Propagated
Zero on the Top 100 and the Ultimate List
AIxploria’s Top 100 and Ultimate List pages are server-rendered, so they can be counted. At 11:34 UTC on 14 September, 7 hours and 22 minutes after publication, “MAI-Transcribe-2” appeared zero times on both. That is normal: every card in this series started at zero.
Six hits on the Free AI page
The Free AI page tells a different story. All six hits for MAI-Transcribe-2 sit inside a block headed “List of Tools verified by AIxploria team”, next to DeepSeek-V4.1-Flash, Floor Plan Maker, Muse, fal.live, Weather Next 3 and Claude Academy. Earlier articles in this series did not count this page, because its main list loads in the browser, but that block is part of the server-rendered HTML. It is a recent-tools block, not a ranking.
Gemini 3.5 Transcribe after 18 days
Google’s rival card, published on 27 August, now reads 12 on the Top 100 and 10 on the Ultimate List. On the Top 100, every one of its text hits sits in the hourly trending panel rather than in the ranked list itself. Eighteen days of head start have bought Gemini 3.5 Transcribe visibility in a panel, which is worth remembering when you read a zero for a seven-hour-old card.
The series readings on 14 September
| Card | Card published | Age at reading | Top 100 hits | Ultimate List hits |
|---|---|---|---|---|
| Gemini 3.5 Transcribe | 27 August | About 18 days | 12 | 10 |
| Claude Fable 5.1 | 3 September | About 11 days | 16 | 10 |
| Gemini 3.8 Flash | 3 September | About 11 days | 16 | 10 |
| GPT-6 Astra | 6 September | About 8 days | 0 | 0 |
| Lyria 3.5 | 6 September | About 8 days | 0 | 0 |
| Weather Next 3 | 6 September | About 8 days | 0 | 0 |
| Anyvids | 10 September | About 4 days | 0 | 7 |
| Google Pics | 14 September | About 7 hours | 0 | 0 |
| MAI-Transcribe-2 | 14 September | About 7 hours | 0 | 0 |
Claude Fable 5.1 and Gemini 3.8 Flash have read 16 and 10 at every check since 6 September, and GPT-6 Astra, Lyria 3.5 and Weather Next 3 still read zero after eight days. The seven Anyvids hits remain in promotional containers.
How MAI-Transcribe-2 Compares With Earlier Cards
Upstream lag across eight cards
The time between a vendor’s announcement and the AIxploria card varies widely. Each bar below is the gap recorded in that card’s article, measured from the vendor’s own publication; the Gemini 3.5 Transcribe figure is the maximum the dates allow.
A sixth badge shape
Every card we have tested has produced a different kind of rank story. GPT-6 Astra’s badge named a category that did not contain it. Lyria 3.5’s badged slot held its predecessor. Weather Next 3 ranked below its own predecessor, Anyvids below the models it resells, and Google Pics behind Google’s older Whisk. MAI-Transcribe-2 adds a sixth shape: a duplicate ordinal, where the badge names a slot that another card both claims and occupies.
Two transcription cards, 18 days apart
The two big-vendor transcription models make a useful pair. Gemini 3.5 Transcribe got its card within 27 and a half hours, a Freemium tag and a badge of #8 that matches its slot. MAI-Transcribe-2 waited ten and a half days and got a wrong pricing tag and a duplicated #25. On Artificial Analysis the order is reversed: MAI-Transcribe-2 leads on accuracy, 2.0% against 2.6%, on median speed, 285.8x against 87.8x, and on price, $0.10 against $0.30 an hour. Directory rank is not a benchmark.
What MAI-Transcribe-2 Means for Your Business
If you transcribe calls, meetings or dictation
Test MAI-Transcribe-2 on your own recordings rather than on FLEURS. Run the same files through verbatim and clean styles, feed in a keyword list of product and staff names, and check code switching if your customers mix languages. If you need speaker labels on anything longer than 15 minutes, build and price the two-step pipeline before you compare vendors. Meeting assistants such as the Agora meeting copilot built on GPT-Live-1 show how quickly transcription becomes a product feature rather than a standalone purchase.
If you build voice agents
A transcript is only the first stage of a voice agent. Intent detection, summarisation and routing are natural language processing tasks further down the chain, and they only work as well as the words they receive. Microsoft’s June post listed a native streaming API as “what’s next”, and the MAI-Transcribe-2 pages do not announce one, although the documentation covers input transcription in Voice Live. Our piece on orchestration for AI agents in customer service covers the rest of that stack.
If you are budgeting for 2027
Price MAI-Transcribe-2 at a rate you can defend after 31 December, not at $0.10. Microsoft has published no post-promotion price, and the family has changed version three times in 154 days. Write the switching cost into the plan: reference recordings, a comparison harness and a second supplier you have already tested. That is standard vendor management for any preview-stage model.
If you shortlist tools from directories
Treat this card as a lead, not a verdict. It describes the model reasonably well, and it still carried a badge another tool holds, a pricing tag its own FAQ contradicts and three errors about defaults, limits and Whisper. Our AI models and tools hub tracks the wider field, and an AI strategy review is the place to turn a shortlist into a decision.
How to Check a Directory Card Before Acting on It
Follow the badge and open the card in that slot
A badge links to its category. Open the page and look at the card sitting at the badge’s number. On this listing, six controls matched their badges, and one of them, the card in slot 25, was a different tool carrying the same badge as the new listing.
Check the pricing tag against the FAQ
Tags drive filtered views, so they are what skimming readers act on. When the tag says Free and the FAQ says “No”, believe the FAQ and then check the vendor.
Read the parameter table, not the pros list
Defaults and limits live in API documentation. On this card the pros list said “by default” for two features the documentation switches off, and the cons list gave a file limit Microsoft does not publish.
Recompute the vendor’s ratios from the benchmark
When a vendor quotes a benchmark, open the benchmark. Here, dividing the published medians turned “10x faster” into 7.7x, which is still impressive and far more useful for planning.
Frequently Asked Questions
Is MAI-Transcribe-2 free?
No. The card’s Free tag is wrong, and its own FAQ says so. MAI-Transcribe-2 costs $0.10 per hour of audio as a limited-time price through 31 December 2026. You can try it without code in the MAI Playground, and creating an Azure account is free.
Is the #25 rank badge accurate?
Not as far as we could test. Six control cards matched their badge positions exactly, but position 25 is Whisper Large V3 Turbo, which carries the same “#25 in Transcriber” badge. MAI-Transcribe-2 does not appear in the first 36 places, and pages four onwards sit behind a Cloudflare challenge.
Can MAI-Transcribe-2 label speakers in a one-hour meeting?
Not in a single request during the preview. Microsoft says diarization requests fail for recordings of about 15 minutes and longer. The documented workaround is to transcribe without diarization, keep word-level timestamps and run a separate diarization step.
How does MAI-Transcribe-2 compare with Gemini 3.5 Transcribe?
On Artificial Analysis, MAI-Transcribe-2 is more accurate at 2.0% AA-WER against 2.6%, faster at a median of 285.8x against 87.8x, and cheaper at $0.10 an hour against $0.30. On AIxploria, Gemini 3.5 Transcribe ranks higher, at #8 with a badge that matches its slot.
What will MAI-Transcribe-2 cost from January 2027?
Microsoft has not said. The launch post and the Azure pricing page both describe $0.10 an hour as a discounted, limited-time price through the end of 2026, and neither publishes the rate that follows. Budget for a higher figure and check the Azure pricing page before renewing.
References and Further Reading
MAI-Transcribe-2 — AIxploria
Transcriber category — AIxploria
Whisper Large V3 Turbo — AIxploria
MAI-Transcribe-2 model page — Microsoft AI
MAI-Transcribe in Azure Speech (preview) — Microsoft Learn
Transcriptions – Transcribe REST API reference — Microsoft Learn
Azure Speech pricing — Microsoft Azure
Introducing MAI-Transcribe-1.5 — Microsoft AI
State of the Art Speech Recognition with MAI-Transcribe-1 — Microsoft AI
Introducing MAI-Image-2.5-Pro and MAI-Voice-2-Flash — Microsoft AI
Speech to Text (ASR) Providers Leaderboard — Artificial Analysis
MAI-Transcribe 2 — OpenRouter
whisper/transcribe.py — OpenAI Whisper on GitHub
More AI coverage: explore Progressive Robot's AI Models, Tools & Releases hub — hands-on reviews, setup guides and benchmarks in one place.