Gemini 3.5 Transcribe is now listed on AIxploria, the AI tools directory that catalogues thousands of AI sites across more than fifty categories. The new card files Google’s speech-to-text model under Transcriber, where it already sits at number eight, prices it “Freemium”, rates it 4.5 out of 5 and stamps it with a Verified Tool badge and a “new AI tool” flag. For a model Google announced only the day before, that is a fast arrival on one of the most-browsed AI directories on the web.
The timing is the story. Google published the announcement on 26 August 2026, calling this its most precise speech-to-text model yet and confirming that it replaces Chirp 3, the transcription model that has quietly powered Google Cloud speech work for years. Within a day the directory card appeared — which means the model is now being discovered by people who were never reading Google’s model blog at all.
Our AI models and tools hub tracks releases like this one, and we ran the same exercise on the two previous arrivals in this series: Cloudflare OS and, last week, Claude Academy. This article does what a directory card cannot: it explains what the model actually does, how accurate it is against independent measurement, what it costs per hour of audio, where you can use it today, and which parts of the announcement do not yet apply to you.
Every figure quoted here comes from Google’s announcement, the Gemini API documentation and pricing pages, the AIxploria listing, or the independent Artificial Analysis leaderboard, all checked on 27 August 2026. Where a number is our arithmetic on those figures, we say so.
Table of contents
- What Gemini 3.5 Transcribe Actually Is
- Gemini 3.5 Transcribe on AIxploria: What the Listing Says
- How Accurate Is Gemini 3.5 Transcribe, Really?
- What Gemini 3.5 Transcribe Costs to Run
- Where You Can Use Gemini 3.5 Transcribe Today
- Gemini 3.5 Transcribe vs the Alternatives AIxploria Lists
- What Gemini 3.5 Transcribe Means for Your Business
- Where Gemini 3.5 Transcribe Falls Short
- Frequently Asked Questions
- References and Further Reading
What Gemini 3.5 Transcribe Actually Is
Gemini 3.5 Transcribe is Google’s new speech recognition model: a system that takes raw audio and returns clean, formatted, punctuated text rather than a stream of lowercase words. It is sold as a Gemini model, called through the Gemini API with a model ID, and priced per minute of audio. That framing is the shift — transcription has moved from a separate speech product into the same family as the rest of Google’s frontier models.
A model that edits while it listens
The headline capability is what Google calls smart transcription. Conventional speech recognition writes down what it hears, filler words and all; this one cleans as it goes. It removes the “ums” and “ahs”, resolves self-corrections in place — Google’s own example is “let’s meet Tuesday, no, Wednesday” landing as a single clean instruction — and applies automatic structured formatting, so a rambled thought comes back as paragraphs and lists instead of one unbroken block. It also handles live language switching mid-sentence, which is a real problem in multilingual offices and a genuine weakness of older systems.
The end of Chirp 3
Chirp 3 was Google’s previous transcription model, and Gemini 3.5 Transcribe supersedes it outright. Google reports that time to final transcription improves by 70 per cent against Chirp 3 — an important number, because it is the delay between someone finishing a sentence and the corrected text settling on screen, which is what makes dictation feel responsive or sluggish. On the FLEURS benchmark the new model records 5.50 per cent word error rate streaming and 5.04 per cent non-streaming.
The capability list that matters
Most of the feature list is standard for a 2026 speech model. Three entries are not: the disfluency editing described above, custom vocabulary at a scale that covers real domain jargon, and function calling — the model can hand a task off to another Gemini model mid-transcription, so speaking “make me an image of that” inside a dictation session actually produces one.
| Capability | What Gemini 3.5 Transcribe does | Documented limit |
|---|---|---|
| Language coverage | Detects and transcribes automatically, no language flag needed | More than 85 languages |
| Speaker diarization | Attributes segments to distinct speaker labels | Up to 8 speakers; 3+ experimental |
| Word-level timestamps | Start and end offset for every recognised word | May reduce overall accuracy |
| Custom vocabulary | Biases towards your product names and jargon | Up to 1,000 terms, best near 100 |
| Audio length per request | Single file transcription | 1 hour, or 30 minutes with diarization or timestamps |
| Function calling | Delegates tasks to other Gemini models mid-session | macOS app now, API support to follow |
Gemini 3.5 Transcribe on AIxploria: What the Listing Says
AIxploria is a discovery directory, not an evaluator, and its card is short. It is still worth reading closely, because it is how a large audience will first meet this model.
The entry itself
The listing sits in the Transcriber category, where it currently ranks number eight, priced Freemium, rated 4.5 out of 5 and marked as a Verified Tool with the “new AI tool” flag. The description is accurate and admirably plain: “A next-generation speech-to-text model that detects more than 85 languages, identifies different speakers, adds precise timestamps, and adapts to industry-specific vocabulary. Capable of converting raw audio into clean, formatted text, even in the presence of background noise.” The card carries a trend counter reading 9,994 and 68 upvotes, and its “useful links” point straight at Google AI Studio and the Agent Platform rather than at a marketing page.
Read the vote count, not the star rating
Here is the detail almost nobody checks. That 4.5 out of 5 is computed from four votes. The rating is not wrong, but it is not evidence either — it is four people’s launch-week enthusiasm, and it will move materially the first time a hundred people rate it. The same caveat applied when Cloudflare OS arrived on the directory with a 4.5, and when Claude Academy arrived with a 4.6. Treat the badge as a visibility signal and the star as noise until the sample grows.
What placement in Transcriber actually buys
Transcriber is a busy, commodity-heavy category on AIxploria — the alternatives the card itself suggests range from a free daily-minutes subtitle tool to a YouTube-to-blog converter. Landing at number eight in that company puts a frontier model from Google directly alongside single-purpose utilities, which is exactly the comparison most buyers are unequipped to make. A directory can tell you a tool exists; it cannot tell you that two entries on the same page differ by an order of magnitude in what they are for.
How Accurate Is Gemini 3.5 Transcribe, Really?
Accuracy in speech recognition is measured as word error rate: the share of words the system gets wrong, so lower is better. Google quotes two figures for Gemini 3.5 Transcribe, both attributed to the independent Artificial Analysis benchmark: 4.0 per cent for streaming audio and 2.6 per cent for pre-recorded files, across diverse real-world conditions.
Putting 2.6 per cent into human terms
A word error rate is an abstraction until you convert it. Ordinary speech runs at roughly 150 words per minute, so an hour of talk is about 9,000 words. At 2.6 per cent, that hour comes back with roughly 234 wrong words — our arithmetic on Google’s figure. That is a good transcript. It is not a transcript you can publish, bill from or put in front of a regulator without a human reading it, and no vendor in this category is currently selling one that is.
Where it sits on the independent leaderboard
Artificial Analysis publishes a public leaderboard, and this is where the marketing claim meets the field. Gemini 3.5 Transcribe entered at number five on the non-streaming table with its 2.6 per cent — a strong debut, and not the top. ElevenLabs Scribe v2 leads the named commercial field at 2.2 per cent, and a preview model, Fun-Realtime-ASR, sits at 1.7 per cent.
The gap that matters for most buyers is not the one above Gemini 3.5 Transcribe but the one below it. Whisper Large v3, still the default open model in thousands of internal pipelines, sits at 4.1 per cent — which is 58 per cent more errors than 2.6 per cent, on our arithmetic. That is the upgrade decision in one number.
What word error rate does not measure
Two systems with identical scores can feel completely different in use, because word error rate counts substitutions, insertions and deletions equally and cares nothing for which words. It does not measure whether the errors land on proper nouns, drug names or account numbers. It does not measure punctuation, paragraphing or speaker attribution — the things that decide whether a transcript is readable. And it does not measure the disfluency editing that is Gemini 3.5 Transcribe’s actual differentiator, because a benchmark scoring against a verbatim reference will mark a cleaned-up “um” as a deletion. Test on your own audio before you believe any of these numbers, including the good ones.
What Gemini 3.5 Transcribe Costs to Run
Google publishes the price openly on the Gemini API pricing page, in both token and per-minute form, which is more transparency than this category usually offers.
The published price
For the file-based model, gemini-3.5-transcribe, input audio is $0.003 per minute and text output is $0.002 per minute, giving a blended $0.005 per minute — $0.30 per hour of audio on our arithmetic. The streaming variant, gemini-3.5-transcribe-live, runs $0.005 in and $0.004 out, a blended $0.009 per minute, or $0.54 per hour. Both carry a free tier for evaluation.
The arithmetic on a real workload
Assume a modest professional-services firm recording 1,000 hours of client calls and meetings a year — about four hours a working day. At $0.30 per hour, the pre-recorded model costs $300 a year to transcribe all of it. Push the same volume through the live model for real-time captions and it is $540. Those figures are ours, computed from Google’s published per-minute rates, and they are the reason the build-versus-buy conversation around transcription has changed: the model is no longer the expensive part of the system.
Where the cheap option stops being cheap
Whisper at $1.15 per 1,000 minutes looks unanswerable until you price the rest of the job. It gives you a verbatim transcript and nothing else: no cleaned disfluencies, no automatic formatting, no jargon biasing, and diarization only if you bolt on a second system and maintain it. The extra $3.85 per 1,000 minutes that Gemini 3.5 Transcribe charges buys the post-processing you would otherwise build, plus a 1.5-point accuracy gain. On 1,000 hours a year that difference is about $231 — our arithmetic, and less than a day of an engineer’s time.
| Model | Provider | Word error rate | Per 1,000 min |
|---|---|---|---|
| Scribe v2 | ElevenLabs | 2.2% | $3.67 |
| MAI-Transcribe-1.5 | Microsoft Azure | 2.4% | $6.00 |
| Gemini 3.5 Transcribe | 2.6% | $5.00 | |
| Voxtral Small | Mistral | 2.8% | $4.00 |
| Universal-3 Pro | AssemblyAI | 3.1% | $3.50 |
| GPT Transcribe | OpenAI | 3.3% | $4.50 |
| Whisper Large v3 | fal.ai | 4.1% | $1.15 |
Where You Can Use Gemini 3.5 Transcribe Today
This is the part the directory card omits entirely, and it is the part that decides whether the announcement is relevant to you this week or next quarter. Availability is split across three audiences, and only one of them has full access right now.
Developer surfaces
Developers get it first. Gemini 3.5 Transcribe is in public preview through the Gemini API, reachable in Google AI Studio and Google Antigravity. Pre-recorded audio goes through the Interactions API with the model ID gemini-3.5-transcribe; real-time streaming uses the Live API over WebSockets with gemini-3.5-transcribe-live. Public preview is the operative phrase — it means the interface can still change, so treat anything you build now as a prototype rather than production.
Consumer surfaces
Consumers meet it without knowing its name. It is already live behind Rambler, the Gboard feature on Android, in selected countries and languages, and behind “Speak to Window” in the Gemini app on macOS, in English. Chrome support is announced but not shipped: the promise is talk-to-type in any web field, which is the feature most likely to change daily habits when it lands. Google also describes the model rolling out across Search Live, Docs, Keep and Gmail.
Enterprise surfaces
Enterprise access is the least complete. Gemini 3.5 Transcribe is in public preview on the Gemini Enterprise Agent Platform, while Gemini Enterprise for Customer Experience — the contact-centre route, and the one with the clearest business case — is listed as coming soon with no date attached.
| Surface | Who it is for | Status |
|---|---|---|
| Gemini API via AI Studio | Developers | Public preview |
| Google Antigravity | Developers | Public preview |
| Rambler on Gboard | Android users | Live, selected countries |
| Gemini app on macOS | Desktop users | Live, English |
| Chrome talk-to-type | Everyone | Announced, not shipped |
| Enterprise Agent Platform | Enterprise builders | Public preview |
| Enterprise for CX | Contact centres | Coming soon, no date |
Gemini 3.5 Transcribe vs the Alternatives AIxploria Lists
The directory’s own “alternatives” panel is a useful reality check, because it shows what the model is being shelved next to.
Two different products wearing the same label
The suggested alternatives include ElevenLabs Scribe v2, FineVoice, a free subtitles generator, a YouTube-to-blog converter and a meeting-notes agent. Only the first is a peer. The rest are applications built on top of speech recognition — finished tools with an interface, a workflow and a price per user. Gemini 3.5 Transcribe is a component you call from code, or a capability embedded invisibly in Google’s own apps. Choosing between them is not a comparison; it is a decision about whether you are buying a product or building one.
The one genuine peer
ElevenLabs Scribe v2 is the direct competitor, and on the directory card it is rated 4.6 to Google’s 4.5 while advertising 150-millisecond live latency and support for over 90 languages. On the independent leaderboard it is more accurate at 2.2 per cent and cheaper at $3.67 per 1,000 minutes. What Gemini 3.5 Transcribe offers against that is the disfluency editing, the function calling into other Gemini models, and — for anyone already on Google Cloud — one vendor, one bill and one data processing agreement. If you have no such attachment, the case for it is narrower than the announcement suggests.
What Gemini 3.5 Transcribe Means for Your Business
A cheap, accurate, formatting-aware transcription model is not an AI story so much as an operations story. Three areas change first.
Meeting notes stop being a licence line
Most firms pay per seat for a meeting-notes tool. At $0.30 an hour of audio, the transcription underneath that tool is close to free, which reframes the spend as payment for the interface, the integrations and the retention policy — all legitimate things to buy, but worth pricing honestly. Where a firm already has developers and a workflow platform, building the capture path directly on Gemini 3.5 Transcribe is now a small project rather than a programme.
Contact centres and the CX pipeline
Transcription is the raw material for everything a modern contact centre does with its calls: quality monitoring, compliance checking, sentiment analysis, agent coaching. Cheaper, better transcription raises the ceiling on all of it — though as we argued in our piece on orchestration as the new CX challenge, the model was rarely the bottleneck. The same caution applies to the voice-agent wave we covered when Ringg raised from Peak XV: better ears do not by themselves make a better agent.
Accessibility and the record-keeping question
Live captions, searchable archives and accessible recordings all get cheaper and better at once, and that is unambiguously good. The record-keeping consequence is less comfortable. When transcription costs $0.30 an hour, organisations transcribe everything — and every one of those transcripts is a discoverable record subject to UK GDPR, with a lawful basis, a retention period and a subject access obligation attached. Decide the retention policy before you switch the pipeline on, not after the first request lands.
Do the readiness work first
None of this pays off in a firm that has not sorted out who owns the output, where it is stored and who checks it. That is the same conclusion our AI readiness assessment reaches for every capability of this kind: the technology is rarely the constraint, and a model this cheap simply moves the constraint somewhere less convenient.
Where Gemini 3.5 Transcribe Falls Short
An honest read of a launch needs a limitations list, and this one has five entries worth stating plainly.
First, it is a preview. Public preview means the API surface can change and there is no stability commitment, so production plans should carry that risk explicitly. Second, it is not the accuracy leader — two named commercial models and one preview model beat 2.6 per cent, and the marketing language does not make that obvious.
Third, the availability gaps are real: Chrome support is unshipped, the contact-centre product has no date, and Rambler is limited to selected countries and languages. Fourth, the length limits bite sooner than expected — an hour per request drops to thirty minutes the moment you turn on speaker labels or word timestamps, so long recordings need chunking logic you have to write. Fifth, the diarization guarantee is narrower than the feature list implies: attribution beyond three speakers is explicitly experimental, which covers most real meetings.
None of these is disqualifying. They are the ordinary boundaries of a public preview, and knowing them is part of using Gemini 3.5 Transcribe well.
Frequently Asked Questions
Is Gemini 3.5 Transcribe free?
There is a free tier for evaluation, and the consumer features built on it — Rambler on Gboard, dictation in the Gemini app — cost nothing extra. Paid API use is $0.003 per minute of audio input plus $0.002 per minute of text output, a blended $0.005 per minute. AIxploria’s “Freemium” label is therefore accurate.
How does it compare with Whisper?
On the Artificial Analysis non-streaming leaderboard, Whisper Large v3 records 4.1 per cent word error rate against 2.6 per cent here — about 58 per cent more errors on our arithmetic. Whisper is far cheaper and can be self-hosted, which still matters for data residency, but it returns verbatim text with no formatting, no cleanup and no built-in speaker labels.
Can it tell speakers apart in a meeting recording?
Yes, up to eight speaker labels, but Google marks attribution beyond three speakers as experimental. Enabling diarization also cuts the maximum audio length per request from one hour to thirty minutes, so a long meeting needs splitting.
What replaced what, exactly?
Gemini 3.5 Transcribe supersedes Chirp 3, Google’s previous transcription model. Google reports a 70 per cent improvement in time to final transcription against it, alongside better word error rates on the FLEURS benchmark.
What exactly is AIxploria?
An AI tools directory that catalogues thousands of sites across more than fifty categories, with a running feed of new arrivals. A listing there is a visibility signal, not a quality audit — and on this card the 4.5 rating rests on four votes, so judge the model on the benchmark numbers instead.
References and Further Reading
Intelligent transcription with Gemini 3.5 Transcribe — Google
Audio transcription — Gemini API documentation
Gemini 3.5 Transcribe on AIxploria
The Ultimate List of Best AI Tools — AIxploria
Speech to Text (ASR) Providers Leaderboard — Artificial Analysis
More AI coverage: explore Progressive Robot's AI Models, Tools & Releases hub — hands-on reviews, setup guides and benchmarks in one place.