Eleven v4 is ElevenLabs’ new flagship text-to-speech model, launched on 28 September 2026 alongside Eleven v4 Turbo, a low-latency version built for voice agents. Within half an hour of the announcement, Artificial Analysis ranked Eleven v4 first on its Provider Voice arena leaderboard, the blind-listening contest most people mean when they talk about the “TTS Arena”, with an Elo score of 1,319.
ElevenLabs calls Eleven v4 its “most emotive” model and says it is “preferred by ~75% of listeners in blind head-to-head tests.” Turbo, it says, starts speaking in about 150 milliseconds. Those are strong claims, and they come from different tests with different methods.
This article explains what ElevenLabs launched, what the leaderboard figures actually show, how to read the vendor’s own numbers, what Eleven v4 costs at list and launch prices, and which businesses should switch. We have also worked out the per-hour cost of speech for each model, because per-character pricing hides it.
Table of contents
- What ElevenLabs Launched With Eleven v4
- How Eleven v4 Topped the TTS Arena Leaderboard
- Reading ElevenLabs’ Own Claims About Eleven v4
- Eleven v4 Turbo and Real-Time Voice Agents
- How Eleven v4 Compares With Its Rivals
- What Eleven v4 Costs
- Voice Cloning, Consent and Risk
- Should Your Business Switch to Eleven v4?
- Eleven v4 FAQs
- References
What ElevenLabs Launched With Eleven v4
The release splits ElevenLabs’ newest generation into two models with two jobs. Eleven v4 is the quality model for produced audio: audiobooks, dubbing, ads and games. Turbo carries the same expressive range at conversational speed for phone agents and live assistants. Both are available in ElevenCreative, ElevenAgents and the ElevenAPI, with model IDs eleven_v4 and eleven_v4_turbo.
Two models, two jobs
The table compares the new models with the ElevenLabs models they sit beside, using the company’s own documentation and API price list.
| Model | Built for | Latency | Languages | API list price per 1K characters |
|---|---|---|---|---|
| v4 | Expressive produced audio | Not quoted | 90+ | $0.08 ($0.022 until 12 Oct) |
| v4 Turbo | Real-time agents | ~100 ms inference | 90+ | $0.04 ($0.011 until 12 Oct) |
| Eleven v3 | Dramatic performance | Not quoted | 70+ | $0.08 |
| Eleven v3 Conversational | Real-time conversation | ~280 ms | 70+ | $0.04 |
| Flash v2.5 | Cheapest, fastest | ~75 ms | 32 | $0.04 |
ElevenLabs’ latency figures exclude application and network time. Flash v2.5 remains the fastest model on paper, so Turbo is not a replacement for it on speed alone. Its case is that it is fast enough while sounding far more human.
What changed from Eleven v3
ElevenLabs says Eleven v4 is built on “an entirely new architecture” that reads a script “the way a voice actor would”, aware of who is speaking, what just happened and how a line should land. The practical changes are easier to list:
- More languages. 90+ against 70+ for v3.
- Longer generations. 10,000 characters per request, up from 5,000.
- Better direction. Inline audio tags such as [laughs], [whispers] or [door slams] are followed “more reliably than v3”. TestingCatalog reports that SSML break tags are disabled, so pauses are directed in natural language instead.
- Stable speakers. Regenerating a line keeps the same voice, and “context stitching” keeps long projects consistent.
- Professional Voice Clones return. They were not supported in v3. Instant clones now need about 10 seconds of audio.
- Faster generation. Artificial Analysis measured 73.4 characters per second of generation time for the new model, against 42.5 for Eleven v3.
How Eleven v4 Topped the TTS Arena Leaderboard
Artificial Analysis runs a Speech Arena in which listeners hear the same line from two anonymous models and pick the better one. Votes feed an Elo rating, the same system used to rank chess players. On launch day Eleven v4 took first place on its Provider Voice leaderboard, which compares models using each provider’s own voices.
Which “TTS Arena” this is
Coverage has used “TTS Arena” loosely. The ranking behind the ElevenLabs headline is Artificial Analysis’s Provider Voice Arena. A separate community project on Hugging Face also calls itself TTS Arena. Anyone quoting Eleven v4’s position should name the Artificial Analysis leaderboard, since the two use different voices, prompts and voter pools.
The leaderboard on 29 September
The table shows the top six entries and ElevenLabs’ previous models on the Provider Voice leaderboard, as published by Artificial Analysis on 29 September 2026.
| Rank | Model | Elo (95% CI) | Samples | Price per 1M characters |
|---|---|---|---|---|
| 1 | ElevenLabs v4 | 1,319 (±19) | 1,674 | $80.0 |
| 2 | Cartesia Sonic 3.6 | 1,276 (±16) | 1,946 | $49.0 |
| 3 | Google Gemini 3.8 Flash TTS | 1,267 (±16) | 2,198 | $16.5 |
| 4 | Alibaba Qwen-Audio-3.0-TTS-Plus | 1,258 (±16) | 1,593 | $19.3 |
| 5 | Inworld Realtime TTS-2 | 1,246 (±17) | 1,365 | $20.8 |
| 6 | Google Gemini 3.8 Flash-Lite TTS | 1,241 (±16) | 2,161 | $11.0 |
| 13 | ElevenLabs v3 Conversational | 1,197 (±15) | 1,930 | $50.0 |
| 18 | ElevenLabs Eleven v3 | 1,169 (±11) | 4,401 | $100.0 |
The lead is statistically clear. Eleven v4’s lower bound (1,319 minus 19, or 1,300) sits above Sonic 3.6’s upper bound (1,276 plus 16, or 1,292), which is why Artificial Analysis shows Eleven v4’s rank range as first and first only. Artificial Analysis also says Eleven v4 ranks first in all four categories it tests: customer service, assistants, knowledge sharing and entertainment.
Two caveats. Turbo is not on the leaderboard, so its quality has not been ranked independently. And Artificial Analysis lists Eleven v3 at $100 per million characters, above the $80 ElevenLabs’ own API page shows, so compare prices from the vendor’s price list.
What a 43-point lead means in practice
Elo gaps translate into expected head-to-head win rates using the standard formula, 1 / (1 + 10^(-gap/400)). The chart applies it to Eleven v4’s published lead over three rivals.
In other words, the new model wins a clear majority of blind comparisons, but against the best rivals it is closer to 56 to 44 than to a landslide. The gap over ElevenLabs’ own previous model is much larger.
Controlled voice and pronunciation
Artificial Analysis runs two more tests. In its Controlled Voice arena, where every model speaks with the same custom voice, it ranks second with an Elo of 1,157, behind Alibaba’s Qwen-Audio-3.1-TTS-Plus at 1,178 and well ahead of Eleven v3 at 1,073. On its Pronunciation Robustness benchmark, it scored 91.7%, the highest the firm has measured, against 89.5% for Gemini 3.8 Flash TTS and 85.6% for Eleven v3.
Reading ElevenLabs' Own Claims About Eleven v4
ElevenLabs’ launch materials cite three kinds of evidence. Only one is independent, and it helps to keep them apart.
“Preferred by ~75% of listeners”
This figure comes from ElevenLabs’ own blind tests, not from Artificial Analysis. Its footnote says graders heard the same line from the new model and one competitor at a time (Cartesia Sonic 3.6, Inworld TTS-2 and two Gemini 3.8 models), judged which was more expressive and more natural, and counted ties as half. That design and the arena’s design differ, so the 75% and the arena’s implied 56% are not contradictory. They measure different things, with different voices and graders.
Latency is vendor-measured
ElevenLabs says Turbo reaches first speech in a median 150 ms, against 262 ms for Cartesia Sonic 3.6 and 814 ms for OpenAI’s GPT-4o mini TTS. It measured this itself in September 2026, with network latency removed and Turbo running over WebSocket streaming. Treat those as the best case. Your real figure will include your network, your language model and your telephony provider.
Eleven v4 Turbo and Real-Time Voice Agents
Most of the commercial interest is in Turbo, because voice agents live or die on response time. A caller notices a pause long before they notice a slightly flat intonation.
Bidirectional streaming
Turbo accepts text as a language model produces it and returns audio before the sentence is finished. That removes the wait for a full sentence of text before speech can begin, often the biggest delay in an agent loop. Professional Voice Clones work the same way on both models, so a brand voice can stay consistent across a call.
Where the speech model sits in an agent
A voice agent is a chain. Speech recognition turns the caller’s words into text, a language model handles the natural language processing and decides what to say, and a speech model such as Turbo turns the reply back into audio. Each link adds delay, so a fast speech model only helps if the other two are fast as well.
150 milliseconds in context
ElevenLabs says Turbo “can respond faster than the average pause between two people talking.” Conversation research supports the benchmark it is aiming at: a 2009 study of ten languages found that speakers typically begin their turn within about 200 milliseconds of the previous speaker finishing. A 150 ms first-speech time leaves little room for everything else in the pipeline, which is why teams building AI agents for phone lines should measure the whole round trip, not the speech model alone.
What launch customers say
ElevenLabs published quotes from four early users, and each points to a different use case:
- Salesforce. Ryan Peterson, who runs product for Agentforce Voice, says Turbo gives “faster, more natural responses” that let customers take the product into “even more use cases”.
- BeyondWords. Co-founder Patrick O’Flaherty, whose company turns publishers’ journalism into audio, says Eleven v4 gives publishers “more ways to make sure those voices feel engaging, familiar, and distinctly their own.”
- Rosebud. Chief executive Chrys Bader says voices are “not only more expressive, but more controllable”.
- Accenture. Kyle Gudmundson, a music and audio lead, says his first session “felt like I was running a session with talent in the booth.”
These are launch-day testimonials, chosen by the vendor, so treat them as signals of where the model is being tried rather than as independent reviews.
How Eleven v4 Compares With Its Rivals
The arena table shows how crowded the top of the text-to-speech market has become. Five companies sit within 80 Elo points of each other, and most released a new model in the past two months.
Cartesia Sonic 3.6
Cartesia’s model, released in August 2026, was second on the arena on 29 September at 1,276 and costs $49 per million characters on Artificial Analysis’s figures. ElevenLabs’ own latency test puts it at 262 ms to first speech against Turbo’s 150 ms. For teams that already run Cartesia in production, the new ElevenLabs release is a reason to re-test rather than an automatic switch.
Google’s Gemini 3.8 TTS models
Google has two models in the top six: Gemini 3.8 Flash TTS at 1,267 and $16.50 per million characters, and Gemini 3.8 Flash-Lite TTS at 1,241 and $11. They are far cheaper than Eleven v4 at list price, and Gemini 3.8 Flash TTS came within 2.2 points of it on pronunciation robustness (89.5% against 91.7%).
Alibaba, Inworld and the rest
Alibaba’s Qwen-Audio-3.0-TTS-Plus (1,258) and Inworld’s Realtime TTS-2 (1,246) complete the top five, both at around $20 per million characters. Alibaba’s newer Qwen-Audio-3.1-TTS-Plus leads the Controlled Voice arena. The practical lesson is that price and quality now vary independently, so benchmark the two or three candidates that fit your budget on your own scripts.
What Eleven v4 Costs
ElevenLabs charges by character, which makes comparisons hard. We converted list prices into the cost of one hour of speech, assuming 150 spoken words a minute at six characters per word including spaces. That is 900 characters a minute, or 54,000 an hour.
List price and the launch discount
At list price Eleven v4 costs $0.08 per 1,000 characters through the API, the same as v3, and Turbo costs $0.04. Until 12 October both carry a 72% launch discount on the API, bringing them to $0.022 and $0.011. The chart shows the resulting cost per hour of speech, calculated as 54 x the price per 1,000 characters.
At list price, Eleven v4 costs about 1.6 times Sonic 3.6 and nearly five times Gemini 3.8 Flash TTS per hour. The quality lead has to be worth that premium for your use case. During the launch window, Eleven v4 undercuts Sonic 3.6, and Turbo undercuts Gemini 3.8 Flash TTS.
Plans and credits
App users pay through monthly plans rather than per character. ElevenLabs lists Free (10,000 credits), Starter ($6, 30,000), Creator ($22, 121,000), Pro ($99, 600,000), Scale ($299, 1.8 million) and Business ($990, 6 million), with Enterprise priced on request. For two weeks, Creator plans and above can use up to twice their monthly text-to-speech credits on Eleven v4 in the web and mobile apps without it counting against their balance.
A worked example for a voice agent
Take a service line handling 10,000 calls a month, with the agent speaking for two minutes per call. That is 20,000 minutes, or 18 million characters at 900 a minute. On Turbo at list price, 18,000 x $0.04 is $720 a month. At the launch price it is $198. The same volume on Eleven v4 at list price would cost $1,440, which is why Turbo, not v4, is the realistic choice for high-volume calls.
Voice Cloning, Consent and Risk
Better cloning is part of the pitch. Instant clones need only about 10 seconds of audio, and Professional Voice Clones perform with the model’s full emotional range across its languages. TestingCatalog reports that every cloned voice needs verified owner consent and that generated audio is covered by ElevenLabs’ AI Speech Classifier, which can identify its output.
That matters because voice cloning disputes are already in court. Our report on a Japanese anime actor fighting TikTok over AI voice cloning shows how quickly a cloned voice becomes a legal problem. If you clone a staff member or presenter, get written consent that covers the new model and every language you plan to use.
Cloned voices are also a cybersecurity risk
Criminals already use cloned voices to impersonate executives and relatives on the phone. As cloning improves, voice alone becomes a weaker proof of identity. Businesses that approve payments or password resets by phone should add a second check, such as a call-back to a known number, whatever voice tools they use themselves.
Should Your Business Switch to Eleven v4?
The answer depends on what you produce. The table matches common uses to the model that fits them best.
| Use case | Best fit | Why |
|---|---|---|
| Audiobooks and long narration | v4 | 10,000-character requests, stitching, stable speakers |
| Dubbing and localisation | v4 | 90+ languages with the original voice kept |
| Phone and chat voice agents | v4 Turbo | Streaming, ~150 ms first speech, half the price |
| High-volume, cost-first alerts | Flash v2.5 or a cheaper rival | Lowest latency and price; expression matters less |
| Games and character work | v4 | Audio tags, sound effects, multi-speaker scenes |
A migration checklist
Switching is mostly a change of model ID, but test before you move production traffic:
- Retrain your clones. TestingCatalog reports that older Instant and Professional Voice Clones need retraining for v4.
- Replace SSML breaks. Move pauses into natural-language audio tags.
- Re-check pronunciations. Run your product names and jargon through the new model, using IPA phonemes where needed.
- Measure end-to-end latency. Time the whole agent loop on your own network and telephony.
- Lock in the launch price. Run your largest back-catalogue jobs before the discount ends on 12 October.
If you are new to the platform, our step-by-step guide on how to set up ElevenLabs covers accounts, API keys and voices. For the wider picture of voice in customer service, see our intelligent automation services.
Eleven v4 FAQs
Is Eleven v4 really number one on the TTS Arena?
Yes, on Artificial Analysis’s Provider Voice leaderboard, where it had an Elo of 1,319 on 29 September 2026, ahead of Cartesia Sonic 3.6 at 1,276. On the Controlled Voice leaderboard it ranks second.
What is the difference between Eleven v4 and Eleven v4 Turbo?
Eleven v4 is tuned for quality in produced audio. Turbo keeps its expressive range but streams in real time, with about 150 ms to first speech, and costs half as much per character.
How much does Eleven v4 cost?
Through the API, $0.08 per 1,000 characters at list price and $0.022 until 12 October. Turbo is $0.04, or $0.011 during the launch offer. App plans start free with 10,000 monthly credits.
How many languages does Eleven v4 support?
More than 90, up from 70+ for Eleven v3, and a cloned voice can speak other languages with a native accent.
Can I use Eleven v4 commercially on the free plan?
No. ElevenLabs’ pricing page lists a commercial licence from the $6 Starter plan upwards. The free plan’s 10,000 monthly credits are for trying the models.
Is Eleven v4 Turbo ranked on the leaderboard?
Not yet. Artificial Analysis ranked Eleven v4 on launch day, but Turbo had no arena entry on 29 September, so its quality has only been described by ElevenLabs.
Do my existing voice clones work with Eleven v4?
All 17,500+ library voices work, but TestingCatalog reports that older Instant and Professional Voice Clones need retraining.
References
Introducing Eleven v4, our most emotive model (ElevenLabs)
Eleven v4 and Eleven v4 Turbo Text to Speech models (ElevenLabs)
Models (ElevenLabs Documentation)
ElevenAPI Pricing (ElevenLabs)
Provider Voice Arena Leaderboard (Artificial Analysis)
Eleven v4 takes first place on the Provider Voice leaderboard (Artificial Analysis on X)
ElevenLabs launches Eleven v4 and v4 Turbo voice models (TestingCatalog)
Universals and cultural variation in turn-taking in conversation (PNAS, via PMC)
More AI coverage: explore Progressive Robot's AI Models, Tools & Releases hub — hands-on reviews, setup guides and benchmarks in one place.