ElevenLabs released Eleven v4 and Eleven v4 Turbo on September 28, 2026. Eleven v4 is its new expressive text-to-speech model, and v4 Turbo is a faster version of the same technology for voice agents. Both speak more than 90 languages, up from 70-plus on Eleven v3, and both list at the same price as the v3 models they replace, with a 72% launch discount until October 12.
Two of the launch claims need reading closely. Turbo’s speed is given as about 100 milliseconds in one place and about 150 in another, and the two figures measure different stages. The claim that listeners prefer v4 about 75% of the time comes from ElevenLabs’ own blind tests. A third claim, first place on Artificial Analysis’ leaderboard, we re-read on launch day and it holds.
- 01Eleven v4 ranks first on Artificial Analysis' provider voice leaderboard.1319 Elo on September 28, 43 points clear of Cartesia Sonic 3.6. v4 Turbo is not yet listed there.
- 02Turbo's ~100 ms is inference time; ~150 ms is time to the first audible speech.Neither figure includes the speech recognition and language model stages that sit in front of text-to-speech in a voice agent.
- 03The ~75% preference figure is ElevenLabs' own test against four named rivals.Its chart shows a 65% to 81% win rate depending on the competitor, with ties counted as half a win.
- 04List prices match v3: $0.08 per 1,000 characters for v4, $0.04 for Turbo.Until October 12 they are $0.022 and $0.011. No OpenRouter route exists, so catalog comparisons cannot include either model.
01 — The releaseWhat shipped
- API model IDsPer ElevenLabs' models documentation.
- eleven_v4 · eleven_v4_turbo
- LanguagesEleven v3 supports 70+.
- 90+
- Characters per request, v4About 10 minutes of audio; v3 allows 5,000.
- 10,000
- Instant voice cloneProfessional Voice Clones are also supported on v4.
- 10 s of audio
- API surfacev4 via the Text to Dialogue API; Turbo via the Text to Dialogue WebSocket.
- Dialogue endpoints
According to the launch post, v4 is built on a new architecture and is designed to read tone, pacing, emotion and context from the text. It follows inline audio tags, such as a laugh, an accent or a background sound, more closely than earlier models, and pronunciations written in the International Phonetic Alphabet behave more reliably. Voices keep their identity across a long production, and a voice recorded in one language can speak another with a native accent.
Both models are available in ElevenLabs’ agent product (ElevenAgents), its creative tools and its API. ElevenLabs says it optimised v4 Turbo and ElevenAgents jointly, which is worth knowing if you assemble your own agent from separate vendors: some of the gains it describes are for its own stack.
02 — The latencyTwo latency figures that measure different stages
A voice agent listens, transcribes, decides what to say with a language model, then speaks. Text-to-speech latency is only the last step, and vendors measure even that step in different ways. The models documentation gives v4 Turbo a median inference latency of about 100 ms and marks the figure as excluding application and network latency. The launch post’s footnote gives a second number, about 150 ms, for the median time from sending a request to hearing speech, measured over WebSocket streaming.
| Model | Stated | What it measures |
|---|---|---|
| Eleven v4 Turbo | ~100 ms | Median inference latency, excluding application and network latency (documentation and blog headline) |
| Eleven v4 Turbo | ~150 ms | Median time from request to audible speech over WebSocket streaming, network latency removed (blog footnote) |
| Eleven v3 Conversational | ~280 ms | Same basis as the ~100 ms figure (documentation) |
| Eleven Flash v2.5 | ~75 ms | Same basis as the ~100 ms figure; faster but less expressive (documentation) |
The 150 ms figure is the more useful one for planning, because it is what a caller hears. ElevenLabs says it measured it with identical scripts and default settings against Cartesia Sonic 3.6, xAI’s text-to-speech model, Google’s Gemini 3.8 Flash-Lite TTS and OpenAI’s GPT-4o mini TTS, with network latency removed for all of them. Its chart puts the rivals at 262 to 814 ms on the same measure.
Turbo is not the fastest model ElevenLabs sells. Flash v2.5 is listed at about 75 ms on the same basis as the 100 ms figure. The change is that Turbo brings v4’s expressive delivery down to a speed that a live conversation can use, where v3 Conversational was listed at about 280 ms.
03 — The evidenceThe ranking claims, one independent and one vendor-run
First on Artificial Analysis. ElevenLabs cites Artificial Analysis’ Provider Voice Arena, where listeners vote blind between two clips and each model uses its provider’s own voices. We read the leaderboard on September 28. Eleven v4 leads, and it is the only model that ranks first across its whole confidence range.
| Rank | Model | Elo | Listed price |
|---|---|---|---|
| 1 | ElevenLabs Eleven v4 | 1319 ± 19 | $80 |
| 2 | Cartesia Sonic 3.6 | 1276 ± 16 | $49 |
| 3 | Google Gemini 3.8 Flash TTS | 1267 ± 16 | $16.5 |
| 4 | Alibaba Qwen-Audio-3.0-TTS-Plus | 1258 ± 16 | $19.3 |
| 5 | Inworld Realtime TTS-2 | 1246 ± 17 | $20.8 |
| 6 | Google Gemini 3.8 Flash-Lite TTS | 1241 ± 16 | $11 |
Two limits apply. The v4 rating rests on 1,674 votes and its interval, plus or minus 19 points, is the widest in the top six. And v4 Turbo does not appear on the leaderboard yet, so the ranking says nothing about the model most voice agents would use.
Preferred about 75% of the time. This figure comes from ElevenLabs’ own blind preference tests against Cartesia Sonic 3.6, Inworld TTS-2 and Google’s Gemini 3.8 Flash and Flash-Lite TTS. Each comparison played one line from v4 and from a rival, unlabelled, and graders picked the more expressive and the more natural of the two. Ties counted as half a win. The chart in the post shows v4 winning between 65% and 81% of tests depending on the competitor. The company does not say how many graders took part or how the lines were chosen, so treat it as a vendor claim.
Arena Elo measures which clip strangers prefer on short lines. Whether callers stay on the line is a different question, which our post on voice benchmarks and listener preference covers. Test v4 Turbo on your own scripts and callers before trusting either number.
04 — The invoiceThe price: same list price as v3, discounted until October 12
Eleven v4
The same list price as Eleven v3. Until October 12 it is $0.022, 72% off. ElevenLabs equates 1,000 characters with about a minute of audio.
Eleven v4 Turbo
The same list price as v3 Conversational and Flash. Until October 12 it is $0.011. This is the model for live voice agents.
As an illustration of scale, an agent that speaks 1,000 minutes a month generates roughly a million characters. At Turbo’s list price that is about $40 a month in text-to-speech charges, and about $11 during the discount. The volume is invented for the example; the rates are from ElevenLabs’ API pricing page on September 28.
For comparison, Artificial Analysis lists Google’s Gemini 3.8 Flash TTS at $16.50 per million characters and Cartesia Sonic 3.6 at $49, against $80 for Eleven v4. The discount lasts two weeks, so budget on the list price. Google’s rate is promotional too and doubles on January 1, 2027, as our Gemini 3.8 Flash TTS post explains.
One gap for anyone who compares models through a router: there is no ElevenLabs route in the OpenRouter catalog, which we checked at 21:28 UTC on September 28. Neither v4 model can be priced or called that way.
05 — The agentsWhat a voice agent gains, and what still needs testing
For a support or sales agent, three changes matter more than the ranking. First, expressive speech no longer forces a slow model: the launch post frames the old choice as fast but flat versus expressive but slow, and Turbo is ElevenLabs’ answer to it. Second, the model reads the whole conversation, so replies respond to what was just said rather than sounding like separate recordings. Third, better accent control means one brand voice can serve several markets.
Before switching a live agent, check four things:
- The endpoint. The documentation lists v4 on the Text to Dialogue API and Turbo on the Text to Dialogue WebSocket. Confirm your integration uses those before you change the model ID.
- End-to-end delay. Measure the full turn, including transcription and the language model, from your own region.
- Pronunciation. Run your product names, medical or legal terms through the IPA support and the audio tags.
- Cloned voices. If you use a voice clone, compare it on v4 with your current model; ElevenLabs says speaker similarity is much better, and your listeners will notice any change.
Our September text-to-speech ranking predates v4 and will be updated with it.
06 — Conclusionv4 leads on quality; Turbo’s case rests on tests you run yourself
Run v4 Turbo on your own call scripts during the discount, and decide on the list price rather than the launch price
Eleven v4’s first place is independent and current. The Turbo speed and the 75% preference figure are ElevenLabs’ own measurements, and neither covers the whole delay a caller hears. A week of side-by-side tests on real scripts will settle more than any of the launch numbers.