AI DevelopmentNew Release6 min readPublished September 28, 2026

eleven_v4 · eleven_v4_turbo · 90+ languages · two latency numbers that measure different things

Eleven v4 and v4 Turbo: What Changed for Voice Agents

ElevenLabs' Eleven v4 tops Artificial Analysis' voice arena and v4 Turbo claims ~100ms. The two latency figures, the vendor tests and the v4 price.

DA
Digital Applied Team
Research and practical guidance
ReleasedSeptember 28, 2026
Sources checkedSeptember 28, 2026

ElevenLabs released Eleven v4 and Eleven v4 Turbo on September 28, 2026. Eleven v4 is its new expressive text-to-speech model, and v4 Turbo is a faster version of the same technology for voice agents. Both speak more than 90 languages, up from 70-plus on Eleven v3, and both list at the same price as the v3 models they replace, with a 72% launch discount until October 12.

Two of the launch claims need reading closely. Turbo’s speed is given as about 100 milliseconds in one place and about 150 in another, and the two figures measure different stages. The claim that listeners prefer v4 about 75% of the time comes from ElevenLabs’ own blind tests. A third claim, first place on Artificial Analysis’ leaderboard, we re-read on launch day and it holds.

Key takeaways
  1. 01
    Eleven v4 ranks first on Artificial Analysis' provider voice leaderboard.1319 Elo on September 28, 43 points clear of Cartesia Sonic 3.6. v4 Turbo is not yet listed there.
  2. 02
    Turbo's ~100 ms is inference time; ~150 ms is time to the first audible speech.Neither figure includes the speech recognition and language model stages that sit in front of text-to-speech in a voice agent.
  3. 03
    The ~75% preference figure is ElevenLabs' own test against four named rivals.Its chart shows a 65% to 81% win rate depending on the competitor, with ties counted as half a win.
  4. 04
    List prices match v3: $0.08 per 1,000 characters for v4, $0.04 for Turbo.Until October 12 they are $0.022 and $0.011. No OpenRouter route exists, so catalog comparisons cannot include either model.

01 — The releaseWhat shipped

API model IDsPer ElevenLabs' models documentation.
eleven_v4 · eleven_v4_turbo
LanguagesEleven v3 supports 70+.
90+
Characters per request, v4About 10 minutes of audio; v3 allows 5,000.
10,000
Instant voice cloneProfessional Voice Clones are also supported on v4.
10 s of audio
API surfacev4 via the Text to Dialogue API; Turbo via the Text to Dialogue WebSocket.
Dialogue endpoints

According to the launch post, v4 is built on a new architecture and is designed to read tone, pacing, emotion and context from the text. It follows inline audio tags, such as a laugh, an accent or a background sound, more closely than earlier models, and pronunciations written in the International Phonetic Alphabet behave more reliably. Voices keep their identity across a long production, and a voice recorded in one language can speak another with a native accent.

Both models are available in ElevenLabs’ agent product (ElevenAgents), its creative tools and its API. ElevenLabs says it optimised v4 Turbo and ElevenAgents jointly, which is worth knowing if you assemble your own agent from separate vendors: some of the gains it describes are for its own stack.

02 — The latencyTwo latency figures that measure different stages

A voice agent listens, transcribes, decides what to say with a language model, then speaks. Text-to-speech latency is only the last step, and vendors measure even that step in different ways. The models documentation gives v4 Turbo a median inference latency of about 100 ms and marks the figure as excluding application and network latency. The launch post’s footnote gives a second number, about 150 ms, for the median time from sending a request to hearing speech, measured over WebSocket streaming.

Source: ElevenLabs models documentation and Eleven v4 launch post, September 28, 2026. All figures are vendor-reported; ElevenLabs labels only the v4 Turbo figures as medians.
ModelStatedWhat it measures
Eleven v4 Turbo~100 msMedian inference latency, excluding application and network latency (documentation and blog headline)
Eleven v4 Turbo~150 msMedian time from request to audible speech over WebSocket streaming, network latency removed (blog footnote)
Eleven v3 Conversational~280 msSame basis as the ~100 ms figure (documentation)
Eleven Flash v2.5~75 msSame basis as the ~100 ms figure; faster but less expressive (documentation)

The 150 ms figure is the more useful one for planning, because it is what a caller hears. ElevenLabs says it measured it with identical scripts and default settings against Cartesia Sonic 3.6, xAI’s text-to-speech model, Google’s Gemini 3.8 Flash-Lite TTS and OpenAI’s GPT-4o mini TTS, with network latency removed for all of them. Its chart puts the rivals at 262 to 814 ms on the same measure.

Turbo is not the fastest model ElevenLabs sells. Flash v2.5 is listed at about 75 ms on the same basis as the 100 ms figure. The change is that Turbo brings v4’s expressive delivery down to a speed that a live conversation can use, where v3 Conversational was listed at about 280 ms.

03 — The evidenceThe ranking claims, one independent and one vendor-run

First on Artificial Analysis. ElevenLabs cites Artificial Analysis’ Provider Voice Arena, where listeners vote blind between two clips and each model uses its provider’s own voices. We read the leaderboard on September 28. Eleven v4 leads, and it is the only model that ranks first across its whole confidence range.

Source: Artificial Analysis Provider Voice Arena Leaderboard, read September 28, 2026. Elo with 95% confidence interval; API price per million characters as Artificial Analysis lists it.
RankModelEloListed price
1ElevenLabs Eleven v41319 ± 19$80
2Cartesia Sonic 3.61276 ± 16$49
3Google Gemini 3.8 Flash TTS1267 ± 16$16.5
4Alibaba Qwen-Audio-3.0-TTS-Plus1258 ± 16$19.3
5Inworld Realtime TTS-21246 ± 17$20.8
6Google Gemini 3.8 Flash-Lite TTS1241 ± 16$11

Two limits apply. The v4 rating rests on 1,674 votes and its interval, plus or minus 19 points, is the widest in the top six. And v4 Turbo does not appear on the leaderboard yet, so the ranking says nothing about the model most voice agents would use.

Preferred about 75% of the time. This figure comes from ElevenLabs’ own blind preference tests against Cartesia Sonic 3.6, Inworld TTS-2 and Google’s Gemini 3.8 Flash and Flash-Lite TTS. Each comparison played one line from v4 and from a rival, unlabelled, and graders picked the more expressive and the more natural of the two. Ties counted as half a win. The chart in the post shows v4 winning between 65% and 81% of tests depending on the competitor. The company does not say how many graders took part or how the lines were chosen, so treat it as a vendor claim.

Why rankings and preference disagree

Arena Elo measures which clip strangers prefer on short lines. Whether callers stay on the line is a different question, which our post on voice benchmarks and listener preference covers. Test v4 Turbo on your own scripts and callers before trusting either number.

04 — The invoiceThe price: same list price as v3, discounted until October 12

v4
Eleven v4
$0.08 per 1,000 characters

The same list price as Eleven v3. Until October 12 it is $0.022, 72% off. ElevenLabs equates 1,000 characters with about a minute of audio.

Narration, dubbing
Turbo
Eleven v4 Turbo
$0.04 per 1,000 characters

The same list price as v3 Conversational and Flash. Until October 12 it is $0.011. This is the model for live voice agents.

Voice agents

As an illustration of scale, an agent that speaks 1,000 minutes a month generates roughly a million characters. At Turbo’s list price that is about $40 a month in text-to-speech charges, and about $11 during the discount. The volume is invented for the example; the rates are from ElevenLabs’ API pricing page on September 28.

For comparison, Artificial Analysis lists Google’s Gemini 3.8 Flash TTS at $16.50 per million characters and Cartesia Sonic 3.6 at $49, against $80 for Eleven v4. The discount lasts two weeks, so budget on the list price. Google’s rate is promotional too and doubles on January 1, 2027, as our Gemini 3.8 Flash TTS post explains.

One gap for anyone who compares models through a router: there is no ElevenLabs route in the OpenRouter catalog, which we checked at 21:28 UTC on September 28. Neither v4 model can be priced or called that way.

05 — The agentsWhat a voice agent gains, and what still needs testing

For a support or sales agent, three changes matter more than the ranking. First, expressive speech no longer forces a slow model: the launch post frames the old choice as fast but flat versus expressive but slow, and Turbo is ElevenLabs’ answer to it. Second, the model reads the whole conversation, so replies respond to what was just said rather than sounding like separate recordings. Third, better accent control means one brand voice can serve several markets.

Before switching a live agent, check four things:

  • The endpoint. The documentation lists v4 on the Text to Dialogue API and Turbo on the Text to Dialogue WebSocket. Confirm your integration uses those before you change the model ID.
  • End-to-end delay. Measure the full turn, including transcription and the language model, from your own region.
  • Pronunciation. Run your product names, medical or legal terms through the IPA support and the audio tags.
  • Cloned voices. If you use a voice clone, compare it on v4 with your current model; ElevenLabs says speaker similarity is much better, and your listeners will notice any change.

Our September text-to-speech ranking predates v4 and will be updated with it.

06 — Conclusionv4 leads on quality; Turbo’s case rests on tests you run yourself

You run a live voice agent on ElevenLabs
Trial v4 Turbo during the discount, but measure the full turn time and cost at the $0.04 list price.
eleven_v4_turbo
You produce narration, dubbing or ads
Test v4 against v3 on a long script. The 10,000-character limit and steadier voices across long generations are the practical gains.
eleven_v4
You use another provider today
Compare against Gemini 3.8 Flash TTS and Cartesia Sonic 3.6 on your own lines. v4 lists at about 1.6 times Cartesia’s price per character and nearly five times Gemini’s.
Blind test on your scripts
What to do this week

Run v4 Turbo on your own call scripts during the discount, and decide on the list price rather than the launch price

Eleven v4’s first place is independent and current. The Turbo speed and the 75% preference figure are ElevenLabs’ own measurements, and neither covers the whole delay a caller hears. A week of side-by-side tests on real scripts will settle more than any of the launch numbers.

Digital Applied

Build a voice agent people stay on the line for.

We design, test and run voice agents, from model choice and latency budgets to the scripts and evaluation sets that show whether they work.

Voice model testingLatency budgetsAgent design
Your next project

A voice agent you can measure

  • →End-to-end turn time from your region
  • →Blind tests on your own scripts
  • →Cost at list price, not launch price
Questions and answers

The questions we get about Eleven v4

Eleven v4 is the higher-quality model for narration, dialogue and dubbing. v4 Turbo uses the same technology tuned for speed, for live voice agents, and costs half as much per character.
Digital Applied newsletter

Deep dives on AI, marketing and development.

Practical guides and fresh insights by email. No recycled takes.

Related dispatches

Continue reading

AI Development

Best Text-to-Speech Models, September 2026: Ranked, Priced

92 text-to-speech models ranked by blind listener vote, with the price per million characters, open-weight licences and the job each leader fits.

September 24, 2026 · 9 minRead
AI Development

Gemini 3.8 Flash TTS: Voice Cloning and a Price That Doubles

Gemini 3.8 Flash TTS is generally available with voice cloning from a 30-second sample. The promotional price ends December 31 and doubles on January 1, 2027.

September 23, 2026 · 5 minRead
AI Development

Decision Models: AI That Returns Probabilities, Not Text

Four publishers now list decision models on OpenRouter that return probabilities for typed questions, not prose, at free to $0.05 per million input tokens.

September 28, 2026 · 6 minRead
AI Development

Migrating to Claude Sonnet 5.5: Every Breaking Change

Claude Sonnet 5.5 rejects thinking disabled, forced tool choice and three more settings. The exact errors and fixes, grouped by the model you run today.

September 28, 2026 · 6 minRead
AI Development

Computer-Use Agents: Microsoft vs Anthropic vs Google

Microsoft GA, Anthropic public beta, and Google Gemini preview — OSWorld scores now 78% across frontier models above the ~72% human baseline. Routing guide.

May 22, 2026 · 16 minRead
AI Development

Agent Computer Use: Enterprise Automation Playbook

Enterprise playbook for deploying computer-use agents — a 40-point guardrails checklist spanning identity, audit, action boundaries, failures, and compliance.

May 22, 2026 · 17 minRead