By listener preference on the Artificial Analysis Speech Arena, read September 25, 2026, the best text-to-speech model is Cartesia's Sonic 3.6 at 1279 Elo. Second is Google's Gemini 3.8 Flash TTS, released this week, at 1265 and about a third of the price. On the other independent board, Voice Arena's US-English leaderboard, the cheaper Gemini 3.8 Flash-Lite TTS holds first place. The confidence intervals of the top four on Artificial Analysis overlap, so "best" is a tie broken by price and by the job you need done.
This page is the dated table: 92 models ranked by blind pairwise vote, every vendor's price normalised to dollars per million characters and per minute of audio, the licences of the 16 open-weight models, and a router by job. It is a snapshot. Both boards move weekly, and Google's launch prices double on January 1, 2027.
- 01Sonic 3.6 leads on votes; Gemini 3.8 Flash TTS is close behind at a third of the price.Artificial Analysis lists Sonic 3.6 at $49 per million characters and Gemini 3.8 Flash TTS at $16.49, with a 14-point Elo gap inside both models' 17-point confidence intervals.
- 02The two independent boards disagree at the top, and both are right.Voice Arena's US-English board, built from real production-call prompts, puts Gemini 3.8 Flash-Lite first by confidence-interval rank. Artificial Analysis, across all categories and accents, puts it sixth. Test on your own text.
- 03The strongest open-weight models cannot be used commercially.Breeze TTS 2 (rank 9) is research and non-commercial; Voxtral TTS is CC BY-NC. The best commercially usable open models, NVIDIA's Magpie and Kokoro 82M, sit near rank 48 at 1063 Elo.
- 04Vendor latency figures are claims; of the two boards, only Voice Arena measures them.Artificial Analysis publishes no time-to-first-audio for TTS right now. Voice Arena's measured medians run from 123 ms (Simba 3.2) to 890 ms (Gemini 3.1 Flash TTS). Cartesia's "sub-90ms" and ElevenLabs' "~75ms" are vendor statements.
01 — The tableThe ranked table: 22 of 92 models by listener vote
The Artificial Analysis text-to-speech leaderboard ranks models by Elo from blind pairwise votes in its Speech Arena, with a 95% confidence interval, across all prompt categories and both US and UK accents. The page carries no as-of date, so we date the snapshot to the day we parsed it. Price is Artificial Analysis' normalised figure in dollars per million characters at the endpoint it tracks. The table stops at rank 22, at 1130 Elo, tied with ranks 23 and 24; other vendors appear in the price table.
| Model | Elo ± CI | Votes | $ per 1M chars |
|---|---|---|---|
| 1. Sonic 3.6 (Cartesia) | 1279 ± 17 | 1,796 | 49.00 |
| 2. Gemini 3.8 Flash TTS (Google) | 1265 ± 17 | 2,044 | 16.49 |
| 3. Qwen-Audio-3.0-TTS-Plus (Alibaba) | 1259 ± 16 | 1,477 | 27.59 |
| 4. Realtime TTS-2 (Inworld) | 1246 ± 18 | 1,269 | 20.83 |
| 5. Simba 3.2 (Speechify) | 1239 ± 14 | 2,432 | 6.58 |
| 6. Gemini 3.8 Flash-Lite TTS (Google) | 1239 ± 16 | 2,056 | 11.03 |
| 7. Luna TTS (VUI Labs) | 1229 ± 14 | 2,581 | 15.00 |
| 8. Realtime TTS-2 Flash (Inworld) | 1210 ± 15 | 1,751 | 10.42 |
| 9. Breeze TTS 2 (BreezeBlue, open weights) | 1205 ± 16 | 1,440 | 34.00 |
| 10. Gemini 3.1 Flash TTS (Google) | 1202 ± 12 | 3,484 | 18.31 |
| 11. StepAudio 2.5 TTS (StepFun) | 1199 ± 16 | 1,344 | 85.00 |
| 12. v3 Conversational (ElevenLabs) | 1196 ± 15 | 1,889 | 50.00 |
| 13. Sonic 3.5 (Cartesia) | 1185 ± 12 | 3,039 | 49.00 |
| 14. Lightning V3.1 Pro (Smallest.ai) | 1174 ± 14 | 1,972 | 19.50 |
| 15. TTS Real-Time v2 (Soniox) | 1173 ± 12 | 4,671 | 14.23 |
| 16. Speech 2.8 HD (MiniMax) | 1170 ± 11 | 4,460 | 100.00 |
| 17. Eleven v3 (ElevenLabs) | 1169 ± 11 | 4,390 | 100.00 |
| 18. Falcon 2 (Murf AI) | 1159 ± 15 | 1,785 | 10.00 |
| 19. Speech 2.8 Turbo (MiniMax) | 1149 ± 11 | 4,370 | 60.00 |
| 20. Gradium TTS (Gradium) | 1149 ± 14 | 1,726 | 47.22 |
| 21. S2.1 Pro (Fish Audio) | 1140 ± 13 | 2,226 | 15.00 |
| 22. SpaceXAI TTS (xAI) | 1130 ± 15 | 1,258 | 15.00 |
Read the intervals before the ranks. Sonic 3.6 at 1279 ± 17 and Gemini 3.8 Flash TTS at 1265 ± 17 overlap; so do Qwen-Audio-3.0-TTS Plus at 1259 ± 16 and Inworld's Realtime TTS-2 at 1246 ± 18. Simba 3.2 and Gemini 3.8 Flash-Lite TTS share 1239 exactly. Across the top six, price varies far more than Elo, a more than seven-fold spread: Simba 3.2 at $6.58 per million characters against Sonic 3.6 at $49. Artificial Analysis' own frontier note names three models on the quality-versus-price line: Sonic 3.6, Gemini 3.8 Flash TTS and Simba 3.2.
What the top ten cost, in dollars per million characters
Artificial Analysis normalised endpoint prices, September 25, 2026. Bars scaled to Sonic 3.6 at $49.One normalisation deserves caution. Google bills its TTS models per audio token, not per character. Gemini 3.8 Flash TTS costs $9 per million output tokens and Gemini 3.1 Flash TTS costs $20, a 2.2× gap, yet Artificial Analysis' per-character figures for the two are $16.49 and $18.31, only 1.1× apart. We print the Artificial Analysis numbers for consistency and Google's own per-minute rate in the price table. Our post on the Gemini 3.8 Flash TTS launch works through the token arithmetic.
02 — The second boardVoice Arena's US-English board, the one Google cited
Google's launch post cited Voice Arena, whose US-English text-to-speech leaderboard uses Bradley-Terry Elo on prompts it says are built verbatim from real production calls. Its ranks come with a confidence-interval range, shown in brackets, and unlike Artificial Analysis it measures median time to first audio itself rather than repeating vendor claims. Rows where it has not yet measured latency say so.
| Rank and model | Elo ± | Time to first audio, P50 | Votes |
|---|---|---|---|
| 1 (1–4) · Gemini 3.8 Flash-Lite TTS | 1087 ± 15 | not measured | 636 |
| 2 (1–5) · Cartesia Sonic-3.6 | 1068 ± 13 | 341 ms | 1,004 |
| 3 (1–5) · Gemini 3.8 Flash TTS | 1061 ± 12 | not measured | 1,161 |
| 4 (1–6) · Inworld Realtime TTS 2 (preview) | 1058 ± 15 | 169 ms | 849 |
| 5 (2–5) · Gemini 3.1 Flash TTS | 1057 ± 6 | 890 ms | 6,123 |
| 6 (6–8) · Cartesia Sonic-3.5 | 1033 ± 7 | 250 ms | 4,049 |
| 7 (6–8) · Speechify Simba 3.2 | 1032 ± 6 | 123 ms | 4,534 |
| 8 (5–8) · ElevenLabs v3 Conversational | 1031 ± 15 | not measured | 837 |
| 9 (9–14) · xAI Grok TTS | 996 ± 6 | 354 ms | 6,019 |
| 10 (9–14) · Murf Falcon 2 | 993 ± 11 | 528 ms | 1,550 |
| 11 (9–16) · Fish Audio S2.1 Pro | 992 ± 16 | 283 ms | 802 |
| 12 (9–16) · Maya-2-Global | 988 ± 9 | not measured | 1,662 |
| 13 (9–16) · ElevenLabs v3 | 986 ± 6 | 588 ms | 5,413 |
| 14 (9–16) · Gradium TTS | 985 ± 7 | 236 ms | 3,225 |
| 15 (11–17) · Microsoft Azure Dragon HD Omni | 975 ± 6 | 485 ms | 5,861 |
| 16 (11–17) · Hithink Speech 2.6 | 971 ± 10 | not measured | 1,479 |
| 17 (15–17) · Smallest Lightning 3.1 Pro | 962 ± 10 | 264 ms | 1,589 |
| 18 (18–19) · Fish Audio S2 Pro | 939 ± 7 | 267 ms | 5,290 |
| 19 (18–19) · OpenAI gpt-4o-mini-tts | 929 ± 7 | 812 ms | 5,205 |
| 20 · Cartesia Sonic-3 | 858 ± 10 | not measured | 3,200 |
The two boards agree on the cluster and disagree on the order. Sonic 3.6, Gemini 3.8 Flash TTS and Inworld's TTS-2 are top-four on both; Simba 3.2 is fifth on one and seventh on the other. The largest divergence at the top is Gemini 3.8 Flash-Lite TTS, first here by confidence-interval rank and sixth on Artificial Analysis. Its latency is not yet measured on Voice Arena, so the fastest measured models on this board are Simba 3.2 at 123 ms, Inworld's TTS 2 at 169 ms and Gradium at 236 ms. The Hugging Face TTS Arena, the third public board, is missing the September leaders, including Sonic 3.6 and Gemini 3.8, and was not used. Our guide to reading voice-AI benchmarks explains why arena votes and vendor demos diverge.
03 — The pricesEvery vendor's price, normalised two ways
Vendors bill in four different units: per thousand characters, per million characters, per million audio tokens, and per UTF-8 byte. The table shows each vendor's own unit first, then our normalisation to dollars per million characters, then a per-minute figure. Where a vendor states its own per-minute rate we use it and say so. Where we derive it, we assume 750 to 1,000 characters per minute of speech, the range implied by Cartesia's and Hume's published ratios, so a $15 per million model costs about $0.011 to $0.015 a minute.
| Model | Vendor's own unit | $ per 1M characters | $ per minute of audio |
|---|---|---|---|
| Google Gemini 3.8 Flash TTS | $0.50/M text tokens in, $9.00/M audio tokens out, through December 31, 2026; $1.00 / $18.00 from January 1, 2027 | 16.49 (Artificial Analysis normalisation; see note) | $0.0135 (vendor: $0.00225 per 10 seconds); $0.027 from 2027 |
| Google Gemini 3.8 Flash-Lite TTS | $0.50 / $6.00 per M tokens through December 31, 2026; $1.00 / $12.00 from 2027 | 11.03 (Artificial Analysis normalisation) | $0.009 (vendor: $0.0015 per 10 seconds); $0.018 from 2027 |
| OpenAI gpt-4o-mini-tts | $0.60/M text tokens in, $12.00/M audio tokens out | not on the Artificial Analysis board | not stated on the pricing page |
| ElevenLabs Eleven v3 | $0.10 per 1,000 characters | 100 | $0.075–0.10 (derived) |
| ElevenLabs v3 Conversational | $0.05 per 1,000 characters | 50 | $0.0375–0.05 (derived) |
| Cartesia Sonic 3.6 | Credit plans: Pro $5 for 100K credits; Scale $299 for 8M | 37.4–50 (derived from plans); 49 on Artificial Analysis | about $0.038 on the Pro plan (derived: about 133 minutes for $5) |
| Inworld Realtime TTS-2 / TTS-2 Flash | $25 / $15 per M characters pay-as-you-go; lower on $25 and $100 plans | 25 / 15 list; 20.83 / 10.42 on Artificial Analysis | $0.019–0.025 / $0.011–0.015 (derived) |
| Microsoft MAI-Voice-2 (Azure Speech) | "starts at $22 USD per 1M characters" | 22 | $0.017–0.022 (derived) |
| Fish Audio S2.1 Pro | $15 per 1M UTF-8 bytes (bytes, not characters) | 15 for Latin script; more for other scripts | $0.011–0.015 (derived) |
| MiniMax Speech 2.8 HD / Turbo | $100 / $60 per M characters | 100 / 60 | $0.075–0.10 / $0.045–0.06 (derived) |
| xAI Grok Voice TTS 1.0 | $15.00 per 1M characters | 15 | $0.011–0.015 (derived) |
| Deepgram Aura-2 | $0.030 per 1,000 characters ($0.027 on Growth) | 30 | $0.0225–0.03 (derived) |
| Deepgram Flux TTS | $0.045 per 1,000 characters from September 13, 2026 | 45 | $0.034–0.045 (derived) |
| Hume Octave 2 | Plan overage $0.15 falling to $0.05 per 1,000 characters | 50–150 | $0.05–0.15 (vendor ratio: about 1,000 characters a minute) |
| Mistral Voxtral TTS | $0.016 per 1,000 characters | 16 | $0.012–0.016 (derived) |
| Kokoro 82M v1.0 (hosted) | $0.62–0.65 per 1M characters on DeepInfra and Replicate | 0.62–0.65 | under $0.001 (derived) |
Three prices carry a date. Google's Gemini API pricing page lists the 3.8 TTS rates as promotional through December 31, 2026, with both input and output doubling on January 1, 2027. Deepgram's Flux TTS was free until September 12 and moved to $0.045 per thousand characters on September 13. And Alibaba's international price for Qwen-Audio-3.0-TTS could not be found on a primary page; the $27.59 figure is Artificial Analysis' endpoint price, and OpenRouter lists the same route at $20. Our Qwen-Audio-3.0-TTS post covers that model's arrival.
04 — The licencesOpen weights: the strongest ones are not for sale
Sixteen of the 92 models publish weights. The licence column comes from each model's Hugging Face tag on September 25, 2026; the commercial column is Artificial Analysis' classification, which is not a legal opinion and in one case conflicts with the tag. The pattern is clear: the open models that score well are non-commercial, and the commercially usable ones score around 1063, about 200 Elo below the leaders.
| Model | Licence | Commercial use | Elo (rank) |
|---|---|---|---|
| Breeze TTS 2 (BreezeBlue, 3.47B) | BreezeBlue research and non-commercial licence | No | 1205 (rank 9) |
| Fish Audio S2 Pro | Other; paid commercial licence | No | 1122 (rank 27) |
| StepFun Step Audio EditX | Not tagged | No | 1095 |
| Mistral Voxtral 4B TTS | CC BY-NC 4.0 | No | 1078 (rank 41) |
| NVIDIA Magpie-Multilingual 357M | Other | Yes | 1063 (rank 48) |
| Kokoro 82M v1.0 | Apache 2.0 | Yes | 1063 (rank 49) |
| Maya1 | Not tagged | Yes | 1044 |
| Chatterbox (Resemble) | MIT | Yes | 1023 (rank 70) |
| Microsoft VibeVoice 1.5B | MIT on Hugging Face; non-commercial per Artificial Analysis | Conflicting | 950 (rank 77) |
| Canopy Labs Orpheus 3B | Apache 2.0 | Not stated | not ranked |
| Sesame CSM-1B | Apache 2.0 | Not stated | not ranked |
| Qwen3-TTS 1.7B CustomVoice | Apache 2.0 | Not stated | not ranked |
Kokoro 82M is the practical answer for a self-hosted or cost-floor deployment: Apache 2.0, 11.7 million Hugging Face downloads, and $0.62 to $0.65 per million characters hosted, which Artificial Analysis names as the cheapest endpoint on its board. Voxtral TTS is the strongest open model by vote from a major lab, but its CC BY-NC licence rules out commercial use without a separate agreement; our Voxtral comparison has the detail.
05 — The routerHow to choose by job
The right model depends on what the audio is for. The router below is our reading of the two boards and the price table; every latency figure in it is Voice Arena's measurement, and every vendor claim is labelled as one.
For a realtime conversational agent the choice is different again, because speech-to-speech models skip the text step. Our voice agent infrastructure reference covers that stack, and our content engine practice builds narration and dubbing pipelines on the models above.
06 — The dutiesWhat you must disclose when the voice is synthetic
Two rules apply to anyone shipping synthetic audio at scale. In the European Union, Article 50 of the AI Act has applied since August 2, 2026: providers of systems that generate synthetic audio must mark outputs in a machine-readable, detectable form, and deployers of deepfakes must disclose them at first exposure. The Commission's Article 50 FAQ, updated July 24, 2026, says that disclosure cannot rely only on embedded marks and must be perceivable, for example through audible labels. Fines run to €15 million or 3% of worldwide turnover. The Digital Omnibus gives systems placed on the market before August 2, 2026 until December 2, 2026 to add the marking.
In the United States, the FCC's declaratory ruling of February 8, 2024 classed AI-generated voices as artificial or prerecorded voices under the Telephone Consumer Protection Act, so an outbound call using any model in this table needs prior express consent.
Google says every audio clip from its Gemini audio models is watermarked with SynthID, and that replicated voices also carry C2PA content credentials and require a recorded verbal consent. OpenAI's custom voices need a consent phrase and are limited to eligible customers. Those marks address the machine-readable half of Article 50; the perceivable disclosure to a listener is still the deployer's job.
A snapshot of two independent listener-vote leaderboards and the vendors' own price pages, with every normalisation shown.
- What was collected
- The full Artificial Analysis text-to-speech board (92 models, Elo, 95% confidence interval, votes, release date, open-weight flag, normalised price); the Voice Arena US-English board (Elo, interval rank, measured median latency, votes); each vendor's list price in its own unit; and the Hugging Face licence tag of every open-weight model.
- Sources
- Artificial Analysis raw page data; Voice Arena's rendered table and site bundle; vendor pricing pages and launch posts; the Hugging Face API; the European Commission's Article 50 FAQ and the FCC's ruling. The Hugging Face TTS Arena was read and excluded as out of date.
- As-of date
- Leaderboards and price pages read on September 25, 2026.
- Normalisation
- Dollars per million characters equals the vendor's per-thousand price times 1,000. Per-minute figures marked derived assume 750 to 1,000 characters a minute; vendor-stated per-minute rates are used where they exist. Google's per-character figures are Artificial Analysis' and are flagged in the text.
- Known limitations
- Not found on a primary page: Alibaba's international TTS price, Microsoft's MAI-Voice-2-Flash price, and any ElevenLabs "v4" model. Artificial Analysis publishes no time-to-first-audio for TTS. Release dates for Qwen-Audio-3.0-TTS differ between sources.
- Refresh
- Refreshed in place when either board's top five changes or a listed price moves; Google's promotional rates end December 31, 2026.
08 — ConclusionThe top three are a statistical tie; price and job break it
Shortlist Sonic 3.6, Gemini 3.8 Flash TTS and Simba 3.2, run your own text through all three, and book the January price change now
Listener votes settle the shortlist, not the choice. A hundred sentences from your own product, scored by your own team, will separate three models the arenas cannot.