AI DevelopmentPricing Tracker9 min readPublished September 24, 2026

92 models · two blind arenas · 16 open-weight · the leader costs three times the runner-up

Best Text-to-Speech Models, September 2026: Ranked, Priced

92 text-to-speech models ranked by blind listener vote, with the price per million characters, open-weight licences and the job each leader fits.

DA
Digital Applied Team
Research and practical guidance
PublishedSeptember 24, 2026
Boards readSeptember 25, 2026

By listener preference on the Artificial Analysis Speech Arena, read September 25, 2026, the best text-to-speech model is Cartesia's Sonic 3.6 at 1279 Elo. Second is Google's Gemini 3.8 Flash TTS, released this week, at 1265 and about a third of the price. On the other independent board, Voice Arena's US-English leaderboard, the cheaper Gemini 3.8 Flash-Lite TTS holds first place. The confidence intervals of the top four on Artificial Analysis overlap, so "best" is a tie broken by price and by the job you need done.

This page is the dated table: 92 models ranked by blind pairwise vote, every vendor's price normalised to dollars per million characters and per minute of audio, the licences of the 16 open-weight models, and a router by job. It is a snapshot. Both boards move weekly, and Google's launch prices double on January 1, 2027.

Key takeaways
  1. 01
    Sonic 3.6 leads on votes; Gemini 3.8 Flash TTS is close behind at a third of the price.Artificial Analysis lists Sonic 3.6 at $49 per million characters and Gemini 3.8 Flash TTS at $16.49, with a 14-point Elo gap inside both models' 17-point confidence intervals.
  2. 02
    The two independent boards disagree at the top, and both are right.Voice Arena's US-English board, built from real production-call prompts, puts Gemini 3.8 Flash-Lite first by confidence-interval rank. Artificial Analysis, across all categories and accents, puts it sixth. Test on your own text.
  3. 03
    The strongest open-weight models cannot be used commercially.Breeze TTS 2 (rank 9) is research and non-commercial; Voxtral TTS is CC BY-NC. The best commercially usable open models, NVIDIA's Magpie and Kokoro 82M, sit near rank 48 at 1063 Elo.
  4. 04
    Vendor latency figures are claims; of the two boards, only Voice Arena measures them.Artificial Analysis publishes no time-to-first-audio for TTS right now. Voice Arena's measured medians run from 123 ms (Simba 3.2) to 890 ms (Gemini 3.1 Flash TTS). Cartesia's "sub-90ms" and ElevenLabs' "~75ms" are vendor statements.

01 — The tableThe ranked table: 22 of 92 models by listener vote

The Artificial Analysis text-to-speech leaderboard ranks models by Elo from blind pairwise votes in its Speech Arena, with a 95% confidence interval, across all prompt categories and both US and UK accents. The page carries no as-of date, so we date the snapshot to the day we parsed it. Price is Artificial Analysis' normalised figure in dollars per million characters at the endpoint it tracks. The table stops at rank 22, at 1130 Elo, tied with ranks 23 and 24; other vendors appear in the price table.

Source: Artificial Analysis Speech Arena, text-to-speech leaderboard, parsed September 25, 2026. Elo with 95% confidence interval; price in USD per million characters as normalised by Artificial Analysis.
ModelElo ± CIVotes$ per 1M chars
1. Sonic 3.6 (Cartesia)1279 ± 171,79649.00
2. Gemini 3.8 Flash TTS (Google)1265 ± 172,04416.49
3. Qwen-Audio-3.0-TTS-Plus (Alibaba)1259 ± 161,47727.59
4. Realtime TTS-2 (Inworld)1246 ± 181,26920.83
5. Simba 3.2 (Speechify)1239 ± 142,4326.58
6. Gemini 3.8 Flash-Lite TTS (Google)1239 ± 162,05611.03
7. Luna TTS (VUI Labs)1229 ± 142,58115.00
8. Realtime TTS-2 Flash (Inworld)1210 ± 151,75110.42
9. Breeze TTS 2 (BreezeBlue, open weights)1205 ± 161,44034.00
10. Gemini 3.1 Flash TTS (Google)1202 ± 123,48418.31
11. StepAudio 2.5 TTS (StepFun)1199 ± 161,34485.00
12. v3 Conversational (ElevenLabs)1196 ± 151,88950.00
13. Sonic 3.5 (Cartesia)1185 ± 123,03949.00
14. Lightning V3.1 Pro (Smallest.ai)1174 ± 141,97219.50
15. TTS Real-Time v2 (Soniox)1173 ± 124,67114.23
16. Speech 2.8 HD (MiniMax)1170 ± 114,460100.00
17. Eleven v3 (ElevenLabs)1169 ± 114,390100.00
18. Falcon 2 (Murf AI)1159 ± 151,78510.00
19. Speech 2.8 Turbo (MiniMax)1149 ± 114,37060.00
20. Gradium TTS (Gradium)1149 ± 141,72647.22
21. S2.1 Pro (Fish Audio)1140 ± 132,22615.00
22. SpaceXAI TTS (xAI)1130 ± 151,25815.00

Read the intervals before the ranks. Sonic 3.6 at 1279 ± 17 and Gemini 3.8 Flash TTS at 1265 ± 17 overlap; so do Qwen-Audio-3.0-TTS Plus at 1259 ± 16 and Inworld's Realtime TTS-2 at 1246 ± 18. Simba 3.2 and Gemini 3.8 Flash-Lite TTS share 1239 exactly. Across the top six, price varies far more than Elo, a more than seven-fold spread: Simba 3.2 at $6.58 per million characters against Sonic 3.6 at $49. Artificial Analysis' own frontier note names three models on the quality-versus-price line: Sonic 3.6, Gemini 3.8 Flash TTS and Simba 3.2.

What the top ten cost, in dollars per million characters

Artificial Analysis normalised endpoint prices, September 25, 2026. Bars scaled to Sonic 3.6 at $49.
Sonic 3.6rank 1
$49.00
Breeze TTS 2rank 9, non-commercial
$34.00
Qwen-Audio-3.0-TTS-Plusrank 3
$27.59
Realtime TTS-2rank 4
$20.83
Gemini 3.1 Flash TTSrank 10
$18.31
Gemini 3.8 Flash TTSrank 2
$16.49
Luna TTSrank 7
$15.00
Gemini 3.8 Flash-Lite TTSrank 6
$11.03
Realtime TTS-2 Flashrank 8
$10.42
Simba 3.2rank 5
$6.58

One normalisation deserves caution. Google bills its TTS models per audio token, not per character. Gemini 3.8 Flash TTS costs $9 per million output tokens and Gemini 3.1 Flash TTS costs $20, a 2.2× gap, yet Artificial Analysis' per-character figures for the two are $16.49 and $18.31, only 1.1× apart. We print the Artificial Analysis numbers for consistency and Google's own per-minute rate in the price table. Our post on the Gemini 3.8 Flash TTS launch works through the token arithmetic.

02 — The second boardVoice Arena's US-English board, the one Google cited

Google's launch post cited Voice Arena, whose US-English text-to-speech leaderboard uses Bradley-Terry Elo on prompts it says are built verbatim from real production calls. Its ranks come with a confidence-interval range, shown in brackets, and unlike Artificial Analysis it measures median time to first audio itself rather than repeating vendor claims. Rows where it has not yet measured latency say so.

Source: Voice Arena, TTS leaderboard, US English, rendered September 25, 2026. Rank shows the confidence-interval range where the site publishes one. Latency is Voice Arena's measured median time to first audio.
Rank and modelElo ±Time to first audio, P50Votes
1 (1–4) · Gemini 3.8 Flash-Lite TTS1087 ± 15not measured636
2 (1–5) · Cartesia Sonic-3.61068 ± 13341 ms1,004
3 (1–5) · Gemini 3.8 Flash TTS1061 ± 12not measured1,161
4 (1–6) · Inworld Realtime TTS 2 (preview)1058 ± 15169 ms849
5 (2–5) · Gemini 3.1 Flash TTS1057 ± 6890 ms6,123
6 (6–8) · Cartesia Sonic-3.51033 ± 7250 ms4,049
7 (6–8) · Speechify Simba 3.21032 ± 6123 ms4,534
8 (5–8) · ElevenLabs v3 Conversational1031 ± 15not measured837
9 (9–14) · xAI Grok TTS996 ± 6354 ms6,019
10 (9–14) · Murf Falcon 2993 ± 11528 ms1,550
11 (9–16) · Fish Audio S2.1 Pro992 ± 16283 ms802
12 (9–16) · Maya-2-Global988 ± 9not measured1,662
13 (9–16) · ElevenLabs v3986 ± 6588 ms5,413
14 (9–16) · Gradium TTS985 ± 7236 ms3,225
15 (11–17) · Microsoft Azure Dragon HD Omni975 ± 6485 ms5,861
16 (11–17) · Hithink Speech 2.6971 ± 10not measured1,479
17 (15–17) · Smallest Lightning 3.1 Pro962 ± 10264 ms1,589
18 (18–19) · Fish Audio S2 Pro939 ± 7267 ms5,290
19 (18–19) · OpenAI gpt-4o-mini-tts929 ± 7812 ms5,205
20 · Cartesia Sonic-3858 ± 10not measured3,200

The two boards agree on the cluster and disagree on the order. Sonic 3.6, Gemini 3.8 Flash TTS and Inworld's TTS-2 are top-four on both; Simba 3.2 is fifth on one and seventh on the other. The largest divergence at the top is Gemini 3.8 Flash-Lite TTS, first here by confidence-interval rank and sixth on Artificial Analysis. Its latency is not yet measured on Voice Arena, so the fastest measured models on this board are Simba 3.2 at 123 ms, Inworld's TTS 2 at 169 ms and Gradium at 236 ms. The Hugging Face TTS Arena, the third public board, is missing the September leaders, including Sonic 3.6 and Gemini 3.8, and was not used. Our guide to reading voice-AI benchmarks explains why arena votes and vendor demos diverge.

03 — The pricesEvery vendor's price, normalised two ways

Vendors bill in four different units: per thousand characters, per million characters, per million audio tokens, and per UTF-8 byte. The table shows each vendor's own unit first, then our normalisation to dollars per million characters, then a per-minute figure. Where a vendor states its own per-minute rate we use it and say so. Where we derive it, we assume 750 to 1,000 characters per minute of speech, the range implied by Cartesia's and Hume's published ratios, so a $15 per million model costs about $0.011 to $0.015 a minute.

Source: each vendor's pricing page or launch post, read September 25, 2026, and Artificial Analysis where marked. Per-minute figures marked derived assume 750 to 1,000 characters per minute.
ModelVendor's own unit$ per 1M characters$ per minute of audio
Google Gemini 3.8 Flash TTS$0.50/M text tokens in, $9.00/M audio tokens out, through December 31, 2026; $1.00 / $18.00 from January 1, 202716.49 (Artificial Analysis normalisation; see note)$0.0135 (vendor: $0.00225 per 10 seconds); $0.027 from 2027
Google Gemini 3.8 Flash-Lite TTS$0.50 / $6.00 per M tokens through December 31, 2026; $1.00 / $12.00 from 202711.03 (Artificial Analysis normalisation)$0.009 (vendor: $0.0015 per 10 seconds); $0.018 from 2027
OpenAI gpt-4o-mini-tts$0.60/M text tokens in, $12.00/M audio tokens outnot on the Artificial Analysis boardnot stated on the pricing page
ElevenLabs Eleven v3$0.10 per 1,000 characters100$0.075–0.10 (derived)
ElevenLabs v3 Conversational$0.05 per 1,000 characters50$0.0375–0.05 (derived)
Cartesia Sonic 3.6Credit plans: Pro $5 for 100K credits; Scale $299 for 8M37.4–50 (derived from plans); 49 on Artificial Analysisabout $0.038 on the Pro plan (derived: about 133 minutes for $5)
Inworld Realtime TTS-2 / TTS-2 Flash$25 / $15 per M characters pay-as-you-go; lower on $25 and $100 plans25 / 15 list; 20.83 / 10.42 on Artificial Analysis$0.019–0.025 / $0.011–0.015 (derived)
Microsoft MAI-Voice-2 (Azure Speech)"starts at $22 USD per 1M characters"22$0.017–0.022 (derived)
Fish Audio S2.1 Pro$15 per 1M UTF-8 bytes (bytes, not characters)15 for Latin script; more for other scripts$0.011–0.015 (derived)
MiniMax Speech 2.8 HD / Turbo$100 / $60 per M characters100 / 60$0.075–0.10 / $0.045–0.06 (derived)
xAI Grok Voice TTS 1.0$15.00 per 1M characters15$0.011–0.015 (derived)
Deepgram Aura-2$0.030 per 1,000 characters ($0.027 on Growth)30$0.0225–0.03 (derived)
Deepgram Flux TTS$0.045 per 1,000 characters from September 13, 202645$0.034–0.045 (derived)
Hume Octave 2Plan overage $0.15 falling to $0.05 per 1,000 characters50–150$0.05–0.15 (vendor ratio: about 1,000 characters a minute)
Mistral Voxtral TTS$0.016 per 1,000 characters16$0.012–0.016 (derived)
Kokoro 82M v1.0 (hosted)$0.62–0.65 per 1M characters on DeepInfra and Replicate0.62–0.65under $0.001 (derived)

Three prices carry a date. Google's Gemini API pricing page lists the 3.8 TTS rates as promotional through December 31, 2026, with both input and output doubling on January 1, 2027. Deepgram's Flux TTS was free until September 12 and moved to $0.045 per thousand characters on September 13. And Alibaba's international price for Qwen-Audio-3.0-TTS could not be found on a primary page; the $27.59 figure is Artificial Analysis' endpoint price, and OpenRouter lists the same route at $20. Our Qwen-Audio-3.0-TTS post covers that model's arrival.

04 — The licencesOpen weights: the strongest ones are not for sale

Sixteen of the 92 models publish weights. The licence column comes from each model's Hugging Face tag on September 25, 2026; the commercial column is Artificial Analysis' classification, which is not a legal opinion and in one case conflicts with the tag. The pattern is clear: the open models that score well are non-commercial, and the commercially usable ones score around 1063, about 200 Elo below the leaders.

Source: Hugging Face model-card licence tags and the Artificial Analysis commercial-use flag, read September 25, 2026. Elo and rank from the Artificial Analysis board on the same date.
ModelLicenceCommercial useElo (rank)
Breeze TTS 2 (BreezeBlue, 3.47B)BreezeBlue research and non-commercial licenceNo1205 (rank 9)
Fish Audio S2 ProOther; paid commercial licenceNo1122 (rank 27)
StepFun Step Audio EditXNot taggedNo1095
Mistral Voxtral 4B TTSCC BY-NC 4.0No1078 (rank 41)
NVIDIA Magpie-Multilingual 357MOtherYes1063 (rank 48)
Kokoro 82M v1.0Apache 2.0Yes1063 (rank 49)
Maya1Not taggedYes1044
Chatterbox (Resemble)MITYes1023 (rank 70)
Microsoft VibeVoice 1.5BMIT on Hugging Face; non-commercial per Artificial AnalysisConflicting950 (rank 77)
Canopy Labs Orpheus 3BApache 2.0Not statednot ranked
Sesame CSM-1BApache 2.0Not statednot ranked
Qwen3-TTS 1.7B CustomVoiceApache 2.0Not statednot ranked

Kokoro 82M is the practical answer for a self-hosted or cost-floor deployment: Apache 2.0, 11.7 million Hugging Face downloads, and $0.62 to $0.65 per million characters hosted, which Artificial Analysis names as the cheapest endpoint on its board. Voxtral TTS is the strongest open model by vote from a major lab, but its CC BY-NC licence rules out commercial use without a separate agreement; our Voxtral comparison has the detail.

05 — The routerHow to choose by job

The right model depends on what the audio is for. The router below is our reading of the two boards and the price table; every latency figure in it is Voice Arena's measurement, and every vendor claim is labelled as one.

A voice agent where the first syllable matters
Measured medians: Simba 3.2 at 123 ms, Inworld Realtime TTS 2 at 169 ms, Gradium at 236 ms. Cartesia claims sub-90 ms, but Voice Arena measures Sonic-3.6 at 341 ms; ElevenLabs Flash claims about 75 ms and is not on Voice Arena.
Simba 3.2 or Inworld TTS-2
Long-form narration or audiobooks, quality first
Sonic 3.6 leads by vote at $49 per million; Gemini 3.8 Flash TTS is within its interval at $16.49 through December. Both boards agree on the pair.
Sonic 3.6 or Gemini 3.8 Flash TTS
Many languages from one endpoint
Google's model page lists 130 languages for Gemini 3.8 Flash TTS and 101 for Flash-Lite; ElevenLabs states 70+; Fish Audio 83. Vendor counts, not tested.
Gemini 3.8 Flash TTS
High volume, cost floor, quality acceptable
Simba 3.2 at $6.58 per million is the cheapest top-six model; Kokoro 82M at $0.65 hosted or self-hosted is the floor, roughly 175 Elo lower.
Simba 3.2, or Kokoro 82M
On-premises, commercial use, no vendor API
Kokoro 82M (Apache 2.0) or NVIDIA Magpie (commercial per Artificial Analysis). Breeze TTS 2 and Voxtral score higher but their licences exclude commercial use.
Kokoro 82M or Magpie
Cloning a specific person's voice
Google requires a recorded verbal consent and marks the output with SynthID and C2PA; OpenAI gates custom voices to eligible customers; Microsoft gates MAI-Voice-2 cloning. Cartesia, Inworld, xAI and Hume offer it on plan.
Depends on consent process

For a realtime conversational agent the choice is different again, because speech-to-speech models skip the text step. Our voice agent infrastructure reference covers that stack, and our content engine practice builds narration and dubbing pipelines on the models above.

06 — The dutiesWhat you must disclose when the voice is synthetic

Two rules apply to anyone shipping synthetic audio at scale. In the European Union, Article 50 of the AI Act has applied since August 2, 2026: providers of systems that generate synthetic audio must mark outputs in a machine-readable, detectable form, and deployers of deepfakes must disclose them at first exposure. The Commission's Article 50 FAQ, updated July 24, 2026, says that disclosure cannot rely only on embedded marks and must be perceivable, for example through audible labels. Fines run to €15 million or 3% of worldwide turnover. The Digital Omnibus gives systems placed on the market before August 2, 2026 until December 2, 2026 to add the marking.

In the United States, the FCC's declaratory ruling of February 8, 2024 classed AI-generated voices as artificial or prerecorded voices under the Telephone Consumer Protection Act, so an outbound call using any model in this table needs prior express consent.

What the vendors do for you

Google says every audio clip from its Gemini audio models is watermarked with SynthID, and that replicated voices also carry C2PA content credentials and require a recorded verbal consent. OpenAI's custom voices need a consent phrase and are limited to eligible customers. Those marks address the machine-readable half of Article 50; the perceivable disclosure to a listener is still the deployer's job.

Methodology

A snapshot of two independent listener-vote leaderboards and the vendors' own price pages, with every normalisation shown.

What was collected
The full Artificial Analysis text-to-speech board (92 models, Elo, 95% confidence interval, votes, release date, open-weight flag, normalised price); the Voice Arena US-English board (Elo, interval rank, measured median latency, votes); each vendor's list price in its own unit; and the Hugging Face licence tag of every open-weight model.
Sources
Artificial Analysis raw page data; Voice Arena's rendered table and site bundle; vendor pricing pages and launch posts; the Hugging Face API; the European Commission's Article 50 FAQ and the FCC's ruling. The Hugging Face TTS Arena was read and excluded as out of date.
As-of date
Leaderboards and price pages read on September 25, 2026.
Normalisation
Dollars per million characters equals the vendor's per-thousand price times 1,000. Per-minute figures marked derived assume 750 to 1,000 characters a minute; vendor-stated per-minute rates are used where they exist. Google's per-character figures are Artificial Analysis' and are flagged in the text.
Known limitations
Not found on a primary page: Alibaba's international TTS price, Microsoft's MAI-Voice-2-Flash price, and any ElevenLabs "v4" model. Artificial Analysis publishes no time-to-first-audio for TTS. Release dates for Qwen-Audio-3.0-TTS differ between sources.
Refresh
Refreshed in place when either board's top five changes or a listed price moves; Google's promotional rates end December 31, 2026.

08 — ConclusionThe top three are a statistical tie; price and job break it

What to do

Shortlist Sonic 3.6, Gemini 3.8 Flash TTS and Simba 3.2, run your own text through all three, and book the January price change now

Listener votes settle the shortlist, not the choice. A hundred sentences from your own product, scored by your own team, will separate three models the arenas cannot.

Digital Applied

Put a voice on your content without guessing at the model.

We build narration, dubbing and voice-agent pipelines on the models in this table, with the licence, consent and disclosure steps handled before launch.

Model shortlistsConsent workflowsDisclosure checks
Your next project

A voice pipeline with a known cost per minute

  • →Three models tested on your text
  • →A price per minute you can budget
  • →Article 50 marking in place
Questions and answers

The questions we get about text-to-speech models

By listener preference on the Artificial Analysis Speech Arena, read September 25, 2026, Cartesia's Sonic 3.6 at 1279 Elo, with Gemini 3.8 Flash TTS at 1265 inside the same confidence interval. On Voice Arena's US-English board Gemini 3.8 Flash-Lite TTS is first by confidence-interval rank. The top three are a statistical tie on Artificial Analysis' rank ranges.
Digital Applied newsletter

Deep dives on AI, marketing and development.

Practical guides and fresh insights by email. No recycled takes.

Related dispatches

Continue reading

AI Development

Gemini 3.8 Flash TTS: Voice Cloning and a Price That Doubles

Gemini 3.8 Flash TTS is generally available with voice cloning from a 30-second sample. The promotional price ends December 31 and doubles on January 1, 2027.

September 23, 2026 · 5 minRead
AI Development

Voice AI Benchmarks and Listener Votes Disagree: 14 Models

Artificial Analysis scores realtime voice systems two ways. The arena system with the best reasoning and turn-taking scores ranks last on listener preference.

September 25, 2026 · 6 minRead
AI Development

Fireworks Ember-1: Kimi K3 Quality With Fewer Tokens?

Fireworks tuned Kimi K3 into Ember-1 and says it matches K3 with 35 to 50% shorter reasoning at the same price. The rows it loses, and the preview caveat.

September 23, 2026 · 5 minRead
AI Development

What AI Labs Actually Disclose About Training Costs

A census of 21 training-cost figures labs published themselves, 2022 to 2026. Six give a dollar amount, none is audited, and each leaves something out.

September 22, 2026 · 8 minRead
AI Development

AI Search Agents Compared: Google, Perplexity, ChatGPT

Google's always-on information agents, Perplexity Pro, and ChatGPT Search compared. Which AI search agent delivers the best research results in 2026?

May 20, 2026 · 14 minRead
AI Development

OpenAI + Dell Codex: On-Premises Enterprise Agents

OpenAI and Dell partner to bring Codex to hybrid and on-premises environments via Dell AI Factory. What changes for enterprise coding workflows.

May 18, 2026 · 12 minRead