AI DevelopmentMethodology6 min readPublished September 25, 2026

Two boards, one vendor, read September 25 · the best-scoring arena system is the least liked

Voice AI Benchmarks and Listener Votes Disagree: 14 Models

Artificial Analysis scores realtime voice systems two ways. The arena system with the best reasoning and turn-taking scores ranks last on listener preference.

DA
Digital Applied Team
Research and practical guidance
PublishedSeptember 25, 2026
Data readSeptember 25, 2026

Artificial Analysis publishes two ways of ranking realtime voice systems, the models that listen and talk back without a separate transcription step. One is an arena: listeners hear two systems answer the same prompt and vote, producing an Elo score. The other is a benchmark set: a speech-reasoning test, a turn-taking test, a time-to-first-audio measurement and a cost per hour. Read on September 25, 2026, the two boards disagree about which systems are good, and the disagreement is not noise.

Of the systems on both boards, the one with the best benchmark scores, Qwen Audio 3.0 Realtime Plus, ranks last on preference, 26th of 26. OpenAI's GPT-Realtime-2.1 outscores the older GPT-Realtime-1.5 on Big Bench Audio (95.8% against 81.4%), ties it on Full Duplex Bench and loses to it on preference. This post joins the two tables, says what each one measures, marks where the confidence intervals make a ranking meaningless, and ends with how to choose a model for a voice agent when the boards point in different directions.

Key takeaways
  1. 01
    Qwen Audio 3.0 Realtime Plus tops the benchmarks among arena systems and sits last on preference.Big Bench Audio 99.2%, Full Duplex Bench 98.4%, the highest of any arena system, behind only StepAudio 3 Realtime (99.7% and 98.9%), which has no arena score. Arena Elo 775 with an interval of 731 to 819, below every other listed system, on 171 votes.
  2. 02
    GPT-Realtime-2.1 High scores below the 1.5 anchor on preference and far above it on Big Bench Audio.Elo 928 (902 to 954) against the anchor's 1000. It also costs $10.75 per hour of input audio on Artificial Analysis' measure, the highest in our table, though the 1.5 anchor itself is listed at $11.44.
  3. 03
    The top of the preference board is a statistical tie.Gemini 3.1 Flash Live (Minimal) at 1096 and Gemini 3.8 Live at 1083 have overlapping intervals, as do the next three. Elo differences inside overlapping intervals are not rankings, and the table marks the intervals so you can see which ones are.
  4. 04
    Task completion is a third measure, and it disagrees with both.Across the fourteen listed rows, task completion ranges from 57.1% to 94.6%, and the preference leader completes 74.6% of tasks against 93.2% for the system one place below it.

01 — The findingThe split, in one pair of arena systems

Best on the benchmarks

Qwen Audio 3.0 Realtime Plus

Alibaba · native speech-to-speech
Big Bench Audio 99.2% · Full Duplex Bench 98.4% · arena Elo 775 (731 to 819, 171 votes) · task completion 77.8% · time to first audio 1.54 s · $4.42 per hour of input audio.
Best on preference

Gemini 3.1 Flash Live (Minimal)

Google · native speech-to-speech
Arena Elo 1096 (1068 to 1123, 841 votes) · task completion 74.6% · Big Bench Audio 71.3% · Full Duplex Bench 72.3% · time to first audio 0.96 s · $1.50 per hour of input audio.

Both columns are from Artificial Analysis' speech-to-speech page, read on September 25, 2026. The preference leader scores 71.3% on Big Bench Audio and 72.3% on Full Duplex Bench, and the arena system with the best benchmark scores has the lowest preference score. That is the whole problem: a buyer reading either board alone would choose a system the other board says is the wrong one.

02 — The dataThe preference board: Elo, interval, votes, task completion

The arena anchors GPT-Realtime-1.5 at 1000 and scores every other system relative to it. The interval column is the 95% confidence interval Artificial Analysis publishes; where two intervals overlap, the two systems are not distinguishable on preference. Systems labelled cascaded run a separate transcription, text model and speech step rather than one native model, which Artificial Analysis marks on the board and we carry over. Rank numbers are the board's own and skip where our read of the board did not capture a row.

Artificial Analysis speech-to-speech arena, read September 25, 2026. Elo relative to GPT-Realtime-1.5 at 1000; 95% confidence intervals as published.
SystemElo95% intervalTask completion
1. Gemini 3.1 Flash Live (Minimal)10961068–112374.6%
2. Gemini 3.8 Live10831053–111393.2%
3. Gemini 3.1 Flash Live (High)10631036–109071.8%
4. GPT-Live-1 (Sol, low)10531026–108190.9%
5. GPT-Live-1 (Astra, medium)10481020–107687.4%
6. Cartesia Line (cascaded)1027990–106477.5%
7. Grok Voice Think Fast 2.0 High1011983–103894.6%
8. GPT-Realtime-1.5 (anchor)1000anchor85.1%
9. ElevenLabs Agents (cascaded)993967–102090.5%
10. Gemini 3.8 Live Extended Thinking (High)990963–101789.1%
12. Amazon Nova 2.0 Sonic (Mar 2026)974948–100057.1%
15. GPT-Realtime-2.1 High928902–95491.5%
18. Deepgram Voice Agent (cascaded)914888–94073.7%
26. Qwen Audio 3.0 Realtime Plus775731–81977.8%

Read the intervals before the ranks. Positions one to five span 1048 to 1096 and every adjacent pair overlaps, so the honest statement is that five systems share the top of the board, and the sixth and seventh overlap the fifth. The one gap that is unmistakable is Qwen's: its upper bound of 819 sits below the lower bound of every other row in this table, though three uncaptured systems at ranks 23 to 25 overlap it. Task completion tells a different story again. The system listeners prefer most completes fewer tasks than ten of the thirteen systems below it, and Grok Voice Think Fast 2.0 High, seventh on preference, has the best completion rate in the table. Our post on Grok Voice Think Fast 2 covers what xAI built it for.

03 — The dataThe benchmark board: reasoning, turn-taking, latency, cost

Big Bench Audio tests whether the system reasons correctly about what it heard. Full Duplex Bench tests turn-taking: whether the system yields, interrupts and resumes the way a conversation needs. Time to first audio is the seconds until the system starts speaking. The cost column is Artificial Analysis' figure for an hour of input audio on the vendor's list price. Scores are as published; a dash means the page lists no value.

Artificial Analysis speech-to-speech benchmarks, read September 25, 2026. BBA is Big Bench Audio; FDB is Full Duplex Bench; TTFA is time to first audio in seconds; cost is per hour of input audio.
SystemBBA / FDBTTFA (s)$ per hour
StepAudio 3 Realtime99.7% / 98.9%8.83–
Qwen Audio 3.0 Realtime Plus99.2% / 98.4%1.54$4.42
Gemini 3.8 Live Extended Thinking (High)97.7% / 91.9%1.35$3.50
Grok Voice Think Fast 2.0 High97.2% / 95.1%0.70$4.80
GPT-Realtime-2 (High)96.6% / 95.3%1.14$4.14
GPT-Realtime-2.1 High95.8% / 95.7%1.21$10.75
Gemini 3.8 Live91.7% / 96.1%1.18$0.84
GPT-Live-1 (Astra, medium)90.1% / 94.9%1.34$5.83
GPT-Live-1 (Sol, low)89.0% / 97.3%1.24$4.47
Amazon Nova 2.0 Sonic (Mar 2026)88.1% / –1.14$4.90
Deepslate Opal85.0% / 85.7%0.44$6.48
Raon SpeechChat (Krafton)57.5% / 86.2%0.04–
Moshi (open)4.4% / 61.0%––

Cost per hour of input audio, systems in our benchmark table

Artificial Analysis, read September 25, 2026. Bars scaled to GPT-Realtime-2.1 High at $10.75.
GPT-Realtime-2.1 HighOpenAI
$10.75
Deepslate OpalDeepslate
$6.48
GPT-Live-1 (Astra, medium)OpenAI
$5.83
Nova 2.0 SonicAmazon
$4.90
Grok Voice Think Fast 2.0 HighxAI
$4.80
GPT-Live-1 (Sol, low)OpenAI
$4.47
Qwen Audio 3.0 Realtime PlusAlibaba
$4.42
GPT-Realtime-2 (High)OpenAI
$4.14
Gemini 3.8 Live Extended ThinkingGoogle
$3.50
Gemini 3.8 LiveGoogle
$0.84

Two rows are worth a second look. Gemini 3.8 Live scores 91.7% on reasoning, 96.1% on turn-taking, starts speaking in 1.18 seconds, ranks second on preference with the second-best task completion, and costs $0.84 an hour, less than a quarter of the next cheapest listed system. And StepAudio 3 Realtime, top on both benchmarks at 99.7% and 98.9%, takes 8.83 seconds to begin speaking, which disqualifies it from a live conversation whatever it scores. Our post on Gemini 3.8 Live and its extended-thinking variant explains the split between the two Google rows.

04 — The reasonWhy a benchmark and a listener disagree

The two boards measure different things and neither is wrong. A benchmark scores an answer against a key: did the system reason correctly about the audio, did it yield the turn when it should. A listener votes on an experience: did it sound like a person, did the voice fit, did it feel quick, did it say too much. A system can be correct and unpleasant, or pleasant and wrong, and the boards show both cases. The turn-taking benchmark rewards a system that interrupts at exactly the right moment; a listener may prefer one that waits a beat too long and sounds polite.

The arena also has a sample problem the intervals make visible. Qwen's 171 votes produce an interval 88 points wide; the top-ranked system's 841 votes produce one 55 points wide. A new system enters with a wide interval and a small vote count, and its position moves as votes arrive. That is why we quote the interval with every Elo, why a difference inside overlapping intervals is not a ranking, and why the board should be re-read rather than remembered.

Cascaded systems sit on both boards with their own trade-off. A pipeline of transcription, text model and speech synthesis can reason with a stronger text model than any native voice model carries, and it pays for that in latency and in turn-taking, because the pieces do not share a clock. Our reference on voice-agent latency measures defines time to first audio and the other timings the benchmark board uses.

What neither board measures

Your callers, your accents, your background noise, your task. Both boards use their own prompts and their own audio. A voice agent for a contact centre in one region, with one product vocabulary, can rank the same six systems in a different order from either board, and usually does.

05 — The decisionHow to choose when the boards disagree

The agent has to complete a defined task: book, verify, route, resolve
Weight task completion first, then Full Duplex Bench, then preference. A caller who got what they came for forgives a flat voice; a caller who did not does not care how natural it sounded. Grok Voice Think Fast 2.0 High, Gemini 3.8 Live and GPT-Realtime-2.1 High lead this cut.
Completion first
The agent is the brand's voice and the task is open conversation
Weight preference first, but only across systems whose intervals separate. Test the top five as a group with your own script rather than trusting the order, and check task completion so the preferred voice is not the one that fails the caller.
Preference first
Volume is high and margins are thin
Start from cost per hour and work up. Gemini 3.8 Live at $0.84 costs about a quarter of the next-cheapest system in our table while ranking second on preference and completion; confirm the price on Google's own page, because Artificial Analysis normalises list prices into its own per-hour figure.
Cost first
Any of the above, before signing
Record 50 real calls with consent, replay them to each candidate, score completion and listen to the audio. That is the only board built on your callers, and it takes a day.
Your own recordings

Prices for the vendor-side comparison are in our text-to-speech ranking for the synthesis half of a cascaded stack and in our post on Gemini 3.8 Flash TTS for the January price change. Our AI transformation practice runs the 50-call test as the first week of any voice-agent build.

Methodology

A snapshot of one independent evaluator's two boards, joined by system name, with the evaluator's own intervals and units carried through.

What was collected
The Artificial Analysis speech-to-speech arena (Elo, 95% confidence interval, vote count, task completion) and its benchmark table (Big Bench Audio, Full Duplex Bench, time to first audio, cost per hour of input audio), parsed from the page's raw data. Fourteen arena rows and thirteen benchmark rows are shown; eight systems appear on both.
Sources
Artificial Analysis only, for every score. Vendor pricing pages were read for the cost cross-checks in section 05 and are cited in the linked posts. No score in this post was measured by us.
As-of date
Both boards read on September 25, 2026.
Exclusions
Arena rows at ranks 11, 13, 14, 16, 17 and 19 to 25 were not captured in our read of the board and are omitted; their ranks are skipped in the table so the board's numbering is preserved. Text-to-speech models are covered in a separate post. Hugging Face and OpenRouter voice boards were not used.
Known limitations
Elo positions with overlapping intervals are not rankings. Cascaded and native systems are scored on the same boards despite different architectures. The preference leader's benchmark scores (71.3% / 72.3%) are quoted in section 01 but not carried in the benchmark table. Cost per hour is the evaluator's normalisation, not a vendor quote.
Refresh
Refreshed in place when the top five on either board changes or a new native system enters both boards.

06 — ConclusionCorrect and liked are different scores, and the boards prove it

What to do

Pick the board that matches the job, read the intervals before the ranks, and test the top group on your own recorded calls before you sign

The best benchmark scorer on the arena is the least preferred system on it, and the preferred leader scores 71.3% and 72.3% on the two benchmarks. Neither board can choose for you. Fifty of your own calls can.

Digital Applied

Choose a voice model on your callers, not a leaderboard.

We run recorded-call evaluations across candidate voice models, score completion and listener preference on your own audio, and build the agent on the one that wins both.

Recorded-call evaluationsModel selectionVoice-agent builds
Your next project

A voice model chosen on evidence

  • →50 real calls replayed to every candidate
  • →Completion and preference scored on your audio
  • →Cost per hour confirmed on the vendor's page
Questions and answers

The questions we get about voice model rankings

It depends on which board you trust and what the agent does. As of September 25, 2026, five systems share the top of Artificial Analysis' preference arena within overlapping intervals; Qwen Audio 3.0 Realtime Plus has the best benchmark scores of any arena system and ranks last on preference; Gemini 3.8 Live scores well on both and has the lowest cost per hour in our benchmark table.
Digital Applied newsletter

Deep dives on AI, marketing and development.

Practical guides and fresh insights by email. No recycled takes.

Related dispatches

Continue reading

AI Development

What Paying for a Faster AI Model Actually Costs You

A census of 18 speed and priority tiers where one vendor sells the same model faster: the multiple, the vendor's speed claim, and what else changes.

September 25, 2026 · 7 minRead
AI Development

Best Text-to-Speech Models, September 2026: Ranked, Priced

92 text-to-speech models ranked by blind listener vote, with the price per million characters, open-weight licences and the job each leader fits.

September 24, 2026 · 9 minRead
AI Development

Gemini 3.8 Flash TTS: Voice Cloning and a Price That Doubles

Gemini 3.8 Flash TTS is generally available with voice cloning from a 30-second sample. The promotional price ends December 31 and doubles on January 1, 2027.

September 23, 2026 · 5 minRead
AI Development

How Often AI Coding Agents Cheat on Tests: Published Rates

Every published rate of coding agents gaming their tests, with the count or population it is over: METR, ImpossibleBench and a September 2026 paper. 19 rows.

September 17, 2026 · 8 minRead
AI Development

AI Agent Memory 2026: Vector, Graph, Episodic Update

AI agent memory architectures compared after Code with Claude London — Anthropic Dreaming, Memory Tool, Google Memory Bank, vector, graph, episodic patterns.

May 24, 2026 · 16 minRead
AI Development

AI Agent Governance: Policy and Compliance 2026 Guide

AI agent governance framework for enterprises — access control, audit trails, data residency, and compliance with EU AI Act and SOC 2 requirements.

May 23, 2026 · 20 minRead