Artificial Analysis publishes two ways of ranking realtime voice systems, the models that listen and talk back without a separate transcription step. One is an arena: listeners hear two systems answer the same prompt and vote, producing an Elo score. The other is a benchmark set: a speech-reasoning test, a turn-taking test, a time-to-first-audio measurement and a cost per hour. Read on September 25, 2026, the two boards disagree about which systems are good, and the disagreement is not noise.
Of the systems on both boards, the one with the best benchmark scores, Qwen Audio 3.0 Realtime Plus, ranks last on preference, 26th of 26. OpenAI's GPT-Realtime-2.1 outscores the older GPT-Realtime-1.5 on Big Bench Audio (95.8% against 81.4%), ties it on Full Duplex Bench and loses to it on preference. This post joins the two tables, says what each one measures, marks where the confidence intervals make a ranking meaningless, and ends with how to choose a model for a voice agent when the boards point in different directions.
- 01Qwen Audio 3.0 Realtime Plus tops the benchmarks among arena systems and sits last on preference.Big Bench Audio 99.2%, Full Duplex Bench 98.4%, the highest of any arena system, behind only StepAudio 3 Realtime (99.7% and 98.9%), which has no arena score. Arena Elo 775 with an interval of 731 to 819, below every other listed system, on 171 votes.
- 02GPT-Realtime-2.1 High scores below the 1.5 anchor on preference and far above it on Big Bench Audio.Elo 928 (902 to 954) against the anchor's 1000. It also costs $10.75 per hour of input audio on Artificial Analysis' measure, the highest in our table, though the 1.5 anchor itself is listed at $11.44.
- 03The top of the preference board is a statistical tie.Gemini 3.1 Flash Live (Minimal) at 1096 and Gemini 3.8 Live at 1083 have overlapping intervals, as do the next three. Elo differences inside overlapping intervals are not rankings, and the table marks the intervals so you can see which ones are.
- 04Task completion is a third measure, and it disagrees with both.Across the fourteen listed rows, task completion ranges from 57.1% to 94.6%, and the preference leader completes 74.6% of tasks against 93.2% for the system one place below it.
01 — The findingThe split, in one pair of arena systems
Qwen Audio 3.0 Realtime Plus
Gemini 3.1 Flash Live (Minimal)
Both columns are from Artificial Analysis' speech-to-speech page, read on September 25, 2026. The preference leader scores 71.3% on Big Bench Audio and 72.3% on Full Duplex Bench, and the arena system with the best benchmark scores has the lowest preference score. That is the whole problem: a buyer reading either board alone would choose a system the other board says is the wrong one.
02 — The dataThe preference board: Elo, interval, votes, task completion
The arena anchors GPT-Realtime-1.5 at 1000 and scores every other system relative to it. The interval column is the 95% confidence interval Artificial Analysis publishes; where two intervals overlap, the two systems are not distinguishable on preference. Systems labelled cascaded run a separate transcription, text model and speech step rather than one native model, which Artificial Analysis marks on the board and we carry over. Rank numbers are the board's own and skip where our read of the board did not capture a row.
| System | Elo | 95% interval | Task completion |
|---|---|---|---|
| 1. Gemini 3.1 Flash Live (Minimal) | 1096 | 1068–1123 | 74.6% |
| 2. Gemini 3.8 Live | 1083 | 1053–1113 | 93.2% |
| 3. Gemini 3.1 Flash Live (High) | 1063 | 1036–1090 | 71.8% |
| 4. GPT-Live-1 (Sol, low) | 1053 | 1026–1081 | 90.9% |
| 5. GPT-Live-1 (Astra, medium) | 1048 | 1020–1076 | 87.4% |
| 6. Cartesia Line (cascaded) | 1027 | 990–1064 | 77.5% |
| 7. Grok Voice Think Fast 2.0 High | 1011 | 983–1038 | 94.6% |
| 8. GPT-Realtime-1.5 (anchor) | 1000 | anchor | 85.1% |
| 9. ElevenLabs Agents (cascaded) | 993 | 967–1020 | 90.5% |
| 10. Gemini 3.8 Live Extended Thinking (High) | 990 | 963–1017 | 89.1% |
| 12. Amazon Nova 2.0 Sonic (Mar 2026) | 974 | 948–1000 | 57.1% |
| 15. GPT-Realtime-2.1 High | 928 | 902–954 | 91.5% |
| 18. Deepgram Voice Agent (cascaded) | 914 | 888–940 | 73.7% |
| 26. Qwen Audio 3.0 Realtime Plus | 775 | 731–819 | 77.8% |
Read the intervals before the ranks. Positions one to five span 1048 to 1096 and every adjacent pair overlaps, so the honest statement is that five systems share the top of the board, and the sixth and seventh overlap the fifth. The one gap that is unmistakable is Qwen's: its upper bound of 819 sits below the lower bound of every other row in this table, though three uncaptured systems at ranks 23 to 25 overlap it. Task completion tells a different story again. The system listeners prefer most completes fewer tasks than ten of the thirteen systems below it, and Grok Voice Think Fast 2.0 High, seventh on preference, has the best completion rate in the table. Our post on Grok Voice Think Fast 2 covers what xAI built it for.
03 — The dataThe benchmark board: reasoning, turn-taking, latency, cost
Big Bench Audio tests whether the system reasons correctly about what it heard. Full Duplex Bench tests turn-taking: whether the system yields, interrupts and resumes the way a conversation needs. Time to first audio is the seconds until the system starts speaking. The cost column is Artificial Analysis' figure for an hour of input audio on the vendor's list price. Scores are as published; a dash means the page lists no value.
| System | BBA / FDB | TTFA (s) | $ per hour |
|---|---|---|---|
| StepAudio 3 Realtime | 99.7% / 98.9% | 8.83 | – |
| Qwen Audio 3.0 Realtime Plus | 99.2% / 98.4% | 1.54 | $4.42 |
| Gemini 3.8 Live Extended Thinking (High) | 97.7% / 91.9% | 1.35 | $3.50 |
| Grok Voice Think Fast 2.0 High | 97.2% / 95.1% | 0.70 | $4.80 |
| GPT-Realtime-2 (High) | 96.6% / 95.3% | 1.14 | $4.14 |
| GPT-Realtime-2.1 High | 95.8% / 95.7% | 1.21 | $10.75 |
| Gemini 3.8 Live | 91.7% / 96.1% | 1.18 | $0.84 |
| GPT-Live-1 (Astra, medium) | 90.1% / 94.9% | 1.34 | $5.83 |
| GPT-Live-1 (Sol, low) | 89.0% / 97.3% | 1.24 | $4.47 |
| Amazon Nova 2.0 Sonic (Mar 2026) | 88.1% / – | 1.14 | $4.90 |
| Deepslate Opal | 85.0% / 85.7% | 0.44 | $6.48 |
| Raon SpeechChat (Krafton) | 57.5% / 86.2% | 0.04 | – |
| Moshi (open) | 4.4% / 61.0% | – | – |
Cost per hour of input audio, systems in our benchmark table
Artificial Analysis, read September 25, 2026. Bars scaled to GPT-Realtime-2.1 High at $10.75.Two rows are worth a second look. Gemini 3.8 Live scores 91.7% on reasoning, 96.1% on turn-taking, starts speaking in 1.18 seconds, ranks second on preference with the second-best task completion, and costs $0.84 an hour, less than a quarter of the next cheapest listed system. And StepAudio 3 Realtime, top on both benchmarks at 99.7% and 98.9%, takes 8.83 seconds to begin speaking, which disqualifies it from a live conversation whatever it scores. Our post on Gemini 3.8 Live and its extended-thinking variant explains the split between the two Google rows.
04 — The reasonWhy a benchmark and a listener disagree
The two boards measure different things and neither is wrong. A benchmark scores an answer against a key: did the system reason correctly about the audio, did it yield the turn when it should. A listener votes on an experience: did it sound like a person, did the voice fit, did it feel quick, did it say too much. A system can be correct and unpleasant, or pleasant and wrong, and the boards show both cases. The turn-taking benchmark rewards a system that interrupts at exactly the right moment; a listener may prefer one that waits a beat too long and sounds polite.
The arena also has a sample problem the intervals make visible. Qwen's 171 votes produce an interval 88 points wide; the top-ranked system's 841 votes produce one 55 points wide. A new system enters with a wide interval and a small vote count, and its position moves as votes arrive. That is why we quote the interval with every Elo, why a difference inside overlapping intervals is not a ranking, and why the board should be re-read rather than remembered.
Cascaded systems sit on both boards with their own trade-off. A pipeline of transcription, text model and speech synthesis can reason with a stronger text model than any native voice model carries, and it pays for that in latency and in turn-taking, because the pieces do not share a clock. Our reference on voice-agent latency measures defines time to first audio and the other timings the benchmark board uses.
Your callers, your accents, your background noise, your task. Both boards use their own prompts and their own audio. A voice agent for a contact centre in one region, with one product vocabulary, can rank the same six systems in a different order from either board, and usually does.
05 — The decisionHow to choose when the boards disagree
Prices for the vendor-side comparison are in our text-to-speech ranking for the synthesis half of a cascaded stack and in our post on Gemini 3.8 Flash TTS for the January price change. Our AI transformation practice runs the 50-call test as the first week of any voice-agent build.
A snapshot of one independent evaluator's two boards, joined by system name, with the evaluator's own intervals and units carried through.
- What was collected
- The Artificial Analysis speech-to-speech arena (Elo, 95% confidence interval, vote count, task completion) and its benchmark table (Big Bench Audio, Full Duplex Bench, time to first audio, cost per hour of input audio), parsed from the page's raw data. Fourteen arena rows and thirteen benchmark rows are shown; eight systems appear on both.
- Sources
- Artificial Analysis only, for every score. Vendor pricing pages were read for the cost cross-checks in section 05 and are cited in the linked posts. No score in this post was measured by us.
- As-of date
- Both boards read on September 25, 2026.
- Exclusions
- Arena rows at ranks 11, 13, 14, 16, 17 and 19 to 25 were not captured in our read of the board and are omitted; their ranks are skipped in the table so the board's numbering is preserved. Text-to-speech models are covered in a separate post. Hugging Face and OpenRouter voice boards were not used.
- Known limitations
- Elo positions with overlapping intervals are not rankings. Cascaded and native systems are scored on the same boards despite different architectures. The preference leader's benchmark scores (71.3% / 72.3%) are quoted in section 01 but not carried in the benchmark table. Cost per hour is the evaluator's normalisation, not a vendor quote.
- Refresh
- Refreshed in place when the top five on either board changes or a new native system enters both boards.
06 — ConclusionCorrect and liked are different scores, and the boards prove it
Pick the board that matches the job, read the intervals before the ranks, and test the top group on your own recorded calls before you sign
The best benchmark scorer on the arena is the least preferred system on it, and the preferred leader scores 71.3% and 72.3% on the two benchmarks. Neither board can choose for you. Fifty of your own calls can.