DevelopmentFramework17 min readPublished August 5, 2026

Independent leaderboard · 2.2% vs 5.2% WER on the same datasets

Voice AI Benchmarks: How to Read Vendor Claims Right

A voice vendor can publish a 1.5–2.0× accuracy claim and an independent leaderboard can rank the same models in an order nobody would guess from the vendor page — without either party being wrong. The claims measure different things. This is how to tell which number answers your question.

DA
Digital Applied Team
Senior strategists · Published August 5, 2026
PublishedAugust 5, 2026
Read time17 min
Sources7 primary sources
Independent WER
2.2%
ElevenLabs Scribe v2, non-streaming leaderboard at the time of writing
vs 5.2% for Nova-3
Same datasets, error ratio
2.4×
Nova-3 word error rate over Scribe v2
Speech-to-Speech Index
82.9%
Grok Voice Think Fast 2.0, vendor-reproduced chart
+7.2 vs Think Fast 1.0
Vendor-run accuracy claim
1.5–2.0×
xAI's own harness; no absolute WER published

Voice AI benchmarks have a specific failure mode that text-model leaderboards mostly avoid: two credible parties can publish numbers about the same models, disagree completely, and both be reporting honestly. One measured a full-duplex conversation. The other measured batch transcription. Neither page says so in the headline, and the gap between those two things is where most procurement mistakes get made.

The live example is xAI's Grok Voice Think Fast 2.0 announcement of July 29, 2026, which carries two very different kinds of claim in the same post: a reproduced third-party chart with attributed scores, and a separate accuracy multiplier from xAI's own evaluation harness. Read quickly, they blur into one impression. Read carefully, only one of them is something you can independently check.

This guide separates the two, sets them against Artificial Analysis's independent speech-to-text leaderboard, unpacks what word error rate actually counts, opens up the composite Speech-to-Speech Quality Index, and ends with a harness you can run on your own audio. Every figure below is attributed to the page it came from and labelled with how it was produced.

Key takeaways
  1. 01
    A vendor multiplier and a leaderboard score are not comparable.xAI's 1.5–2.0× transcription-accuracy claim was produced on its own unpublished harness across thousands of short phrases in 24 languages. No absolute word error rate was published for it, so it cannot be placed on any independent leaderboard.
  2. 02
    The same two models rank far apart on a shared dataset.On Artificial Analysis's non-streaming speech-to-text leaderboard at the time of writing, ElevenLabs Scribe v2 sits at 2.2% word error rate and Deepgram Nova-3 at 5.2% — 3.0 percentage points apart, meaning Nova-3's error rate is roughly 2.4× Scribe v2's on identical audio.
  3. 03
    Composite indexes hide which component moved.The Speech-to-Speech Quality Index combines speech reasoning, conversational dynamics and agentic performance. On the July 29 chart the top two models are separated by 1.2 points on reasoning and 0.6 points on dynamics, but 10.8 points on agentic tasks.
  4. 04
    Realistic audio costs real capability.τ-Voice runs 278 customer-service tasks through noise mixing, telephony compression and packet loss. Audio-native models retain roughly 79% of their text-benchmark task completion when moved to full-duplex voice — about a fifth of capability lost in the channel.
  5. 05
    Shared branding is not shared architecture.A separate xAI-branded entry, Grok Speech to Text, appears on that same non-streaming speech-to-text leaderboard at 4.0% word error rate. No primary source available at the time of writing confirms whether it is the same model as Grok Voice Think Fast 2.0's built-in transcription. Do not treat them as one product.

01The ClaimOne announcement, two different kinds of number.

xAI announced Grok Voice Think Fast 2.0 on July 29, 2026. The post contains a benchmark table with a claimed Artificial Analysis Speech-to-Speech Quality Index of 82.9% for the new model, up from 75.7% for Think Fast 1.0, alongside GPT-Realtime-2.1 (High) at 79.1% and Gemini 3.1 Flash (High) at 69.5%. Those competitor figures are sourced from and attributed to Artificial Analysis rather than generated by xAI — an important detail, because it means the numbers are checkable against a third party.

Elsewhere in the same post sits a claim of an entirely different character. xAI states that in its own evaluation across thousands of short phrases in 24 languages, Grok Voice Think Fast 2.0 demonstrated a 1.5–2.0× improvement in accuracy relative to Deepgram Nova 3 and ElevenLabs Scribe v2, and a 1.4× improvement relative to Think Fast 1.0 — widening to roughly 10× in noisy and telephony-compressed conditions. That evaluation is vendor-run on an unpublished harness. No absolute word error rate is published for any of the three models in that specific comparison, which is what makes it impossible to place on a shared scale.

Checkable
Reproduced third-party chart
AA Speech-to-Speech Quality Index

Scores for four models on a named, published index, with the competitor figures attributed to Artificial Analysis. You can open the index methodology, see what it is composed of, and check whether the reproduction matches the live page.

82.9 · 79.1 · 75.7 · 69.5
Not checkable
Vendor-run accuracy multiplier
xAI's own harness, unpublished

A relative improvement across thousands of short phrases in 24 languages, widening in noisy and telephony-compressed audio. The audio set, the scoring configuration and the absolute error rates behind the multiplier are not published.

1.5–2.0× · ~10× in noise

Both claims can be true at once, and the second one may well be accurate. The problem is that a multiplier without a denominator cannot be ranked. If Nova 3 scores 5.2% word error rate on some test and a competitor is 2× more accurate on that same test, the competitor lands near 2.6% — but nothing in the announcement tells you the harness resembles the leaderboard, so the arithmetic is hypothetical rather than reported. We walked through the launch itself in our earlier look at Grok Voice Think Fast 2.0; this piece is about how to weigh what that announcement said.

For completeness on the commercial side: xAI listed Grok Voice Think Fast 2.0 at $0.08 per minute of audio as an announced API list price on July 29, 2026, and set out a schedule to point the grok-voice-latest alias from Think Fast 1.0 to 2.0 on August 5, 2026. Both are announced positions from that post rather than observed behaviour, and API list prices are the one number in this article that a vendor can change without republishing a benchmark.

“Grok Voice Think Fast 2.0 outperforms even dedicated, state of the art transcription models when it comes to accuracy.”— xAI, Grok Voice Think Fast 2.0 announcement, July 29, 2026

02Independent MeasurementWhat the independent leaderboard shows instead.

Artificial Analysis runs a separate speech-to-text leaderboard that scores transcription models on identical datasets. It is the closest thing this category has to a shared yardstick, and it produces an ordering that a reader of vendor pages alone would not predict.

On the non-streaming index at the time of writing, ElevenLabs Scribe v2 records 2.2% word error rate and Deepgram Nova-3 records 5.2%. That is a gap of 3.0 percentage points, or put as a ratio, Nova-3 makes roughly 2.4 word errors for every one Scribe v2 makes on the same audio. xAI's claim names both products in a single breath, as though they occupy similar ground. On this leaderboard they do not.

Word error rate on identical datasets · lower is better

Source: Artificial Analysis non-streaming speech-to-text leaderboard, a live page checked at the time of writing. Bars scale to the highest error rate shown.
Fun-Realtime-ASR-previewTop of the leaderboard across all 55 evaluated models
1.7%
ElevenLabs Scribe v2Named as a comparison target in xAI's accuracy claim
2.2%
Grok Speech to Text (SpaceXAI)Separate xAI-branded STT entry · identity unresolved
4.0%
Deepgram Nova-3Also named in xAI's claim · highest error rate of the four
5.2%

Two further numbers on that page are worth carrying with you. A separate xAI-branded product, Grok Speech to Text, appears at 4.0% word error rate — between Nova-3 and Scribe v2, and 2.3 points behind the leaderboard leader. And the overall top result across all 55 evaluated models is Fun-Realtime-ASR-preview at 1.7%, half a point ahead of Scribe v2. Neither figure appears in any vendor announcement discussed here, which is the point: the independent page has a wider field than any single vendor's chosen comparison set.

Different test, different answer
Nothing in this section accuses anyone of misreporting. A vendor measuring short multilingual phrases on its own harness and an independent evaluator measuring long-form transcription on shared corpora are answering different questions. The error is in the reader, not the writer — treating a relative multiplier and an absolute leaderboard score as points on one scale. The same reflex causes trouble on text leaderboards, as we covered in the guide to benchmark methodology and contamination.

03Side by SideThe same models against four different yardsticks.

We could not find a page that puts a full-duplex conversational model's self-reported transcription claim next to an independent transcription leaderboard while keeping the two categories visibly apart. The table below does exactly that. The two groups are deliberately not averaged or reconciled — they measure different pipelines, on different methodologies, and combining them would manufacture a comparison neither source supports.

Voice models and transcription products grouped by measurement category: full-duplex speech-to-speech models scored on the Artificial Analysis Speech-to-Speech Quality Index as reproduced in xAI’s July 29, 2026 chart, and transcription-only products scored on the Artificial Analysis non-streaming word error rate leaderboard checked at the time of writing. Columns show the headline score, the computed gap to that category’s leader, time to first audio, the vendor’s own accuracy claim, and what the score does not tell you.
Model / productHeadline scoreGap vs category leaderTime to first audioVendor’s own accuracy claimWhat the score does not tell you
Full-duplex speech-to-speech · AA Speech-to-Speech Quality Index, as reproduced in xAI’s July 29, 2026 chart
Grok Voice Think Fast 2.082.9%Category leader0.70s1.5–2.0× more accurate than Nova 3 and Scribe v2 on xAI’s own harness; no absolute WER publishedWhich component drove the index — the sub-scores diverge sharply
GPT-Realtime-2.1 (High)79.1%−3.8 ptsNot shown — the cell is blank in xAI’s chartNone quoted in this chartIts latency, which this reproduction simply does not report
Grok Voice Think Fast 1.075.7%−7.2 pts1.25sBaseline for xAI’s claimed 1.4× transcription improvementWhether the generational gain is uniform or concentrated in one component
Gemini 3.1 Flash (High)69.5%−13.4 pts2.98sNone quoted in this chartHow much of the deficit is agentic task work rather than speech quality
Transcription only · AA non-streaming word error rate, live leaderboard checked at the time of writing
Fun-Realtime-ASR-preview1.7% WERCategory leader of 55 modelsNot measured on this indexNone quoted hereStreaming behaviour, cost, or suitability for live agents
ElevenLabs Scribe v22.2% WER+0.5 ptsNot measured on this indexNone on this page — named as a target in xAI’s claimWhich error types make up the 2.2%, and how costly they are
Grok Speech to Text (SpaceXAI)4.0% WER+2.3 ptsNot measured on this indexNone quoted for this specific entryWhether this is the same model as Think Fast 2.0’s built-in transcription — unresolved
Deepgram Nova-35.2% WER+3.5 ptsNot measured on this indexUp to 36% lower WER than Whisper on select datasets — Deepgram methodology page, November 3, 2025Which datasets were selected, and how the harness was configured

The gap column is computed, not quoted: each row is the difference between that row’s headline score and the best score inside its own group. Read across the top group and the spread is 13.4 points from leader to laggard. Read across the bottom group and the spread is 3.5 percentage points of word error rate — which sounds smaller until you convert it to a ratio and find the bottom row making roughly three times as many errors as the top row on identical audio. Percentage points and ratios tell different stories about the same two numbers, and vendors reliably quote whichever is more flattering.

04The MetricWhat word error rate counts — and what it flattens.

Word error rate is a single arithmetic expression: WER = (S + I + D) / N, where S is substitutions, I is insertions, D is deletions, and N is the number of words in the reference transcript. It is Levenshtein edit distance computed at the word level rather than the character or phoneme level. Deepgram’s own methodology page also defines the two metrics that travel with it: word accuracy rate, which is simply 1 − WER, and character error rate, which applies the same counting logic per character.

The formula treats every error identically, and that is the whole problem. Deepgram’s own worked example makes it concrete. Against the reference “can you transfer five thousand dollars to my savings account” — ten words — a system that outputs “can you transfer five hundred dollars my savings account” has committed one substitution (thousand becomes hundred) and one deletion (the word to). Two errors over ten reference words is a 20% word error rate. The arithmetic is correct and the transaction is wrong by an order of magnitude. A generic WER number does not flag that severity at all.

Worked example
WER on an order-of-magnitude error
20%

One substitution plus one deletion over a ten-word reference. The dollar figure changed from five thousand to five hundred, and the metric scored it identically to a dropped preposition.

(1 + 0 + 1) / 10 = 0.20
Error classes counted
Substitutions, insertions, deletions
3

All three are summed and divided by the reference word count. There is no weighting for entity type, no premium on numbers or names, and no notion of which words a downstream system actually needed.

WER = (S + I + D) / N
The standing caution
Lower WER is not better understanding
2003

A Microsoft Research study found that optimising directly for the understanding objective produced higher task accuracy than optimising for word error rate alone. Cited on the canonical WER reference as the caution against treating it as a complete proxy.

Wang, Acero & Chelba · IEEE ASRU

That 2003 finding is more relevant now than when it was published, because almost nothing in a 2026 voice stack consumes a transcript directly. The transcript feeds an agent that extracts an intent, a reference number, an amount, a date. A system with a slightly worse word error rate that never mangles a currency figure is worth more than a lower-WER system that occasionally does — and no leaderboard column captures that. The same distinction between headline accuracy and consequence shows up across model evaluation generally; we explored the reliability side of it in our study of hallucination-rate benchmarks.

Vendor honesty, quoted
Deepgram’s own developer page states the limitation plainly: “WER treats all errors equally, but not all errors carry equal weight.” It is worth noticing that the clearest warning about the metric’s ceiling comes from a vendor — and that the same page hedges its own Nova-3 claim with the phrase “on select datasets.” That hedge is the single most useful phrase to search for in any benchmark announcement.

05Composite IndexesInside the Speech-to-Speech Quality Index: three benchmarks in a trench coat.

The 82.9% headline is not a measurement. It is an aggregate of exactly three component benchmarks: Speech Reasoning, measured on Big Bench Audio; Conversational Dynamics, measured on a Full Duplex Bench subset; and Agentic Performance, measured on τ-Voice. Artificial Analysis is explicit that a model must have results on all three to receive an index score at all — models missing a component are excluded from the index entirely rather than scored as zero, which quietly shapes who appears on the chart in the first place.

Big Bench Audio, the reasoning component, is a 1,000-question dataset adapted from Big Bench Hard, spanning four categories at 250 questions each — Formal Fallacies, Navigate, Object Counting and Web of Lies. The model receives audio input, must generate audio output, and each answer is graded correct or incorrect. That is a narrower thing than “speech reasoning” sounds like.

Now look at where the models actually separate. On the July 29 chart, Grok Voice Think Fast 2.0 leads GPT-Realtime-2.1 on Speech Reasoning by 1.2 points (97.2% against 96.0%) and trails it on Conversational Dynamics by 0.6 points (95.1% against 95.7%). Both models are above 95% on both components. The entire competitive story lives in the third component, where the spread runs from 56.5% down to 37.7%.

Agentic performance · where the index actually separates

Source: agentic performance on τ-Voice, as reproduced in xAI's July 29, 2026 announcement chart
Grok Voice Think Fast 2.0Highest agentic score on the chart
56.5%
Grok Voice Think Fast 1.0Prior generation · 4.4 points behind 2.0
52.1%
GPT-Realtime-2.1 (High)10.8 points behind the leader on this component
45.7%
Gemini 3.1 Flash (High)18.8 points behind the leader — the widest gap on the chart
37.7%

Here is the piece of arithmetic that makes the structure legible. For the two models where the July 29 chart lists all three component scores, an equal-weight average reproduces the published index exactly. Think Fast 2.0: 97.2 plus 95.1 plus 56.5 is 248.8, divided by three is 82.93, which rounds to the published 82.9. GPT-Realtime-2.1: 96.0 plus 95.7 plus 45.7 is 237.4, divided by three is 79.13, which rounds to the published 79.1. Two independent checks landing on the decimal is strong evidence that the components carry equal weight in this aggregate, though we would treat that as consistent-with rather than confirmed until the weighting is stated on the methodology page.

If that reading holds, it has a real consequence. A component where every serious model already scores above 95 contributes exactly as much to the headline as the component where models range across nearly 19 points. Two thirds of the index is effectively saturated; one third does all the discriminating. A buyer who cares about conversational feel and a buyer who cares about completing transactions are reading the same 82.9 and should be reading different columns.

Two methodology asterisks worth knowing
Artificial Analysis states in its own methodology notes: “Models must have results for all three datasets to receive an index score.” It also flags a trial-count inconsistency inside its τ-Voice reporting: most models are scored on an average of three trials, but some are scored on fewer — one model on a single trial, another on two. Even an independent leaderboard carries per-model asterisks. A third scope limit: the cost metric is normalised to cost per hour of input audio, computed on a fixed 40-question Big Bench Audio subset, and explicitly excludes cached-token discounts and tool-call costs.

That last point deserves an example, because it is where index reading turns into budgeting error. xAI’s announced API list price of $0.08 per minute of audio multiplies out to $4.80 per sixty minutes. It would be natural to set that against a leaderboard’s cost-per-hour-of-input-audio column and treat the two as the same quantity. They are not: the leaderboard figure is derived from a fixed 40-question subset with cached-token discounts and tool calls removed, while a production voice agent bills for output audio and tool use as well. The arithmetic is right and the comparison is still wrong — the same species of mistake we catalogued for coding leaderboards in the SWE-bench Live leaderboard analysis.

06Grounded Evaluationτ-Voice: what happens when the audio gets realistic.

τ-Voice is the component doing the discriminating, so it is worth understanding on its own terms. It is a benchmark from Sierra Research, published as an arXiv preprint in March 2026 and announced in plain language on Sierra’s blog on May 1, 2026. It runs 278 customer-service tasks — retail, airline and telecom domains — inherited directly from τ-bench’s text evaluation. That inheritance is the design decision that matters: because the tasks are the same, you can measure what a model loses purely by moving from text to full-duplex voice.

The audio is not clean studio speech. τ-Voice stacks environmental noise mixing (street noise, car horns, background speech not directed at the agent), speaker personas generated across a diverse range, telephony compression using the G.711 μ-law codec at 8 kHz, and packet and frame drops modelled with a Gilbert-Elliott model — on top of natural turn-taking that includes interruptions and backchannels. That is a deliberate reconstruction of a real phone line, not a simulation of a podcast.

Two headline findings come out of it. Voice-agent task completion rose from roughly 30% to roughly 67% between August 2025 and April 2026 — eight months, about 37 percentage points, a little over a doubling. And audio-native models now retain roughly 79% of their text-benchmark task completion when moved to full-duplex voice. Both numbers are more useful than any index score, because the second one quantifies the channel penalty that vendor marketing almost never puts a figure on: about a fifth of a model’s demonstrated text capability does not survive the trip through a phone line.

Failure mode
Background noise
Street noise, horns, undirected speech

Mixed into the input rather than filtered out. Contributes measurably to degraded task completion, and is the condition where xAI claims its own harness shows the widest advantage — a claim that remains vendor-run and unpublished.

Stacked with codec loss, not tested alone
Failure mode
Accent diversity
Varied speaker personas

Synthesised across a range of speakers rather than a narrow demographic. This is the axis most likely to differ between a benchmark corpus and your actual customer base, which is why an internal replication matters more here than anywhere else.

Corpus choice drives the result
Failure mode
Turn-taking dynamics
Interruptions and backchannels

The full-duplex problem proper: knowing when the caller has finished, when an interruption is a correction, and when a short acknowledgement is not a turn at all. Invisible to any batch transcription metric.

Not measurable by WER

Those are the three failure modes named in the material available at the time of writing. We are deliberately not putting a total count on the list — the summary sourcing we could verify names three, and a benchmark paper’s taxonomy is exactly the kind of detail that gets miscounted in secondary coverage. If you are building on τ-Voice, read the preprint rather than any summary of it, including this one.

The broader lesson generalises past voice. Benchmarks that inherit their task set from an established text evaluation are far more informative than benchmarks invented alongside the model they first scored, because inheritance gives you a controlled comparison rather than an absolute number floating free. The same argument applies to agent evaluation on desktops and terminals, which we unpacked in our explainer on OSWorld and Terminal-Bench.

07Product IdentityWhen two products share a brand but maybe not a model.

There is a genuinely unresolved question sitting in the middle of this comparison, and most coverage papers over it. Artificial Analysis’s speech-to-text leaderboard lists an entry called Grok Speech to Text, attributed to SpaceXAI, at 4.0% word error rate. xAI’s announcement describes Grok Voice Think Fast 2.0, a full-duplex conversational model with built-in transcription, and makes a relative accuracy claim about it.

Whether those are the same weights behind two product names, or two separate product lines, is not established by any primary source we could verify at the time of writing. That matters more than it sounds. If they are the same model, then xAI’s transcription claim has an independent check sitting right there at 4.0% — behind Scribe v2’s 2.2%, which is one of the two products the vendor claim says it beats. If they are different models, the 4.0% figure says nothing whatsoever about Think Fast 2.0 and the vendor claim remains unchecked. The two readings point in opposite directions, so assuming either one is worse than admitting the ambiguity.

Flagged, not resolved
We are stating this as an open question, not a finding. Grok Speech to Text (SpaceXAI) at 4.0% word error rate and Grok Voice Think Fast 2.0 are treated here as separate entries because nothing available at the time of writing confirms they are the same model. If you are evaluating either for production, the correct next step is to ask xAI which endpoint the leaderboard entry corresponds to — a one-line question that resolves the whole thing.
Safe
Compare within one leaderboard

Two rows on the same page, evaluated on the same datasets by the same harness, are directly comparable. Scribe v2 at 2.2% against Nova-3 at 5.2% is a real, like-for-like result you can act on.

Compare rows, not pages
Unsafe
Assume shared branding means shared weights

A vendor can ship a dedicated transcription endpoint and a conversational model under one brand with different architectures, different training data and different tuning. Ask which artefact was measured before you carry a number across.

Ask the vendor
Unsafe
Fill a blank cell from memory

GPT-Realtime-2.1's time-to-first-audio was left blank in xAI's reproduction of the chart. The honest move is to report the omission. Substituting a number from another source, another version, or recall creates a comparison that no source supports.

Report the gap
Unsafe
Reconcile an index with a WER leaderboard

The Speech-to-Speech Quality Index composites three conversational benchmarks. The word error rate leaderboard scores batch transcription. Averaging, ranking or blending across the two produces a number that describes nothing that exists.

Keep them separate

08Read It YourselfThe harness that settles it: your audio, one configuration.

The most useful methodology guidance in this whole space is published by a vendor with a stake in the outcome, which is either ironic or reassuring depending on your temperament. Deepgram’s speech-to-text benchmarks page sets out a fair-comparison procedure that any team can run, and it is the right skeleton for an internal evaluation.

Four controls do most of the work. Assemble production-realistic audio from public corpora — LibriSpeech for clean speech, Common Voice for accent variety, TED-LIUM for single-speaker long form — supplemented with your own recordings. Build one identical containerised harness per vendor, so the only variable is the model. Log time-to-first-byte and final latency identically across vendors, because latency comparisons collapse the moment two stopwatches start at different points. Standardise scoring with jiwer, the open-source Python library, to eliminate the formatting and punctuation differences that otherwise move word error rate by more than the model choice does.

To that we would add three checks the vendor page does not cover. Weight your errors by business consequence rather than counting them flat — build a small set of held-out utterances containing the entities that actually cost money in your domain, and score those separately. Test at the noise level your callers actually generate, not at studio quality, because that is where the τ-Voice numbers say capability disappears. And record a retrieval date every time you cite a live leaderboard, because those pages update continuously and a figure without a date is not a citation.

Corpus assembly
Public corpora, one purpose each
3

LibriSpeech covers clean read speech, Common Voice covers accent variety, TED-LIUM covers single-speaker long form. Each isolates a different failure surface; together they stop one flattering dataset from carrying the whole result.

Plus your own recordings
Scoring discipline
One scorer across every vendor
1

The open-source jiwer library applied identically to every transcript removes the formatting and punctuation skew that silently rewards whichever vendor happens to match your reference style. Without it, you are partly benchmarking text normalisation.

jiwer, identical configuration
Channel penalty
Assume the phone line takes a cut
79%

Audio-native models retain roughly 79% of text-benchmark task completion in full-duplex voice, per τ-Voice. Budget for the gap between a model's demo and its behaviour on a compressed, noisy, interrupted call.

Test at real noise levels

None of this requires a research team. A week of engineering produces a harness that outranks every vendor chart for your specific decision, because it is scored on the audio you actually receive. If you want that built and operated alongside the rest of your measurement stack, our AI and digital transformation engagements start with exactly this kind of comparative evaluation, and our analytics practice handles the instrumentation and reporting layer around it. If your interest is dictation rather than conversational agents, the trade-offs run differently — we compared that field in the open-source voice dictation roundup.

Looking ahead, the pressure in this category runs toward more grounded evaluation, not less. τ-Voice moved the field from clean audio to modelled telephony in a single step, and the capability retention figure it produced is the kind of number that becomes a procurement question once buyers know it exists. Expect the next round of vendor announcements to start quoting channel-degraded results directly — and expect the harnesses behind those results to remain unpublished unless buyers make publication a condition of the conversation.

09ConclusionEvery benchmark number is an answer to a question you did not ask.

Reading voice benchmarks, August 2026

A score without its harness is a sentence fragment.

The July 29 announcement is a useful specimen precisely because nothing in it is dishonest. xAI attributed its competitor figures to a third party, described its own evaluation as its own, and quoted a relative improvement rather than inventing an absolute. Every problem in this article comes from reading those claims together as though they sat on one scale.

The discipline that fixes it is small. Ask what was measured — batch transcription or a live conversation. Ask who ran the harness and whether it is published. Ask what the composite is made of, and check whether the components that separate the field are the ones you care about. Ask whether two similarly-named products are actually the same model. Where a source leaves a cell blank, report the blank.

None of that requires distrusting vendors. It requires treating a benchmark score the way you would treat a quoted price without a currency: not wrong, just incomplete until you know what it was denominated in. The teams that build a small internal harness on their own audio stop needing to adjudicate between vendor pages at all — they have the only number that was ever going to decide it.

Benchmark voice AI on your own audio

The only voice benchmark that decides your build is the one run on your own audio.

Our team builds internal evaluation harnesses for voice and language models — production-realistic corpora, identical containerised runs, consequence-weighted scoring — so procurement decisions rest on your audio rather than a vendor chart.

Free consultationExpert guidanceTailored solutions
What we work on

Model evaluation engagements

  • Internal WER harnesses on your own call recordings
  • Consequence-weighted scoring for entities that cost money
  • Latency instrumentation across competing vendors
  • Voice-agent task evaluation beyond transcription accuracy
  • Procurement briefs that separate vendor claims from measurement
FAQ · Reading voice AI benchmarks

The questions worth asking before you believe a number.

Word error rate is the standard transcription accuracy metric, computed as WER = (S + I + D) / N — substitutions plus insertions plus deletions, divided by the number of words in the reference transcript. It is Levenshtein edit distance applied at the word level rather than the character or phoneme level. Two related metrics travel with it: word accuracy rate, which is simply 1 − WER, and character error rate, which applies the same counting logic per character. The formula's defining limitation is that it weights every error identically. Deepgram's own methodology page puts it directly: WER treats all errors equally, but not all errors carry equal weight. A dropped preposition and a mangled currency figure score the same, which is why a raw WER number is a starting point for evaluation rather than a conclusion.
Related dispatches

Continue exploring benchmark literacy.