Speech-to-speech agents crossed a threshold this week. xAI — which now brands itself SpaceXAI — released Grok Voice Think Fast 2.0 on July 29, 2026: a model that listens, reasons and speaks inside a single network rather than chaining speech-to-text into a language model into text-to-speech. It scores 82.9 on the Artificial Analysis Speech-to-Speech Quality Index, up from 75.7, and it starts talking back in 0.70 seconds.
The number that will get quoted, though, is a different one. xAI states that Think Fast 2.0 transcribes 1.5 to 2.0 times more accurately than Deepgram Nova 3 and ElevenLabs Scribe v2, with the gap widening to roughly ten times in noisy conditions. If that holds, a single voice model just absorbed a job that has belonged to dedicated transcription engines for a decade — and every CRM that logs call transcripts has a vendor decision to reconsider.
This guide takes that claim at face value and then shows you exactly where the evidence stops. We decompose the quality index into its three sub-benchmarks, put Grok next to the model that actually outscores it, separate four pricing surfaces that vendors and coverage routinely conflate, and end with the architectural question that matters more than any leaderboard: whether your CRM needs a great conversation or a defensible transcript.
- 01Think Fast 2.0 is fast first, and top-tier second.82.9 on the Artificial Analysis Speech-to-Speech Quality Index at 0.70s time-to-first-audio. Qwen Audio 3.0 Realtime Plus scores higher at 84.1 — but takes about 4 seconds to start speaking, roughly 5.7 times longer.
- 02Almost all of the gain came from one sub-benchmark.The index is an equal-weighted mean of Speech Reasoning, Conversational Dynamics and Agentic Performance. Conversational Dynamics jumped 77.8 to 95.1 and accounts for about 79% of the raw sub-score improvement. Speech reasoning barely moved.
- 03The transcription claim is credible but unaudited.xAI reports a 1.5 to 2.0 times word-error-rate improvement over Deepgram Nova 3 and ElevenLabs Scribe v2 on its own eval of thousands of short phrases in 24 languages. No absolute word error rate is published, and no independent party has replicated it.
- 04Word error rate moves by roughly 4.8 times on configuration alone.The same Deepgram Nova-3 model has been measured at 5.26% on Deepgram's own batch audio and 25.3% by Coval in a latency-optimised voice-agent configuration — a 4.8-fold spread. Any single WER figure is a statement about a test harness, not a model.
- 05Four pricing surfaces, one product family.The raw model API is $0.08/min. Think Fast 1.0 was $0.05/min. Agent Builder advertised $0.05/min at its July 1 launch. xAI also sells dedicated speech-to-text at $0.10/hr REST. Whether the platform price moves when the default model flips on August 5 is unconfirmed.
01 — What ShippedA model that reasons while it speaks.
Grok Voice Think Fast 2.0 landed on July 29, 2026 as xAI’s next-generation voice model, positioned on three axes: intelligence, transcription accuracy, and conversational capability. The architectural bet is that a single speech-to-speech network beats a cascade of specialists on the things a live conversation actually rewards — turn-taking, interruption handling, and the gap between the end of your sentence and the start of the reply.
The headline latency figure moved from 1.25 seconds to 0.70 seconds to first audio. That is the number a caller feels. Everything else on the spec sheet is downstream of it: xAI also reports that median reasoning tokens per response fell to 0.4 times the Think Fast 1.0 baseline, so the model reaches an answer with roughly 60% less internal deliberation than its predecessor while scoring higher.
xAI describes tuning the model toward the patterns of real human conversation — shorter sentences, one question at a time, less filler. That is a design goal rather than a measured claim, and no benchmark number is attached to it. It does, however, explain where the benchmark movement came from, which we get to in the next section.
The model also has a no-code sibling. About four weeks earlier, xAI launched Grok Voice Agent Builder, the knowledge-base-and-connectors layer that sits on top of this same model family; we covered xAI’s no-code Voice Agent Builder at launch. That distinction matters later, because the two products are priced on different surfaces.
Grok Voice Think Fast 2.0
Single-model speech in, speech out. 0.70s to first audio, 82.9 AA quality index. Supports function calling, web and X search, Collections search, and remote MCP tool calls inside the voice session. Rate limits: 10 concurrent sessions per team, 120-minute maximum session length.
Grok Voice Think Fast 1.0
1.25s to first audio, 75.7 AA quality index. The August 5 flip moves the grok-voice-latest alias to 2.0; xAI's stated instruction for staying on this generation is to pin grok-voice-think-fast-1.0 before then. Cheaper per minute, and meaningfully weaker on conversational dynamics and agentic tasks.
02 — Benchmark AnatomyWhere the 7.2 points actually came from.
The Artificial Analysis Speech-to-Speech Quality Index is not a single test. It is an equal-weighted mean — one third each — of three sub-benchmarks: Speech Reasoning (Big Bench Audio), Conversational Dynamics (Full Duplex Bench), and Agentic Performance (τ-Voice). A model has to post a valid score on all three to appear in the index at all.
That structure makes the headline decomposable, and decomposing it is more useful than the headline. Below are the sub-scores xAI published for both generations, with the change and each sub-benchmark’s share of the total raw improvement recomputed from those figures.
| Sub-benchmark | TF 2.0 | TF 1.0 | Change | Share of gain |
|---|---|---|---|---|
| Speech ReasoningBig Bench Audio | 97.2 | 97.1 | +0.1 | 0.5% |
| Conversational DynamicsFull Duplex Bench | 95.1 | 77.8 | +17.3 | 79.4% |
| Agentic Performanceτ-Voice Bench | 56.5 | 52.1 | +4.4 | 20.2% |
| Quality Indexequal-weighted mean of the three | 82.9 | 75.7 | +7.2 | — |
Sub-scores as published by xAI, citing Artificial Analysis. Share of gain = each sub-benchmark’s point change divided by the sum of the three changes (0.1 + 17.3 + 4.4 = 21.8). Shares are rounded to one decimal and therefore sum to 100.1%. Index rows are the published values; the equal-weighted mean of the 2.0 sub-scores is 82.93 and of the 1.0 sub-scores is 75.67, both consistent with the published figures.
Read that table twice, because it reframes the release. Speech reasoning was already effectively saturated — 97.1 to 97.2 is noise. Agentic performance improved meaningfully in relative terms but is still the weakest leg by a wide margin at 56.5. Nearly four fifths of the visible improvement is Conversational Dynamics: turn-taking, barge-in, knowing when to stop talking. This is not a model that got smarter. It is a model that got better at holding a conversation, which for a phone line is arguably the more valuable upgrade.
It also sets a ceiling worth noticing. Full Duplex Bench at 95.1 leaves very little headroom on that leg. The next generation’s index movement will have to come from agentic performance — the 56.5 — which is the sub-benchmark that measures whether a voice agent can actually complete a task rather than merely sound good doing it. For a CRM front-end, that is the number to track across the next two releases.
03 — The FieldFastest among the leaders, not the highest scorer.
xAI’s own comparison chart puts Think Fast 2.0 against GPT-Realtime-2.1 and Gemini 3.1 Flash. It does not include Qwen Audio 3.0 Realtime Plus, which on the same Artificial Analysis index scores 84.1 — higher than Grok’s 82.9. That omission is worth naming, and it is also worth not overreading, because the Qwen model takes about 4 seconds to produce its first audio against Grok’s 0.70. That is a real engineering trade, not a gotcha.
Speech-to-Speech Quality Index · leaderboard snapshot
Source: Artificial Analysis Speech-to-Speech Index, retrieved August 2026Quality alone is the wrong lens for a phone line, so the table below adds latency and Artificial Analysis’s own hourly price column, plus a latency multiple computed against Grok Voice Think Fast 2.0’s 0.70 seconds. No model wins all three columns.
| Model | Quality index | Time to first audio | Latency vs Grok 2.0 | AA price / hr |
|---|---|---|---|---|
| Qwen Audio 3.0 Realtime Plus | 84.1 | 4.02s | 5.7× | $4.42 |
| Grok Voice Think Fast 2.0 | 82.9 | 0.70s | 1.0× | $4.80 |
| GPT-Realtime-2.1 (High) | 79.1 | 1.21s | 1.7× | $10.75 |
| GPT-Realtime-2 (High) | 77.2 | 1.14s | 1.6× | $4.14 |
| Qwen Audio 3.0 Realtime Flash | 76.3 | 4.16s | 5.9× | $4.77 |
| Grok Voice Think Fast 1.0 | 75.7 | 1.25s | 1.8× | $3.00 |
| Gemini 3.1 Flash (High) | 69.5 | 2.99s | 4.3× | $1.75 |
| Gemini 3.1 Flash (Minimal) | 56.6 | 0.96s | 1.4× | $1.50 |
Index, latency and hourly price as listed on the Artificial Analysis Speech-to-Speech leaderboard, retrieved August 2026. The hourly column is Artificial Analysis’s own blended figure, not a vendor-published hourly rate: OpenAI and Google meter realtime audio by token, so their rows are derived rather than quoted. xAI’s two rows do match its published per-minute pricing. Latency vs Grok 2.0 = each model’s time to first audio divided by 0.70s, rounded to one decimal. GPT-Realtime-2.1’s latency is shown as a dash in xAI’s own chart; the 1.21s figure here is Artificial Analysis’s, not xAI’s.
Three readings fall out of it. First, the top of the board is tight: 1.2 index points separate Qwen Plus from Grok 2.0, which is well inside the range where prompt design and voice selection will matter more than model choice. Second, latency is where the field genuinely splits — Grok 2.0 is roughly 5.7 times faster to first audio than the model that outscores it, and that difference is audible on a phone call in a way that 1.2 index points is not. Third, price does not track quality: GPT-Realtime-2.1 carries the highest hourly figure in the set at $10.75 — an Artificial Analysis blended number, not an OpenAI list rate — and sits third on quality.
Our reading is that speech-to-speech has stopped being a capability race and become a positioning one. Gemini 3.1 Flash Minimal is on the board at an AA-blended $1.50 an hour and 0.96 seconds — genuinely fast and genuinely cheap, at 56.6 quality. Qwen Plus takes the opposite corner. Grok 2.0 is the only entry currently sitting near the top of both the quality and the latency columns at once, and that combination, not the 82.9, is what it is actually selling. For the fuller picture of how these pieces assemble into a production system, our voice agent infrastructure stack reference covers why latency is the KPI that ends up governing everything else.
04 — The Transcription ClaimA bold claim with no absolute number attached.
This is the part of the release that deserves the most scrutiny, and it is worth being precise about what xAI actually said, because it is the opposite of what a voice model’s vendor usually concedes.
"Grok Voice Think Fast 2.0 outperforms even dedicated, state of the art transcription models when it comes to accuracy."— xAI, Grok Voice Think Fast 2.0 announcement, July 29, 2026
The supporting detail: an evaluation across thousands of short phrases in 24 different languages showing a 1.5 to 2.0 times improvement in word error rate relative to Deepgram Nova 3 and ElevenLabs Scribe v2, and a 1.4 times improvement over Think Fast 1.0. xAI further states that the gap widens to roughly ten times in noisy settings, and frames that as deliberate: the model was tuned for real-world audio with substantial background noise and telephony compression.
Take it seriously. A speech-to-speech model that ingests raw audio never passes through a lossy text bottleneck, so there is a plausible mechanism for it to hear things a transcription-first pipeline drops. The noisy-audio result is exactly where you would expect that advantage to show up.
Now the parts that are missing, all of which are checkable rather than speculative. The comparison models were chosen by xAI. The language set was chosen by xAI. The phrase set was chosen by xAI, and short phrases are a specific and relatively forgiving transcription regime. The results are published as relative multipliers on a chart with no visible axis values — no absolute word error rate appears anywhere. And as of publication, no independent benchmark, academic group or competing vendor has replicated the comparison.
| Who measured | Model measured | Test conditions | Absolute WER |
|---|---|---|---|
| xAI (vendor) | Grok Voice Think Fast 2.0 | Thousands of short phrases, 24 languages. Vendor-selected comparison set. Relative multipliers only. | Not published |
| Artificial Analysis | Deepgram Nova-3 | AA-WER v2. Duration-weighted average across three datasets, about 8 hours of audio. | 5.2% |
| Artificial Analysis | ElevenLabs Scribe v2 | Same AA-WER v2 harness and datasets. | 2.2% |
| Coval (third party) | Deepgram Nova-3 | Latency-optimised voice-agent configuration rather than batch transcription. | 25.3% |
| Deepgram (vendor) | Deepgram Nova-3 | Vendor-claimed median on batch audio. | 5.26% |
Sources: xAI announcement (July 29, 2026); Artificial Analysis speech-to-text index, retrieved August 2026; Coval and Deepgram figures via secondary reporting. Spread on the same model = 25.3 ÷ 5.26 = 4.81, i.e. roughly a 4.8× difference between the highest and lowest published Nova-3 figures.
Look at the three Nova-3 rows together. The identical model has been published at 5.2%, 5.26% and 25.3% word error rate — roughly a 4.8-fold spread — depending entirely on whether it was configured for batch accuracy or for live-agent latency. Nobody in that list is lying. They are measuring different things and calling both of them WER.
That is the honest frame for xAI’s claim. The problem is not that a 1.5 to 2.0 times improvement is implausible; it is that a relative multiplier against a vendor-chosen baseline, on a vendor-chosen corpus, with no absolute figure to anchor it, cannot be compared to anything you already measure. It is unfalsifiable in its published form — not wrong, just not yet checkable. And the same logic applies to a self-hosted Whisper transcription layer: the number that matters is the one your own audio produces.
05 — VerificationHow to test it yourself in an afternoon.
The good news about an unaudited claim is that this one is cheap to audit. Word error rate is a mechanical measure — substitutions plus deletions plus insertions, divided by reference words — and the audio you care about is already sitting in your call recordings. The whole exercise costs a few dollars of API time.
The design principle is that you are not trying to reproduce xAI’s benchmark. You are trying to answer a narrower and far more useful question: on our calls, with our accents, product names and line quality, which engine produces the transcript we would be willing to defend in a customer dispute?
Build the reference set
Pull roughly an hour of real recorded calls spanning your actual conditions — mobile callers, speakerphone, accents, hold music bleed, your product and place names. Have a human produce a verbatim reference transcript. This human pass is the only expensive part of the exercise, and it is the part you cannot skip.
Transcribe the same hour on each engine
One hour through Grok Voice Think Fast 2.0's raw API lists at $4.80. Run the same audio through your incumbent engine and through one dedicated alternative. Normalise casing, punctuation and number formatting identically across all outputs before scoring, or you will measure formatting conventions instead of accuracy.
Score WER, then score what breaks deals
Report overall word error rate, then separately count errors on the entities that actually matter: names, addresses, order numbers, dates, amounts, and any consent or compliance language. A model with a better headline WER that fumbles order numbers is the worse model for a CRM.
Re-run on your worst audio
xAI's largest claimed advantage — roughly ten times — is specifically in noisy and telephony-compressed conditions. That is a testable prediction. Isolate your ugliest 15 minutes of recordings and re-score. If the advantage does not widen there, the vendor claim does not describe your traffic.
One methodological warning drawn straight from the table above: score each engine in the configuration you will actually deploy. The 4.8-fold Nova-3 spread exists because batch transcription and live-agent transcription are different products wearing the same model name. If you benchmark in batch mode and deploy in streaming mode, your production accuracy will not resemble your test results.
06 — Pricing SurfacesFour different prices for one product family.
Voice pricing is where budget forecasts go wrong, because the same vendor sells the same capability on several surfaces at different rates and coverage regularly quotes one as another. xAI alone lists a raw speech-to-speech model API, a no-code platform, a dedicated speech-to-text component and a dedicated text-to-speech component. Here is every published rate in one place, labelled by surface.
| Product | Surface | Published price | Per-minute equivalent |
|---|---|---|---|
| xAI | |||
| Grok Voice Think Fast 2.0 | Raw model API (speech-to-speech) | $0.08/min audio ($4.80/hr), plus $0.004 per conversation.item.create text-input event | $0.08 |
| Grok Voice Think Fast 1.0 | Raw model API (speech-to-speech) | $0.05/min audio ($3.00/hr) | $0.05 |
| Grok Voice Agent Builder | No-code platform | $0.05/min with voices included, plus $0.01/min for a provisioned phone number — as advertised at its July 1, 2026 launch | $0.05 + $0.01 |
| xAI Speech-to-Text | Dedicated component API | $0.10/hr REST, $0.20/hr streaming | $0.0017 / $0.0033 |
| xAI Text-to-Speech | Dedicated component API | $15.00 per 1M characters | Not per-minute; depends on characters spoken |
| OpenAI | |||
| gpt-realtime-2.1 (audio) | Standard list, per 1M tokens | $32.00 input, $0.40 cached input, $64.00 output | Token-priced; no vendor per-minute rate |
| gpt-realtime-2.1-mini (audio) | Standard list, per 1M tokens | $10.00 input, $0.30 cached input, $20.00 output | Token-priced; no vendor per-minute rate |
| gpt-live-transcribe / gpt-realtime-whisper | Standard list, dedicated transcription | $0.017/min each | $0.017 |
| gpt-realtime-translate | Standard list, dedicated translation | $0.034/min | $0.034 |
| Gemini 3.1 Flash Live (audio) | Standard list, per 1M tokens | $3.00 audio input, $12.00 audio output | $0.005 in / $0.018 out, per Google’s own stated equivalents |
All rates are vendor standard list as published on xAI, OpenAI and Google pricing pages, retrieved August 2026. None are batch, flex or reseller rates. xAI speech-to-text per-minute equivalents computed as hourly ÷ 60: $0.10 ÷ 60 = $0.0017 and $0.20 ÷ 60 = $0.0033. Grok hourly figures divide exactly: $4.80 ÷ 60 = $0.08 and $3.00 ÷ 60 = $0.05. OpenAI publishes Realtime by token, not by minute; the third-party per-minute estimates in circulation for it are derived, not vendor-published.
One structural note that the table makes visible. Think Fast 2.0 costs 60% more per minute than Think Fast 1.0 ($0.08 against $0.05) while using roughly 60% fewer reasoning tokens per response. Those two figures are unrelated in cause — one is a list price, the other is an internal efficiency measure — but their relationship is the point. Under per-minute billing, a token-efficiency gain accrues to the vendor’s margin, not to your invoice. Efficiency improvements only reach the buyer when pricing is token-metered, which is precisely how OpenAI and Google meter their realtime audio and xAI does not.
If you want a sense of how fast the surrounding component market is repricing, our look at the text-to-speech price war reshaping the voice stack covers the other half of the cascade.
07 — The Real DecisionConversation quality versus a defensible transcript.
Strip away the leaderboard and the choice in front of a team wiring voice into a CRM is architectural. A speech-to-speech model hears audio and emits audio; the reasoning happens inside one network, which is why it can answer in 0.70 seconds. A cascade runs speech-to-text, then a language model, then text-to-speech; every stage adds latency, and the text between stages is the thing you can store, search, redact and put in front of a compliance officer.
xAI is unusual in selling both halves. You can build a cascade entirely on its own stack — dedicated speech-to-text at $0.10 per hour on REST or $0.20 per hour streaming, then a text model, then text-to-speech at $15 per million characters. Run the arithmetic on a thousand minutes of calls and the shape of the trade becomes concrete: the speech-to-speech path bills 1,000 × $0.08 = $80.00, while the transcription leg of a cascade on the streaming tier bills 1,000 ÷ 60 × $0.20 = $3.33, about 4% of the speech-to-speech line. The language and speech legs are extra and depend on your token and character volumes, but transcription itself is not what makes a cascade expensive. Latency is.
Speech-to-speech, end to end
When a caller is waiting, sub-second time to first audio and clean turn-taking decide whether the call survives. Conversational Dynamics at 95.1 is the sub-benchmark that maps to this, and it is where Think Fast 2.0 made almost all of its gain.
Cascaded STT into LLM into TTS
When a transcript is the record — consent capture, financial instructions, anything a customer may dispute — you want an auditable text artefact from a component you can benchmark, version and swap independently. Accept the extra latency.
Test both on your own audio
Vendor multilingual claims are the least transferable of all published numbers, because the language mix and phrase difficulty are chosen by the vendor. Score both architectures on your real call mix before committing. The entity error rate matters more than the headline.
Speak fast, transcribe separately
Nothing stops you running speech-to-speech for the live conversation and a dedicated transcription pass over the recording afterwards for the CRM record. At around $0.0033 per minute on streaming STT, the second pass is a rounding error against the conversation itself.
That last row is the one most teams end up at, and it is worth stating plainly because no vendor will pitch it. The two architectures are not mutually exclusive. Speech-to-speech optimises the caller’s experience; a separate transcription pass optimises the record. They cost different things, fail in different ways, and can be swapped independently. Buying one model to do both jobs is a coupling decision, and coupling is what makes the vendor’s unaudited WER claim load-bearing in the first place.
The other capability worth weighing on the speech-to-speech side: xAI’s speech-to-speech API supports function calling, web and X search, Collections search and remote MCP tool calls inside the live session. That is what turns a voice model into a genuine CRM front-end rather than a talking FAQ — the agent can look a record up and act on it mid-call. Wiring that safely into a live customer database is the work we do in CRM and marketing automation engagements, and the tool layer is usually where the real project sits.
08 — ActionThree things worth doing this week.
There is a hard date attached to this release. xAI announced on July 29 that on August 5, 2026 the grok-voice-latest alias moves from grok-voice-think-fast-1.0 to grok-voice-think-fast-2.0, with no action required to upgrade. Teams that want to stay on 1.0 must pin the explicit model id before that date.
Read that as a cost event as much as a capability one. Anything pointing at the alias moves from a $0.05 to a $0.08 per-minute audio rate on August 5 without a deployment, a changelog entry, or a line in anyone’s sprint. A support line running a few thousand minutes a month absorbs a 60% increase in its audio line item silently. Whether or not you want the new model — and on these numbers most teams will — the pin-or-accept decision should be made deliberately rather than by default.
Two operational constraints belong in the same review. xAI publishes a limit of 10 concurrent speech-to-speech sessions per team and a 120-minute maximum session duration. Ten concurrent calls is a small number for anything resembling a real contact centre, and it is a capacity planning question to resolve before a pilot becomes a rollout, not after.
Looking forward, the sub-benchmark decomposition suggests where the next twelve months go. Conversational dynamics is close to saturated at 95.1 and speech reasoning has been saturated for a generation. The open leg is agentic performance at 56.5 — whether the agent completes the task. We expect the next round of releases to compete there, and we expect the benchmark conversation to shift from how natural the model sounds to how often it finishes the job without a human. That is also the axis on which a voice agent either earns or loses its place in a CRM workflow. If you are scoping that kind of build, our AI transformation engagements start with exactly this sort of comparative evaluation before a single line of integration code gets written.
09 — ConclusionTake the claim seriously. Then test it.
A number you cannot check is not a number you can budget against.
Grok Voice Think Fast 2.0 is a real step. 82.9 on the Artificial Analysis Speech-to-Speech Quality Index at 0.70 seconds to first audio is a combination nothing else on that board currently offers, and the gain is concentrated in exactly the sub-benchmark a phone line cares about. At $0.08 per minute it is competitively priced against GPT-Realtime-2.1, and it is meaningfully faster than the one model that outscores it.
The transcription claim is a different kind of statement. xAI says its model beats dedicated speech-to-text engines by 1.5 to 2.0 times, and roughly ten times in noise. That may well be true. But it is published as a relative multiplier against a vendor-selected baseline, on a vendor-selected corpus, with no absolute word error rate and no independent replication. The same industry has published three different Deepgram Nova-3 figures spanning a roughly 4.8-fold range, all of them honest. That is not a reason to disbelieve xAI. It is a reason to treat any single WER number as a statement about a test harness.
The practical move is unglamorous and cheap. An hour of your own recorded calls, one human reference transcript, and about five dollars of API time answers the only question that matters — whether this model transcribes your callers well enough to log. Decide the architecture on that evidence, keep the live conversation and the stored record as separately swappable pieces, and make the August 5 alias decision on purpose. The leaderboard will move again by October. Your call audio will not.