CRM & AutomationNew Release17 min readPublished August 1, 2026

Speech-to-speech, no cascade · 0.70s to first audio · 82.9 quality index

Grok Voice 2.0 and the State of Speech-to-Speech Agents

xAI shipped Grok Voice Think Fast 2.0 on July 29 — 82.9 on the Artificial Analysis Speech-to-Speech Quality Index, time-to-first-audio down to 0.70 seconds, and $0.08 per minute of audio on the raw API. It also claims its transcription beats dedicated speech-to-text engines. That claim is worth taking seriously, and worth testing yourself before those transcripts land in a CRM.

DA
Digital Applied Team
Senior strategists · Published Aug 1, 2026
PublishedAug 1, 2026
Read time17 min
SourcesxAI, Artificial Analysis, OpenAI, Google, Coval
AA quality index
82.9
Speech-to-Speech Index
+7.2 vs Think Fast 1.0
Time to first audio
0.70s
vendor-reported
from 1.25s
Raw API audio price
$0.08
per minute ($4.80/hr)
TF 1.0 was $0.05
Reasoning tokens
0.4×
P50 vs Think Fast 1.0
~60% fewer

Speech-to-speech agents crossed a threshold this week. xAI — which now brands itself SpaceXAI — released Grok Voice Think Fast 2.0 on July 29, 2026: a model that listens, reasons and speaks inside a single network rather than chaining speech-to-text into a language model into text-to-speech. It scores 82.9 on the Artificial Analysis Speech-to-Speech Quality Index, up from 75.7, and it starts talking back in 0.70 seconds.

The number that will get quoted, though, is a different one. xAI states that Think Fast 2.0 transcribes 1.5 to 2.0 times more accurately than Deepgram Nova 3 and ElevenLabs Scribe v2, with the gap widening to roughly ten times in noisy conditions. If that holds, a single voice model just absorbed a job that has belonged to dedicated transcription engines for a decade — and every CRM that logs call transcripts has a vendor decision to reconsider.

This guide takes that claim at face value and then shows you exactly where the evidence stops. We decompose the quality index into its three sub-benchmarks, put Grok next to the model that actually outscores it, separate four pricing surfaces that vendors and coverage routinely conflate, and end with the architectural question that matters more than any leaderboard: whether your CRM needs a great conversation or a defensible transcript.

Key takeaways
  1. 01
    Think Fast 2.0 is fast first, and top-tier second.82.9 on the Artificial Analysis Speech-to-Speech Quality Index at 0.70s time-to-first-audio. Qwen Audio 3.0 Realtime Plus scores higher at 84.1 — but takes about 4 seconds to start speaking, roughly 5.7 times longer.
  2. 02
    Almost all of the gain came from one sub-benchmark.The index is an equal-weighted mean of Speech Reasoning, Conversational Dynamics and Agentic Performance. Conversational Dynamics jumped 77.8 to 95.1 and accounts for about 79% of the raw sub-score improvement. Speech reasoning barely moved.
  3. 03
    The transcription claim is credible but unaudited.xAI reports a 1.5 to 2.0 times word-error-rate improvement over Deepgram Nova 3 and ElevenLabs Scribe v2 on its own eval of thousands of short phrases in 24 languages. No absolute word error rate is published, and no independent party has replicated it.
  4. 04
    Word error rate moves by roughly 4.8 times on configuration alone.The same Deepgram Nova-3 model has been measured at 5.26% on Deepgram's own batch audio and 25.3% by Coval in a latency-optimised voice-agent configuration — a 4.8-fold spread. Any single WER figure is a statement about a test harness, not a model.
  5. 05
    Four pricing surfaces, one product family.The raw model API is $0.08/min. Think Fast 1.0 was $0.05/min. Agent Builder advertised $0.05/min at its July 1 launch. xAI also sells dedicated speech-to-text at $0.10/hr REST. Whether the platform price moves when the default model flips on August 5 is unconfirmed.

01What ShippedA model that reasons while it speaks.

Grok Voice Think Fast 2.0 landed on July 29, 2026 as xAI’s next-generation voice model, positioned on three axes: intelligence, transcription accuracy, and conversational capability. The architectural bet is that a single speech-to-speech network beats a cascade of specialists on the things a live conversation actually rewards — turn-taking, interruption handling, and the gap between the end of your sentence and the start of the reply.

The headline latency figure moved from 1.25 seconds to 0.70 seconds to first audio. That is the number a caller feels. Everything else on the spec sheet is downstream of it: xAI also reports that median reasoning tokens per response fell to 0.4 times the Think Fast 1.0 baseline, so the model reaches an answer with roughly 60% less internal deliberation than its predecessor while scoring higher.

xAI describes tuning the model toward the patterns of real human conversation — shorter sentences, one question at a time, less filler. That is a design goal rather than a measured claim, and no benchmark number is attached to it. It does, however, explain where the benchmark movement came from, which we get to in the next section.

The model also has a no-code sibling. About four weeks earlier, xAI launched Grok Voice Agent Builder, the knowledge-base-and-connectors layer that sits on top of this same model family; we covered xAI’s no-code Voice Agent Builder at launch. That distinction matters later, because the two products are priced on different surfaces.

Speech-to-speech
Grok Voice Think Fast 2.0
$0.08/min audio · $4.80/hr · raw model API

Single-model speech in, speech out. 0.70s to first audio, 82.9 AA quality index. Supports function calling, web and X search, Collections search, and remote MCP tool calls inside the voice session. Rate limits: 10 concurrent sessions per team, 120-minute maximum session length.

docs.x.ai/developers/models/speech-to-speech
Previous generation
Grok Voice Think Fast 1.0
$0.05/min audio · $3.00/hr · raw model API

1.25s to first audio, 75.7 AA quality index. The August 5 flip moves the grok-voice-latest alias to 2.0; xAI's stated instruction for staying on this generation is to pin grok-voice-think-fast-1.0 before then. Cheaper per minute, and meaningfully weaker on conversational dynamics and agentic tasks.

Pin before Aug 5 to stay on it
One vendor claim to discount
xAI reports that A/B testing Think Fast 2.0 on the Starlink support line produced a significant increase in sales conversion rate and support containment rate. No percentages, no absolute figures, and no baseline were disclosed. Treat it as directional vendor signalling, not evidence — a containment-rate claim with no denominator cannot be compared against anything you already run.

02Benchmark AnatomyWhere the 7.2 points actually came from.

The Artificial Analysis Speech-to-Speech Quality Index is not a single test. It is an equal-weighted mean — one third each — of three sub-benchmarks: Speech Reasoning (Big Bench Audio), Conversational Dynamics (Full Duplex Bench), and Agentic Performance (τ-Voice). A model has to post a valid score on all three to appear in the index at all.

That structure makes the headline decomposable, and decomposing it is more useful than the headline. Below are the sub-scores xAI published for both generations, with the change and each sub-benchmark’s share of the total raw improvement recomputed from those figures.

Grok Voice Think Fast 2.0 versus 1.0 across the three sub-benchmarks of the Artificial Analysis Speech-to-Speech Quality Index, with the point change and each sub-benchmark’s share of the total raw improvement.
Sub-benchmarkTF 2.0TF 1.0ChangeShare of gain
Speech ReasoningBig Bench Audio97.297.1+0.10.5%
Conversational DynamicsFull Duplex Bench95.177.8+17.379.4%
Agentic Performanceτ-Voice Bench56.552.1+4.420.2%
Quality Indexequal-weighted mean of the three82.975.7+7.2

Sub-scores as published by xAI, citing Artificial Analysis. Share of gain = each sub-benchmark’s point change divided by the sum of the three changes (0.1 + 17.3 + 4.4 = 21.8). Shares are rounded to one decimal and therefore sum to 100.1%. Index rows are the published values; the equal-weighted mean of the 2.0 sub-scores is 82.93 and of the 1.0 sub-scores is 75.67, both consistent with the published figures.

Read that table twice, because it reframes the release. Speech reasoning was already effectively saturated — 97.1 to 97.2 is noise. Agentic performance improved meaningfully in relative terms but is still the weakest leg by a wide margin at 56.5. Nearly four fifths of the visible improvement is Conversational Dynamics: turn-taking, barge-in, knowing when to stop talking. This is not a model that got smarter. It is a model that got better at holding a conversation, which for a phone line is arguably the more valuable upgrade.

It also sets a ceiling worth noticing. Full Duplex Bench at 95.1 leaves very little headroom on that leg. The next generation’s index movement will have to come from agentic performance — the 56.5 — which is the sub-benchmark that measures whether a voice agent can actually complete a task rather than merely sound good doing it. For a CRM front-end, that is the number to track across the next two releases.

How the index is built
Artificial Analysis constructs the Speech-to-Speech Quality Index as an equal-weighted average of Speech Reasoning, Conversational Dynamics and Agentic Performance results, and requires a valid score on all three before a model is listed. That matters when reading any single headline figure: a model can move the index several points without improving at the thing you are buying it for. Methodology is published at artificialanalysis.ai/methodology/speech-to-speech-benchmarking.

03The FieldFastest among the leaders, not the highest scorer.

xAI’s own comparison chart puts Think Fast 2.0 against GPT-Realtime-2.1 and Gemini 3.1 Flash. It does not include Qwen Audio 3.0 Realtime Plus, which on the same Artificial Analysis index scores 84.1 — higher than Grok’s 82.9. That omission is worth naming, and it is also worth not overreading, because the Qwen model takes about 4 seconds to produce its first audio against Grok’s 0.70. That is a real engineering trade, not a gotcha.

Speech-to-Speech Quality Index · leaderboard snapshot

Source: Artificial Analysis Speech-to-Speech Index, retrieved August 2026
Qwen Audio 3.0 Realtime PlusHighest index in the set · ~4.02s to first audio
84.1
Grok Voice Think Fast 2.00.70s to first audio · fastest among the leaders
82.9
GPT-Realtime-2.1 (High)1.21s to first audio
79.1
GPT-Realtime-2 (High)1.14s to first audio
77.2
Qwen Audio 3.0 Realtime Flash~4.16s to first audio
76.3
Grok Voice Think Fast 1.01.25s to first audio
75.7
Gemini 3.1 Flash (High)~2.99s to first audio
69.5
Gemini 3.1 Flash (Minimal)0.96s to first audio · cheapest listed
56.6

Quality alone is the wrong lens for a phone line, so the table below adds latency and Artificial Analysis’s own hourly price column, plus a latency multiple computed against Grok Voice Think Fast 2.0’s 0.70 seconds. No model wins all three columns.

Cross-vendor speech-to-speech scorecard comparing quality index, time to first audio, latency multiple relative to Grok Voice Think Fast 2.0, and the hourly price listed by Artificial Analysis.
ModelQuality indexTime to first audioLatency vs Grok 2.0AA price / hr
Qwen Audio 3.0 Realtime Plus84.14.02s5.7×$4.42
Grok Voice Think Fast 2.082.90.70s1.0×$4.80
GPT-Realtime-2.1 (High)79.11.21s1.7×$10.75
GPT-Realtime-2 (High)77.21.14s1.6×$4.14
Qwen Audio 3.0 Realtime Flash76.34.16s5.9×$4.77
Grok Voice Think Fast 1.075.71.25s1.8×$3.00
Gemini 3.1 Flash (High)69.52.99s4.3×$1.75
Gemini 3.1 Flash (Minimal)56.60.96s1.4×$1.50

Index, latency and hourly price as listed on the Artificial Analysis Speech-to-Speech leaderboard, retrieved August 2026. The hourly column is Artificial Analysis’s own blended figure, not a vendor-published hourly rate: OpenAI and Google meter realtime audio by token, so their rows are derived rather than quoted. xAI’s two rows do match its published per-minute pricing. Latency vs Grok 2.0 = each model’s time to first audio divided by 0.70s, rounded to one decimal. GPT-Realtime-2.1’s latency is shown as a dash in xAI’s own chart; the 1.21s figure here is Artificial Analysis’s, not xAI’s.

Three readings fall out of it. First, the top of the board is tight: 1.2 index points separate Qwen Plus from Grok 2.0, which is well inside the range where prompt design and voice selection will matter more than model choice. Second, latency is where the field genuinely splits — Grok 2.0 is roughly 5.7 times faster to first audio than the model that outscores it, and that difference is audible on a phone call in a way that 1.2 index points is not. Third, price does not track quality: GPT-Realtime-2.1 carries the highest hourly figure in the set at $10.75 — an Artificial Analysis blended number, not an OpenAI list rate — and sits third on quality.

Our reading is that speech-to-speech has stopped being a capability race and become a positioning one. Gemini 3.1 Flash Minimal is on the board at an AA-blended $1.50 an hour and 0.96 seconds — genuinely fast and genuinely cheap, at 56.6 quality. Qwen Plus takes the opposite corner. Grok 2.0 is the only entry currently sitting near the top of both the quality and the latency columns at once, and that combination, not the 82.9, is what it is actually selling. For the fuller picture of how these pieces assemble into a production system, our voice agent infrastructure stack reference covers why latency is the KPI that ends up governing everything else.

04The Transcription ClaimA bold claim with no absolute number attached.

This is the part of the release that deserves the most scrutiny, and it is worth being precise about what xAI actually said, because it is the opposite of what a voice model’s vendor usually concedes.

"Grok Voice Think Fast 2.0 outperforms even dedicated, state of the art transcription models when it comes to accuracy."— xAI, Grok Voice Think Fast 2.0 announcement, July 29, 2026

The supporting detail: an evaluation across thousands of short phrases in 24 different languages showing a 1.5 to 2.0 times improvement in word error rate relative to Deepgram Nova 3 and ElevenLabs Scribe v2, and a 1.4 times improvement over Think Fast 1.0. xAI further states that the gap widens to roughly ten times in noisy settings, and frames that as deliberate: the model was tuned for real-world audio with substantial background noise and telephony compression.

Take it seriously. A speech-to-speech model that ingests raw audio never passes through a lossy text bottleneck, so there is a plausible mechanism for it to hear things a transcription-first pipeline drops. The noisy-audio result is exactly where you would expect that advantage to show up.

Now the parts that are missing, all of which are checkable rather than speculative. The comparison models were chosen by xAI. The language set was chosen by xAI. The phrase set was chosen by xAI, and short phrases are a specific and relatively forgiving transcription regime. The results are published as relative multipliers on a chart with no visible axis values — no absolute word error rate appears anywhere. And as of publication, no independent benchmark, academic group or competing vendor has replicated the comparison.

Four published measurements of speech-to-text accuracy, showing who ran each test, what it measured, and the absolute word error rate reported.
Who measuredModel measuredTest conditionsAbsolute WER
xAI (vendor)Grok Voice Think Fast 2.0Thousands of short phrases, 24 languages. Vendor-selected comparison set. Relative multipliers only.Not published
Artificial AnalysisDeepgram Nova-3AA-WER v2. Duration-weighted average across three datasets, about 8 hours of audio.5.2%
Artificial AnalysisElevenLabs Scribe v2Same AA-WER v2 harness and datasets.2.2%
Coval (third party)Deepgram Nova-3Latency-optimised voice-agent configuration rather than batch transcription.25.3%
Deepgram (vendor)Deepgram Nova-3Vendor-claimed median on batch audio.5.26%

Sources: xAI announcement (July 29, 2026); Artificial Analysis speech-to-text index, retrieved August 2026; Coval and Deepgram figures via secondary reporting. Spread on the same model = 25.3 ÷ 5.26 = 4.81, i.e. roughly a 4.8× difference between the highest and lowest published Nova-3 figures.

Look at the three Nova-3 rows together. The identical model has been published at 5.2%, 5.26% and 25.3% word error rate — roughly a 4.8-fold spread — depending entirely on whether it was configured for batch accuracy or for live-agent latency. Nobody in that list is lying. They are measuring different things and calling both of them WER.

That is the honest frame for xAI’s claim. The problem is not that a 1.5 to 2.0 times improvement is implausible; it is that a relative multiplier against a vendor-chosen baseline, on a vendor-chosen corpus, with no absolute figure to anchor it, cannot be compared to anything you already measure. It is unfalsifiable in its published form — not wrong, just not yet checkable. And the same logic applies to a self-hosted Whisper transcription layer: the number that matters is the one your own audio produces.

05VerificationHow to test it yourself in an afternoon.

The good news about an unaudited claim is that this one is cheap to audit. Word error rate is a mechanical measure — substitutions plus deletions plus insertions, divided by reference words — and the audio you care about is already sitting in your call recordings. The whole exercise costs a few dollars of API time.

The design principle is that you are not trying to reproduce xAI’s benchmark. You are trying to answer a narrower and far more useful question: on our calls, with our accents, product names and line quality, which engine produces the transcript we would be willing to defend in a customer dispute?

Sample size
Build the reference set
60min

Pull roughly an hour of real recorded calls spanning your actual conditions — mobile callers, speakerphone, accents, hold music bleed, your product and place names. Have a human produce a verbatim reference transcript. This human pass is the only expensive part of the exercise, and it is the part you cannot skip.

Human-transcribed ground truth
Cost to run
Transcribe the same hour on each engine
$4.80

One hour through Grok Voice Think Fast 2.0's raw API lists at $4.80. Run the same audio through your incumbent engine and through one dedicated alternative. Normalise casing, punctuation and number formatting identically across all outputs before scoring, or you will measure formatting conventions instead of accuracy.

$0.08/min × 60 = $4.80
What to score
Score WER, then score what breaks deals
3

Report overall word error rate, then separately count errors on the entities that actually matter: names, addresses, order numbers, dates, amounts, and any consent or compliance language. A model with a better headline WER that fumbles order numbers is the worse model for a CRM.

Overall WER + entity error rate
Noise condition
Re-run on your worst audio
10×

xAI's largest claimed advantage — roughly ten times — is specifically in noisy and telephony-compressed conditions. That is a testable prediction. Isolate your ugliest 15 minutes of recordings and re-score. If the advantage does not widen there, the vendor claim does not describe your traffic.

Vendor's own strongest claim

One methodological warning drawn straight from the table above: score each engine in the configuration you will actually deploy. The 4.8-fold Nova-3 spread exists because batch transcription and live-agent transcription are different products wearing the same model name. If you benchmark in batch mode and deploy in streaming mode, your production accuracy will not resemble your test results.

06Pricing SurfacesFour different prices for one product family.

Voice pricing is where budget forecasts go wrong, because the same vendor sells the same capability on several surfaces at different rates and coverage regularly quotes one as another. xAI alone lists a raw speech-to-speech model API, a no-code platform, a dedicated speech-to-text component and a dedicated text-to-speech component. Here is every published rate in one place, labelled by surface.

Published voice pricing across xAI, OpenAI and Google, labelled by pricing surface with unit price and per-minute equivalent where one can be computed.
ProductSurfacePublished pricePer-minute equivalent
xAI
Grok Voice Think Fast 2.0Raw model API (speech-to-speech)$0.08/min audio ($4.80/hr), plus $0.004 per conversation.item.create text-input event$0.08
Grok Voice Think Fast 1.0Raw model API (speech-to-speech)$0.05/min audio ($3.00/hr)$0.05
Grok Voice Agent BuilderNo-code platform$0.05/min with voices included, plus $0.01/min for a provisioned phone number — as advertised at its July 1, 2026 launch$0.05 + $0.01
xAI Speech-to-TextDedicated component API$0.10/hr REST, $0.20/hr streaming$0.0017 / $0.0033
xAI Text-to-SpeechDedicated component API$15.00 per 1M charactersNot per-minute; depends on characters spoken
OpenAI
gpt-realtime-2.1 (audio)Standard list, per 1M tokens$32.00 input, $0.40 cached input, $64.00 outputToken-priced; no vendor per-minute rate
gpt-realtime-2.1-mini (audio)Standard list, per 1M tokens$10.00 input, $0.30 cached input, $20.00 outputToken-priced; no vendor per-minute rate
gpt-live-transcribe / gpt-realtime-whisperStandard list, dedicated transcription$0.017/min each$0.017
gpt-realtime-translateStandard list, dedicated translation$0.034/min$0.034
Google
Gemini 3.1 Flash Live (audio)Standard list, per 1M tokens$3.00 audio input, $12.00 audio output$0.005 in / $0.018 out, per Google’s own stated equivalents

All rates are vendor standard list as published on xAI, OpenAI and Google pricing pages, retrieved August 2026. None are batch, flex or reseller rates. xAI speech-to-text per-minute equivalents computed as hourly ÷ 60: $0.10 ÷ 60 = $0.0017 and $0.20 ÷ 60 = $0.0033. Grok hourly figures divide exactly: $4.80 ÷ 60 = $0.08 and $3.00 ÷ 60 = $0.05. OpenAI publishes Realtime by token, not by minute; the third-party per-minute estimates in circulation for it are derived, not vendor-published.

Open question, not a settled one
Grok Voice Agent Builder advertised $0.05 per minute at its July 1 launch — a figure that matched Think Fast 1.0’s raw API rate exactly. The default model underneath it moves to Think Fast 2.0, whose raw rate is $0.08 per minute. We could not find any source, vendor or otherwise, confirming whether the platform price changed, absorbed the difference, or stayed put. If Agent Builder is in your budget model, treat that line as unpriced until xAI restates it — do not assume either direction. Our earlier walkthrough of xAI’s no-code Voice Agent Builder documents the launch-day pricing as it stood.

One structural note that the table makes visible. Think Fast 2.0 costs 60% more per minute than Think Fast 1.0 ($0.08 against $0.05) while using roughly 60% fewer reasoning tokens per response. Those two figures are unrelated in cause — one is a list price, the other is an internal efficiency measure — but their relationship is the point. Under per-minute billing, a token-efficiency gain accrues to the vendor’s margin, not to your invoice. Efficiency improvements only reach the buyer when pricing is token-metered, which is precisely how OpenAI and Google meter their realtime audio and xAI does not.

If you want a sense of how fast the surrounding component market is repricing, our look at the text-to-speech price war reshaping the voice stack covers the other half of the cascade.

07The Real DecisionConversation quality versus a defensible transcript.

Strip away the leaderboard and the choice in front of a team wiring voice into a CRM is architectural. A speech-to-speech model hears audio and emits audio; the reasoning happens inside one network, which is why it can answer in 0.70 seconds. A cascade runs speech-to-text, then a language model, then text-to-speech; every stage adds latency, and the text between stages is the thing you can store, search, redact and put in front of a compliance officer.

xAI is unusual in selling both halves. You can build a cascade entirely on its own stack — dedicated speech-to-text at $0.10 per hour on REST or $0.20 per hour streaming, then a text model, then text-to-speech at $15 per million characters. Run the arithmetic on a thousand minutes of calls and the shape of the trade becomes concrete: the speech-to-speech path bills 1,000 × $0.08 = $80.00, while the transcription leg of a cascade on the streaming tier bills 1,000 ÷ 60 × $0.20 = $3.33, about 4% of the speech-to-speech line. The language and speech legs are extra and depend on your token and character volumes, but transcription itself is not what makes a cascade expensive. Latency is.

Live sales and support lines
Speech-to-speech, end to end

When a caller is waiting, sub-second time to first audio and clean turn-taking decide whether the call survives. Conversational Dynamics at 95.1 is the sub-benchmark that maps to this, and it is where Think Fast 2.0 made almost all of its gain.

Pick speech-to-speech
Regulated or disputed records
Cascaded STT into LLM into TTS

When a transcript is the record — consent capture, financial instructions, anything a customer may dispute — you want an auditable text artefact from a component you can benchmark, version and swap independently. Accept the extra latency.

Pick the cascade
Multilingual, mixed conditions
Test both on your own audio

Vendor multilingual claims are the least transferable of all published numbers, because the language mix and phrase difficulty are chosen by the vendor. Score both architectures on your real call mix before committing. The entity error rate matters more than the headline.

Benchmark, do not assume
Hybrid, in practice
Speak fast, transcribe separately

Nothing stops you running speech-to-speech for the live conversation and a dedicated transcription pass over the recording afterwards for the CRM record. At around $0.0033 per minute on streaming STT, the second pass is a rounding error against the conversation itself.

Often the right answer

That last row is the one most teams end up at, and it is worth stating plainly because no vendor will pitch it. The two architectures are not mutually exclusive. Speech-to-speech optimises the caller’s experience; a separate transcription pass optimises the record. They cost different things, fail in different ways, and can be swapped independently. Buying one model to do both jobs is a coupling decision, and coupling is what makes the vendor’s unaudited WER claim load-bearing in the first place.

The other capability worth weighing on the speech-to-speech side: xAI’s speech-to-speech API supports function calling, web and X search, Collections search and remote MCP tool calls inside the live session. That is what turns a voice model into a genuine CRM front-end rather than a talking FAQ — the agent can look a record up and act on it mid-call. Wiring that safely into a live customer database is the work we do in CRM and marketing automation engagements, and the tool layer is usually where the real project sits.

08ActionThree things worth doing this week.

There is a hard date attached to this release. xAI announced on July 29 that on August 5, 2026 the grok-voice-latest alias moves from grok-voice-think-fast-1.0 to grok-voice-think-fast-2.0, with no action required to upgrade. Teams that want to stay on 1.0 must pin the explicit model id before that date.

Read that as a cost event as much as a capability one. Anything pointing at the alias moves from a $0.05 to a $0.08 per-minute audio rate on August 5 without a deployment, a changelog entry, or a line in anyone’s sprint. A support line running a few thousand minutes a month absorbs a 60% increase in its audio line item silently. Whether or not you want the new model — and on these numbers most teams will — the pin-or-accept decision should be made deliberately rather than by default.

Two operational constraints belong in the same review. xAI publishes a limit of 10 concurrent speech-to-speech sessions per team and a 120-minute maximum session duration. Ten concurrent calls is a small number for anything resembling a real contact centre, and it is a capacity planning question to resolve before a pilot becomes a rollout, not after.

Looking forward, the sub-benchmark decomposition suggests where the next twelve months go. Conversational dynamics is close to saturated at 95.1 and speech reasoning has been saturated for a generation. The open leg is agentic performance at 56.5 — whether the agent completes the task. We expect the next round of releases to compete there, and we expect the benchmark conversation to shift from how natural the model sounds to how often it finishes the job without a human. That is also the axis on which a voice agent either earns or loses its place in a CRM workflow. If you are scoping that kind of build, our AI transformation engagements start with exactly this sort of comparative evaluation before a single line of integration code gets written.

09ConclusionTake the claim seriously. Then test it.

Speech-to-speech agents, August 2026

A number you cannot check is not a number you can budget against.

Grok Voice Think Fast 2.0 is a real step. 82.9 on the Artificial Analysis Speech-to-Speech Quality Index at 0.70 seconds to first audio is a combination nothing else on that board currently offers, and the gain is concentrated in exactly the sub-benchmark a phone line cares about. At $0.08 per minute it is competitively priced against GPT-Realtime-2.1, and it is meaningfully faster than the one model that outscores it.

The transcription claim is a different kind of statement. xAI says its model beats dedicated speech-to-text engines by 1.5 to 2.0 times, and roughly ten times in noise. That may well be true. But it is published as a relative multiplier against a vendor-selected baseline, on a vendor-selected corpus, with no absolute word error rate and no independent replication. The same industry has published three different Deepgram Nova-3 figures spanning a roughly 4.8-fold range, all of them honest. That is not a reason to disbelieve xAI. It is a reason to treat any single WER number as a statement about a test harness.

The practical move is unglamorous and cheap. An hour of your own recorded calls, one human reference transcript, and about five dollars of API time answers the only question that matters — whether this model transcribes your callers well enough to log. Decide the architecture on that evidence, keep the live conversation and the stored record as separately swappable pieces, and make the August 5 alias decision on purpose. The leaderboard will move again by October. Your call audio will not.

Put a voice agent in front of your CRM

A voice agent is only as good as the record it leaves behind.

We design, benchmark and ship voice agents that write into real CRM systems — model selection on your own call audio, transcript fidelity you can defend, and the tool layer that lets an agent actually resolve the call.

Free consultationExpert guidanceTailored solutions
What we work on

Voice agent engagements

  • Model bake-offs scored on your own recorded calls
  • Speech-to-speech vs cascaded pipeline architecture
  • Transcript fidelity and entity-level error auditing
  • CRM write-back, tool calling and MCP integration
  • Per-minute cost modelling across pricing surfaces
FAQ · Grok Voice 2.0 and speech-to-speech

The questions we get every week.

Grok Voice Think Fast 2.0 is xAI's next-generation speech-to-speech voice model, announced on July 29, 2026. Rather than chaining a transcription model into a language model into a text-to-speech model, it processes audio in and produces audio out inside a single network. xAI reports 82.9 on the Artificial Analysis Speech-to-Speech Quality Index, up from 75.7 for Think Fast 1.0, and time-to-first-audio down from 1.25 seconds to 0.70 seconds. Median reasoning tokens per response fell to 0.4 times the previous generation's baseline, roughly 60% fewer. It is available on xAI's WebSocket speech-to-speech API at $0.08 per minute of audio, and it supports function calling, web and X search, Collections search and remote MCP tool calls inside the live voice session.
Related dispatches

Continue exploring voice and CRM.

CRM & Automation

Qwen-Audio-3.0-TTS Tops the Arena at a Third of the Price

Qwen-Audio-3.0-TTS-Plus tops the Artificial Analysis TTS arena (~1,236 Elo) at ~$27.59/1M chars, a third of ElevenLabs — but it is hosted-only, no weights.

July 22, 2026 · 10 minRead
CRM & Automation

Build a WhatsApp Lead-Capture Agent for Your CRM Stack

Build a WhatsApp lead-capture agent on Meta's Cloud API: inbound webhook, in-chat qualification with Flows, and a routed CRM record, with no BSP required.

July 15, 2026 · 12 minRead
CRM & Automation

Klaviyo Social Marketing GA: Instagram DMs Meet Your CRM

Klaviyo Social Marketing is GA: Instagram DMs, comments, and mentions now write to the same customer profile as email, SMS, and order history.

July 10, 2026 · 10 minRead
CRM & Automation

Build a Competitor-Monitoring Agent With MCP Servers

Build a competitor-monitoring agent with MCP tools, open-source change detection and hash-gated noise control — the hard part is signal, not connectivity.

July 8, 2026 · 14 minRead
CRM & Automation

Build Live Client Dashboards with Claude Code Artifacts

Claude Code Artifacts now pull live data through a viewer's own MCP connectors per view. But a connector-backed dashboard can't be a public link on any plan.

July 24, 2026 · 11 minRead
CRM & Automation

SMS Marketing Statistics 2026: 110+ Open and CTR Data

SMS marketing statistics for 2026: 110+ data points on delivery, open and click-through rates, opt-out behavior, and revenue-per-send benchmarks.

April 22, 2026 · 15 minRead