AI DevelopmentNew Release7 min readPublished September 15, 2026

Two models · one price · one design decision

Gemini 3.8 Live: Should a Voice Agent Think While Talking?

Google split its live voice model in two on September 15: one answers at once, one reasons while it speaks. Which to pick, and where each is available.

DA
Digital Applied Team
Research and practical guidance
Editorial dateSeptember 15, 2026
SourcesGoogle blog · Gemini API docs

On September 15, 2026 Google released two live voice models instead of one. Gemini 3.8 Live answers without a reasoning pause and is built, in Google's words, for scale and cost efficiency. Gemini 3.8 Live Extended Thinking reasons and speaks at the same time, saying things like "Let me check that…" while it works. The split turns a model choice into a product choice: should your voice agent sound instant, or should it be allowed to think out loud?

This post is for the builder choosing a live voice model this month. It reads Google's announcement and the Gemini API pricing page as read on September 16, sets out the decision, and prints the benchmark figures with their sources. We have run no latency tests of our own and quote none.

Key takeaways
  1. 01
    Same price, different behaviour.Google's pricing page lists one Live rate covering both models. The choice is about how the agent behaves, not what it costs per token.
  2. 02
    Extended Thinking narrates while it reasons.Google says it uses early verbal cues and live progress narration during multi-step background tasks, so the caller hears activity rather than silence.
  3. 03
    3.8 Live is the default for most agents.Google's model list calls it the default option for low-latency voice experiences without reasoning delays, and it runs tools in the background while talking.
  4. 04
    Availability differs by surface.Both are in the Gemini API and AI Studio today. Enterprise access is private preview. Consumer surfaces differ between the two.

01The releaseWhat Google released

Both models are audio-to-audio, meaning speech goes in and speech comes out without a separate transcription step, over Google's Live API. Google describes 3.8 Live as combining conversational intelligence with fluid dialogue and visual grounding, and 3.8 Live Extended Thinking as built for high-complexity tasks with multi-step reasoning. The model IDs in the API are gemini-3.8-live and gemini-3.8-live-extended-thinking, and both are listed as stable, not preview. The earlier 3.1 Flash Live preview is now marked legacy with a recommendation to move.

Three capabilities Google attributes to 3.8 Live matter for a business agent. It processes visual input in near real time. It "automatically detects and transitions between 97 supported languages mid-conversation". And it executes tools and API calls in the background while continuing to talk, so it can acknowledge a request and keep the conversation going while the task finishes. Extended Thinking adds reasoning on top of that, and Google's description of how it fills the gap is the reason this post exists.

Google's description of Extended Thinking

It "reasons and speaks simultaneously", using early verbal cues such as "Let me check that…" to acknowledge a prompt naturally, and "live progress narration to walk users through multi-step background tasks as they progress". We covered an earlier think-while-talking design in the Grok Voice release. Google's answer to the same problem is two separate models.

02The trade-offThe decision the split creates

A voice agent has one budget that a chat agent does not: the caller's patience with silence. Every second a model spends reasoning before it speaks is a second the caller hears nothing. The industry's first answer was to make models fast enough that the pause disappeared. Google's second answer is to let the model talk through the pause. That is a different product, and it suits different calls.

An instant model suits short, high-volume interactions where the answer is usually retrieval: opening hours, order status, a booking change, a routed transfer. A narrating model suits calls where the caller expects the agent to go and do something that takes several steps: check three systems, apply a rule, come back with a decision. In that setting a spoken "Let me check that" is what a competent human would say, and silence is what a broken one would produce.

The cost of narration is that it is speech the caller has to sit through. If the task takes two seconds, a preamble is padding. If it takes twenty, it is the difference between a caller staying on the line and hanging up. Google's post gives no figures for how long the background tasks take, so the only way to know which side of that line your calls fall on is to measure your own. Our voice agent latency measures reference defines the numbers to collect; this post does not repeat them.

Model A
Gemini 3.8 Live
Answers at once, tools run in the background

Google's default for low-latency voice agents without reasoning delays. Handles visual input, 97 languages, and background tool calls while the conversation continues. Built for scale and cost.

Instant
Model B
Gemini 3.8 Live Extended Thinking
Reasons and speaks at the same time

Google's recommendation when higher background reasoning is needed during a live interaction. Uses verbal cues and progress narration while multi-step tasks run.

Narrating

03RoutingWhich model, when

The routing below is ours. It follows from Google's own descriptions of the two models and from the shape of the call, not from any measurement of the models. Because Google's pricing page lists a single Live rate for both, price does not enter the decision at the per-token level.

Most calls end with one lookup or one routed transfer, and volume is high
3.8 Live. The caller wants an answer, not a narrative, and Google positions this model for scale and cost.
3.8 Live
The agent must consult several systems or apply a policy before it can answer, and callers currently hear silence or hold music
Extended Thinking. Spoken progress replaces the hold, and the deeper reasoning is the point of the call.
Extended Thinking
Calls are mixed and you cannot predict which kind a caller will bring
Start on 3.8 Live and route to Extended Thinking only for the intents your logs show taking longest. Google's post describes both as Live API models, so the surface is the same.
Route

One earlier decision still applies whichever model you pick. OpenAI's live voice API took a different route to the same problem, which we described in our business guide to GPT Live 1, and the choice between vendors is not settled by either vendor's benchmarks. The Grok Voice think-while-talking release is the third design in the same space. Three vendors, three different answers to the silence problem, and no cross-vendor test any of them has published.

04AccessWhere each model is available

Google lists availability by audience. The table reproduces it, with one caution: "private preview" is an invitation, not availability, and "coming soon" carries no date.

Source: Google's announcement of September 15, 2026, as published. Rollouts described as "starting today".
SurfaceGemini 3.8 Live3.8 Live Extended Thinking
Gemini API and Google AI StudioAvailableAvailable
Gemini EnterprisePrivate previewPrivate preview
Gemini Enterprise for Customer ExperienceComing soonComing soon
Google Workspace, business customersNot listedComing soon
Consumer surfacesSearch LiveGemini Live; Workspace Docs for AI Pro and Ultra subscribers; Gmail and Keep for all AI subscribers

On price, the Gemini API pricing page as read on September 16 lists a single paid tier for "Gemini 3.8 Live, Gemini 3.8 Live Extended Thinking, and Gemini 3.1 Flash Live Preview" together. Per million tokens, input is $0.75 for text, $3.00 for audio and $1.00 for image or video; output, including thinking tokens, is $4.50 for text and $12.00 for audio. Google also states the audio rates per minute: $0.005 in and $0.018 out. The page lists a free tier for both models and says free-tier content is used to improve Google's products while paid-tier content is not.

05EvidenceThe benchmark figures, with sources

Google cites five results. Each is either a leaderboard position on a third-party site, dated the day of the announcement and free to move, or a score Google reports for its own model. None is a test of a business call, and none compares the two Gemini models against each other on the same task.

Artificial Analysis Speech to Speech Quality IndexExtended Thinking · third-party index · rank stated by Google
82.6, first
τ-Voice agentic task completionExtended Thinking · Google-reported score
68.6%
τ-Voice-banking (Sierra)Extended Thinking · Google-reported score
35.1%
Big Bench AudioExtended Thinking · Google-reported score
97.7%
Speech Agent Arena3.8 Live · user-preference leaderboard · rank stated by Google
second place

Google also says the models sit on the accuracy-versus-quality frontier of ServiceNow's EVA-Bench for complex workflows, with a note that the run used the Live API on Google's enterprise agent platform. Read all of these as a vendor's selection of favourable results. That is normal, and it is why the routing in section three rests on call shape rather than on these numbers. If you want a figure you can act on, record the silence your current agent produces on your ten most common intents and test both models against it. Our AI transformation practice runs that comparison as the first week of a voice-agent build.

06Next stepSilence is now a design choice, not a limitation

Put it into practice

Sort your call intents by how long the work behind them takes

Google has made the pause optional and given both options the same price. Pull the intents your agent handles, measure how long the work behind each one takes, and put the short ones on 3.8 Live and the long ones on Extended Thinking. Treat the benchmark figures as the vendor's evidence, treat private preview as not yet available, and let your own recordings decide the rest.

Digital Applied

Build a voice agent your callers will stay on the line for.

We prototype on the live model that fits each intent, measure the silence, and ship the version your call recordings justify.

Intent analysisLive model trialsProduction delivery
Your next project

Start with the calls that go quiet

  • List the ten most common intents
  • Measure the work behind each
  • Trial both models on the long ones
Questions and answers

Applying this post

As of September 16, 2026 the Gemini API pricing page lists one paid rate covering Gemini 3.8 Live, Gemini 3.8 Live Extended Thinking and the 3.1 Flash Live preview: $0.75 per million text input tokens, $3.00 audio input, $4.50 text output and $12.00 audio output, with thinking tokens billed as output.
Digital Applied newsletter

Deep dives on AI, marketing and development.

Practical guides and fresh insights by email. No recycled takes.

Related dispatches

Continue reading

AI Development

Cloudflare Blocks AI Agents on Ad Pages: Which Bots Are Hit

From September 15, 2026 new ad-supported Cloudflare domains block AI agents on ad pages and refuse AI training by default. A 20-bot census of who is affected.

September 15, 2026 · 10 minRead
AI Development

TabPFN 3.5 Beats Boosted Trees: When to Use It on Your Data

Prior Labs' TabPFN-3.5 report claims first place on seven tabular benchmarks. What a tabular foundation model is, when to use it, and what the licence allows.

September 15, 2026 · 8 minRead
AI Development

Coding Agents Grew Anthropic's CI 25x: How the Fix Worked

Anthropic says coding agents raised its CI jobs 25x in six months. Three patches bought 70 days, 29 days and under a day. What the redesign teaches.

September 15, 2026 · 7 minRead
AI Development

AI Labs Say They Will Slow Down: What Was Actually Promised

Dario Amodei's pacing essay commits Anthropic to embedded outside evaluators. What is promised, what is only proposed, and what a model buyer should watch.

September 15, 2026 · 8 minRead
AI Development

Preview, Beta, GA: What Vendors Said vs What Coverage Said

Thirty-six AI vendor announcements from 17-22 August 2026, each scored on the vendor's own status word against the word its coverage used, where located.

August 22, 2026 · 27 minRead
AI Development

AI Agent Memory 2026: Vector, Graph, Episodic Update

AI agent memory architectures compared after Code with Claude London — Anthropic Dreaming, Memory Tool, Google Memory Bank, vector, graph, episodic patterns.

May 24, 2026 · 16 minRead