Gemini 3.8 Live: Google’s Voice AI That Thinks While It Talks

Key Takeaways

  • Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking are Google’s newest voice AI models, released September 15, 2026, built to reason and speak at the same time instead of pausing to think
  • Extended Thinking ranks #1 overall on Artificial Analysis’ Speech to Speech Quality Index with a score of 82.6, and leads agentic task completion benchmarks too
  • Both models automatically detect and switch between 97 spoken languages mid-conversation, without you changing any setting
  • Gemini 3.8 Live is live now in Search Live for everyone; Extended Thinking is rolling out in the Gemini app, Gmail, Keep, and Google AI Pro/Ultra Workspace plans
  • Every audio output carries an invisible SynthID watermark so AI-generated speech stays identifiable

Gemini 3.8 Live just changed what a voice conversation with AI can actually feel like. Instead of going silent while it works through a hard question, Google’s newest voice model talks and thinks simultaneously, narrating its progress the way a helpful coworker would rather than leaving you staring at a spinner. Google released it alongside a second version, Gemini 3.8 Live Extended Thinking, on September 15, 2026. This guide covers how the two models differ, what the real benchmarks show, and exactly how to start using Gemini 3.8 Live today.


The Difference Between the Two Models

Google shipped two versions rather than one, and the split matters more than it might first appear. Gemini 3.8 Live answers immediately, without a reasoning pause, and Google built it specifically for scale and cost efficiency — the right pick for high-volume, straightforward conversations. Gemini 3.8 Live Extended Thinking, on the other hand, reasons and speaks at the same time, using early verbal cues like “Let me check that…” to acknowledge a request naturally while it works through something harder in the background.

Both models process visual input in near real-time, so you can point your camera at something and get a spoken answer grounded in what the model actually sees. Both also handle 97 spoken languages, switching automatically mid-conversation without you touching a setting.


How Gemini 3.8 Live Actually Performs

Google backed this launch with real third-party benchmark numbers rather than just marketing claims:

BenchmarkScore
Artificial Analysis Speech to Speech Quality Index82.6 (#1 overall)
τ-Voice agentic task completion68.6%
Sierra’s τ-Voice-banking benchmark35.1%
Big Bench Audio reasoning97.7%

Standard Gemini 3.8 Live took second place in the Speech Agent Arena, a separate ranking that measures direct human preference between voice models, while still remaining highly cost-effective for large-scale use.


Gemini 3.8 Live vs. GPT-Live

Google’s launch positions Gemini 3.8 Live squarely against OpenAI’s GPT-Live family, and the two take a genuinely similar architectural approach — both are full-duplex speech-to-speech models rather than the older cascaded pipeline of separate transcription, reasoning, and speech models chained together.

FeatureGemini 3.8 Live Extended ThinkingGPT-Live
Reasoning while speakingYes, with narrated progress updatesYes, with reasoning-level settings
Language switching97 languages, automatic mid-conversationSupported, fewer languages confirmed
Visual groundingNear real-time camera inputNot a launch headline feature
Video/screen sharing in voice modeYes, via visual contextNot yet at launch

Neither company has published a benchmark suite testing directly against the other’s model, so treat any specific head-to-head claim, including comparisons drawn from separate self-reported test runs, with appropriate caution.


Where You Can Use It Right Now

  • Everyone: Gemini 3.8 Live is live today in Search Live, and Extended Thinking is live in the Gemini app.
  • Google Workspace: Extended Thinking works in Docs Live and Keep Live for Google AI Pro and Ultra subscribers, and in Gmail for all Google AI subscribers.
  • Developers: Both models are available now through the Gemini API and Google AI Studio.
  • Enterprises: Both are in private preview inside Gemini Enterprise, with Gemini Enterprise for Customer Experience support coming soon.

How to Get Started, Step by Step

Step 1: Try It in the Gemini App

Open the Gemini app and start a voice conversation. Extended Thinking already powers this experience, so you don’t need to select a specific model manually.

Step 2: Ask for Something That Needs Real Work

Instead of a simple question, try a multi-step request — planning your day, drafting an email while you talk through the details, or walking through a troubleshooting problem. This is where the “thinking while talking” behavior actually shows up.

Step 3: Point Your Camera at Something

Use visual grounding by showing Gemini something through your camera mid-conversation — a whiteboard sketch, a broken appliance, or a document — and ask a question about what it sees.

Step 4: Build With the API

Developers can access both models directly through the Gemini API or Google AI Studio’s Live interface, referencing them as gemini-3.8-live or gemini-3.8-live-extended-thinking.

Step 5: Use an Infrastructure Partner for Production Apps

For building a real voice product rather than testing casually, Google points developers toward infrastructure partners like LiveKit, Pipecat, Agora, and Vercel, which handle the real-time media streaming plumbing so you can focus on the actual user experience.


Developer Pricing

As of this week, the Gemini API pricing page lists a single Live rate covering both Gemini 3.8 Live and Extended Thinking: $0.75 per million text input tokens, $3.00 per million audio input tokens, $4.50 per million text output tokens, and $12.00 per million audio output tokens, with thinking tokens billed as output. Some reports also cite a simpler per-minute rate of roughly $0.005 for audio input and $0.018 for audio output. So, the price is identical either way you use it — the real decision is which model behavior fits your product, not what it costs.


Watermarking and Transparency

Every piece of audio Gemini 3.8 Live generates carries an invisible SynthID watermark woven directly into the output. This doesn’t affect what you hear, and it exists specifically so AI-generated speech remains detectable later, which matters as voice content becomes harder to distinguish from a real recording by ear alone.


Frequently Asked Questions

Do I need to choose between the two models myself?

Usually not as an everyday user — the Gemini app and other consumer surfaces already route you to the appropriate model. Developers calling the API do choose explicitly between the two model names.

Does Gemini 3.8 Live replace the older Gemini Live model?

It’s the newest generation of Google’s live dialogue models, built for the same purpose but with meaningfully stronger reasoning and language-switching than the prior 3.1 Flash Live generation.

Can it handle a task that takes several minutes to complete?

Yes — Extended Thinking specifically continues working on background function calls and multi-step tasks while narrating progress, rather than requiring you to wait silently or restart the conversation.

Is this available outside English?

Yes, extensively. The model supports 97 spoken languages and switches between them automatically mid-conversation without any manual setting change.


Final Thoughts

Gemini 3.8 Live marks a real architectural shift in how voice AI handles hard problems — narrating its own progress rather than forcing an awkward silence, and grounding answers in what it can actually see through your camera. Whether it’s the better choice over GPT-Live likely comes down to which ecosystem you already build in, since both companies are still testing on their own benchmark suites rather than a shared, neutral comparison. Either way, the direction is clear: voice interfaces are quickly becoming capable enough to handle real, multi-step work, not just quick questions.

Read Next:

ChatGPT vs Gemini vs Claude: Which One Should You Use in 2026 — for the bigger picture beyond just voice, comparing these assistants across the board.

Read Google’s full announcement for complete technical details and demo videos.

Leave a Comment

Your email address will not be published. Required fields are marked *

Scroll to Top