Gemini 3.8 Live voice agents with Gemini sparkle icon

Google Ships Gemini 3.8 Live and Extended Thinking for Voice Agents

Focus keyphrase Gemini 3.8 Live
SEO title Google Ships Gemini 3.8 Live for Voice Agents
Slug gemini-3-8-live-extended-thinking-voice-agents
Meta description Google launched Gemini 3.8 Live and Extended Thinking (Sep 15, 2026): speech-to-speech models with async tools, API pricing, and Workspace rollout.
Category AI News
Tags Gemini 3.8 Live, Gemini Live API, Google DeepMind, voice agents, speech-to-speech, Extended Thinking, Google AI Studio

On September 15, 2026, Google introduced Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, two speech-to-speech models aimed at real-time voice agents that can keep talking while they call tools and reason in the background.

By the morning of September 16, the launch was on the Hacker News front page (“Gemini 3.8 Live and 3.8 Live Extended Thinking”) with hundreds of points. That spike is why this is today’s news piece: a same-day developer and product ship from Google DeepMind, not a recycled model teaser.

The useful framing is builder-facing. These are Live API models with concrete prices, model IDs, and migration notes — plus a consumer/Workspace rollout — not a research-only demo.

Key points

  • What shipped: Gemini 3.8 Live (scale / cost) and Gemini 3.8 Live Extended Thinking (deeper multi-step reasoning while speaking).
  • Primary sources: Google’s product post (Sep 15) and the developer post Build real-time voice applications with Gemini 3.8 Live… (Sep 15).
  • Developer access: Gemini API and Google AI Studio via the Live API; model code gemini-3.8-live (Extended Thinking: gemini-3.8-live-extended-thinking).
  • List price (audio): $0.005/min input and $0.018/min output (Google notes the token-equivalent estimate as $3/1M input and $12/1M output).
  • Headline capabilities: asynchronous (non-blocking) tool/API calls while streaming audio; visual grounding; 97+ languages with mid-conversation switching; alphanumeric precision for codes and IDs.
  • Benchmarks Google cites: Extended Thinking #1 on Artificial Analysis Speech-to-Speech Quality Index (82.6); 68.6% on τ-Voice; 35.1% on Sierra τ-Voice-banking; 97.7% on Big Bench Audio. Treat as vendor-reported.
  • Where it lands for users: Search Live / Gemini Live; Docs Live for Google AI Pro/Ultra; Gmail and Keep Live for Google AI subscribers; Gemini Enterprise private preview.
  • Transparency: SynthID watermark on generated audio; model card linked from Google’s post.

What Gemini 3.8 Live shipped

Google positions the pair as near–real-time dialogue models for voice agents:

Model Role (Google’s framing) Best for
Gemini 3.8 Live Conversational intelligence + fluid dialogue + visual grounding; built for scale and cost Default low-latency voice agents
Gemini 3.8 Live Extended Thinking Higher intelligence + multi-step reasoning while speaking Complex workflows that need background thinking + live narration

Both run through the Gemini Live API. Partner platforms Google names for media streaming include Agora, Fishjam, LiveKit, LangChain, Pipecat, Vercel, and Vision Agents.

Extended Thinking in practice

For harder tasks, Extended Thinking is meant to reason and speak at the same time: early verbal acknowledgements (“Let me check that…”), progress narration while tools run, and configurable background thinking so the conversation does not stall on every tool call.

Google’s demos (on the product page) lean on visual + voice workflows: onboarding with live visuals, chess with camera context, sketch-to-React with spoken feedback, and multi-step bookings with async function calls.

API / migration notes (builders)

From Google’s Gemini 3.8 Live model docs:

  • Stable model string: gemini-3.8-live
  • Inputs: text, images, audio, video · Outputs: text and audio
  • Context: 131,072 input / 65,536 output token limits
  • Migrating from gemini-3.1-flash-live-preview: drop thinking_level / thinking_config; async function calling (behavior: NON_BLOCKING) is the default; proactive audio is always on; affective dialogue config is removed; audio is the supported response modality

We have not re-run Google’s latency or quality benches ourselves. Prefer the Live API guide and model cards for production limits.

What it means for builders

1. Voice agents without a cascade stack — Google pitches native speech-to-speech plus async tools as a cleaner alternative to separate STT → LLM → TTS pipelines. If you already run LiveKit / Pipecat / similar, the partner list is the practical on-ramp.

2. Price is usable for product experiments — $0.005 / $0.018 per minute is concrete enough to model cost for support, coaching, and booking agents before you commit to enterprise preview.

3. Pick the right SKU — default to 3.8 Live for chatty, low-latency UX; reserve Extended Thinking for workflows that need multi-step tool use and spoken progress without freezing the turn.

4. Consumer surface ≠ API surface — Workspace/Search Live access is plan-gated; API access is what matters if you are shipping your own agent.

What to watch

  • How Extended Thinking quality holds up outside Google’s cited benches once third-party evals land.
  • Enterprise Customer Experience and Workspace business availability (still “coming soon” / private preview in the launch post).
  • Whether async tool calling plus visual grounding reduces the classic “voice agent stall” failure mode in real apps.
  • Parallel market context (not this piece’s subject): OpenAI’s GPT-Live-1 / Agents API wave earlier in September, and other speech-agent launches — compare on latency, tool reliability, and price, not press copy.

Sources

9to6AI — clear, honest coverage for people who ship with AI tools.

Leave a Comment

Your email address will not be published. Required fields are marked *