| Focus keyphrase | Gemini 3.8 Live |
|---|---|
| SEO title | Google Ships Gemini 3.8 Live for Voice Agents |
| Slug | gemini-3-8-live-extended-thinking-voice-agents |
| Meta description | Google launched Gemini 3.8 Live and Extended Thinking (Sep 15, 2026): speech-to-speech models with async tools, API pricing, and Workspace rollout. |
| Category | AI News |
| Tags | Gemini 3.8 Live, Gemini Live API, Google DeepMind, voice agents, speech-to-speech, Extended Thinking, Google AI Studio |
On September 15, 2026, Google introduced Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, two speech-to-speech models aimed at real-time voice agents that can keep talking while they call tools and reason in the background.
By the morning of September 16, the launch was on the Hacker News front page (“Gemini 3.8 Live and 3.8 Live Extended Thinking”) with hundreds of points. That spike is why this is today’s news piece: a same-day developer and product ship from Google DeepMind, not a recycled model teaser.
The useful framing is builder-facing. These are Live API models with concrete prices, model IDs, and migration notes — plus a consumer/Workspace rollout — not a research-only demo.
Key points
- What shipped: Gemini 3.8 Live (scale / cost) and Gemini 3.8 Live Extended Thinking (deeper multi-step reasoning while speaking).
- Primary sources: Google’s product post (Sep 15) and the developer post Build real-time voice applications with Gemini 3.8 Live… (Sep 15).
- Developer access: Gemini API and Google AI Studio via the Live API; model code
gemini-3.8-live(Extended Thinking:gemini-3.8-live-extended-thinking). - List price (audio): $0.005/min input and $0.018/min output (Google notes the token-equivalent estimate as $3/1M input and $12/1M output).
- Headline capabilities: asynchronous (non-blocking) tool/API calls while streaming audio; visual grounding; 97+ languages with mid-conversation switching; alphanumeric precision for codes and IDs.
- Benchmarks Google cites: Extended Thinking #1 on Artificial Analysis Speech-to-Speech Quality Index (82.6); 68.6% on τ-Voice; 35.1% on Sierra τ-Voice-banking; 97.7% on Big Bench Audio. Treat as vendor-reported.
- Where it lands for users: Search Live / Gemini Live; Docs Live for Google AI Pro/Ultra; Gmail and Keep Live for Google AI subscribers; Gemini Enterprise private preview.
- Transparency: SynthID watermark on generated audio; model card linked from Google’s post.
What Gemini 3.8 Live shipped
Google positions the pair as near–real-time dialogue models for voice agents:
| Model | Role (Google’s framing) | Best for |
|---|---|---|
| Gemini 3.8 Live | Conversational intelligence + fluid dialogue + visual grounding; built for scale and cost | Default low-latency voice agents |
| Gemini 3.8 Live Extended Thinking | Higher intelligence + multi-step reasoning while speaking | Complex workflows that need background thinking + live narration |
Both run through the Gemini Live API. Partner platforms Google names for media streaming include Agora, Fishjam, LiveKit, LangChain, Pipecat, Vercel, and Vision Agents.
Extended Thinking in practice
For harder tasks, Extended Thinking is meant to reason and speak at the same time: early verbal acknowledgements (“Let me check that…”), progress narration while tools run, and configurable background thinking so the conversation does not stall on every tool call.
Google’s demos (on the product page) lean on visual + voice workflows: onboarding with live visuals, chess with camera context, sketch-to-React with spoken feedback, and multi-step bookings with async function calls.
API / migration notes (builders)
From Google’s Gemini 3.8 Live model docs:
- Stable model string:
gemini-3.8-live - Inputs: text, images, audio, video · Outputs: text and audio
- Context: 131,072 input / 65,536 output token limits
- Migrating from
gemini-3.1-flash-live-preview: dropthinking_level/thinking_config; async function calling (behavior: NON_BLOCKING) is the default; proactive audio is always on; affective dialogue config is removed; audio is the supported response modality
We have not re-run Google’s latency or quality benches ourselves. Prefer the Live API guide and model cards for production limits.
What it means for builders
1. Voice agents without a cascade stack — Google pitches native speech-to-speech plus async tools as a cleaner alternative to separate STT → LLM → TTS pipelines. If you already run LiveKit / Pipecat / similar, the partner list is the practical on-ramp.
2. Price is usable for product experiments — $0.005 / $0.018 per minute is concrete enough to model cost for support, coaching, and booking agents before you commit to enterprise preview.
3. Pick the right SKU — default to 3.8 Live for chatty, low-latency UX; reserve Extended Thinking for workflows that need multi-step tool use and spoken progress without freezing the turn.
4. Consumer surface ≠ API surface — Workspace/Search Live access is plan-gated; API access is what matters if you are shipping your own agent.
What to watch
- How Extended Thinking quality holds up outside Google’s cited benches once third-party evals land.
- Enterprise Customer Experience and Workspace business availability (still “coming soon” / private preview in the launch post).
- Whether async tool calling plus visual grounding reduces the classic “voice agent stall” failure mode in real apps.
- Parallel market context (not this piece’s subject): OpenAI’s GPT-Live-1 / Agents API wave earlier in September, and other speech-agent launches — compare on latency, tool reliability, and price, not press copy.
Sources
- Introducing Gemini 3.8 Live and 3.8 Live Extended Thinking (Google, Sep 15, 2026)
- Build real-time voice applications with Gemini 3.8 Live and 3.5 Transcribe (Google, Sep 15, 2026)
- Gemini 3.8 Live model docs
- Gemini 3.8 Live Extended Thinking model docs
- Hacker News discussion of the Google blog post (front page morning of Sep 16, 2026)
9to6AI — clear, honest coverage for people who ship with AI tools.



