Particle.news

Google Launches Gemini 3.8 Live and Extended Thinking for Conversational Voice Agents

The release signals that voice AI can now carry out background tasks during speech, changing how developers build real-time agents.

Overview

  • Google announced on Sept. 15 that Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking are rolling out to developers via the Gemini API and Google AI Studio and to enterprises in private preview while consumer integrations appear in Search Live and Workspace.
  • The models let a voice agent keep talking as it executes API or tool calls, reason mid-conversation, accept visual inputs, and auto-detect 97 languages to support richer, uninterrupted voice workflows.
  • Extended Thinking performs multi-step in-model reasoning and narrates progress aloud, a design choice that contrasts with competitors that split low-latency voice and backend reasoning and that creates different tradeoffs for latency, control, and visibility.
  • Google published base Live API rates of $0.005 per minute for audio input and $0.018 per minute for audio output and said Extended Thinking will add reasoning token fees and extra input charges for things like live video and documents.
  • Google cites top scores on several vendor-reported benchmarks and embeds SynthID watermarks in generated audio, but buyers should validate performance and cost for real workloads and expect design work around cost, interruption handling, and streaming infrastructure.