Particle.news

Google Expands Gemini With Flash TTS and Live Avatar

The updates move studio-grade voice creation, short-sample cloning with consent, embedded provenance, live animated avatars into developer and enterprise tools.

Overview

  • Gemini 3.8 Flash TTS and Flash-Lite TTS rolled out to developers through the Gemini API and Google AI Studio on Sept. 23, with enterprise API access coming soon and consumer access in Gemini Notebook and Google Vids.
  • Developers can prompt new voices from text or replicate a specific voice from a roughly 10–30 second sample after a required consent recording, which yields a reusable voice_… ID or a short-lived encrypted voice key with project storage limits.
  • Flash TTS offers line-by-line performance control, native two-speaker scene staging, and stability over multi-hour audio for use cases like audiobooks, games, and podcast production while Flash-Lite targets high-volume, lower-cost dubbing and voice agents.
  • Google built provenance and safety into the flow: voice replication requires consent verification and all generated audio and video carry SynthID watermarks and C2PA credentials to help trace and label synthetic content.
  • Gemini 3.8 Live with Live Avatar became available in Gemini Enterprise on Sept. 24, pairing low-latency video with speech for lip-synced, multilingual avatars, and independent benchmarks place Flash TTS near the top while analysts urge testing for latency, cost, and integration trade-offs.