Particle.news

Google Launches Gemini 3.8 Flash TTS With Voice Cloning and Studio Controls

Independent early benchmarks place the model atop pronunciation tests, raising the bar for commercial text‑to‑speech rivals.

Overview

  • Gemini 3.8 Flash TTS and a lighter Flash‑Lite variant were released Wednesday, Sept. 23, through the Gemini API and Google AI Studio and have started rolling into Gemini Notebook and Google Vids.
  • The models can generate new voices from text descriptions, reproduce an authorized voice from roughly a 30‑second sample, and handle two‑speaker dialogue with fine control over accent, pacing, and emotion.
  • Google says the system supports more than 100 languages and dialects and includes over 2,000 production‑ready voices for developers to deploy.
  • Third‑party reports published after the launch show Gemini 3.8 Flash TTS scored 89.5% on the Artificial Analysis Pronunciation Robustness Benchmark and placed highly on Hume AI voice‑design measures, but those results come from vendor‑reported or independent tests that buyers should validate against their own workloads and consent controls.
  • The release tightens competition in the TTS market dominated by players such as ElevenLabs and OpenAI and could speed adoption of synthetic voice features in customer service, agents, and media while raising questions about consent, legal safeguards, integration cost, and real‑world latency.