Overview
- Google has publicly released Gemini 3.5 Transcribe and begun rolling it into products such as Gboard on Pixel 11 and the Mac Gemini app while opening a developer preview through the Gemini API.
- The model moves beyond literal transcription by detecting self‑corrections, removing fillers like “um” and “uh,” and automatically formatting spoken language into cleaner text.
- Gemini 3.5 Transcribe supports over 85 languages, automatic language detection, custom vocabularies, and for pre‑recorded audio it offers speaker identification and word‑level timestamps with reported support for up to three speakers.
- Google reports material performance gains versus its prior Chirp 3 engine, citing about 70% faster end‑to‑end speed and a lower real‑time error rate (about 5.5% versus Chirp 3’s 7.32%), and file transcriptions are currently limited to one hour.
- The change raises practical trade‑offs for users and developers because the model’s automatic editing can alter exact wording, which may be undesirable for legal, journalistic, or other contexts that require verbatim records and could influence how broadly voice replaces typing in everyday apps.