Particle.news

Meta Launches Muse Voice Transcribe for Real-Time Speech

The model uses an RL-trained adaptive delay to deliver low-latency streaming transcription with speaker labeling and long-form multilingual support.

Overview

  • Meta launched Muse Voice Transcribe on Tuesday, September 1, 2026, and made it available through the Meta Model API, Meta AI for Mac, and Muse Code.
  • The model transcribes streaming audio in 80-millisecond chunks, applies an RL-learned “adaptive delay” to trade off latency and accuracy, and performs speaker diarization for more than 20 speakers.
  • Meta says Muse was trained on over 70 languages with 25 validated at launch, supports native code-switching and audio longer than an hour, and offers language and keyword biasing to improve recognition.
  • On Artificial Analysis’s AA-WER Streaming benchmark Muse scored a 3.1% word error rate in English and posted a 17.5% error on speaker-recognition tests, placing it ahead of several competitors on those measures.
  • Developers can access Muse via Meta’s API at $3 per 1,000 audio-minutes, but Meta will not release the model weights, a choice that speeds product integration while limiting local research and fine-tuning.