Particle.news

Meta Releases Muse Voice Transcribe, a Real-Time Multilingual Speech Model

The product gives developers low-cost streaming transcription with built-in speaker separation and an adaptive delay that balances speed and accuracy.

Overview

  • Meta made Muse Voice Transcribe public for developers through the Meta Model API and has integrated it into Meta AI for Mac and Muse Code with a pay-as-you-go price of $3 per 1,000 audio-minutes.
  • The model was trained on more than 70 languages with 25 validated, supports recordings longer than an hour, handles native code-switching, and can separate more than 20 speakers in real time with endpointing.
  • Muse Voice Transcribe processes audio in 80-millisecond chunks as an autoregressive Muse Spark model and uses a reinforcement-learned adaptive delay to trade off per-word latency and accuracy.
  • On the Artificial Analysis AA-WER Streaming English benchmark the model posted a 3.1% word error rate, a lead over several rivals, but that result covers only English and independent multilingual and diarization tests are still needed.
  • Meta confirmed it will not release the model weights, a closed approach that may speed product use inside Meta’s apps but could limit outside research, third-party audits, and some deployment choices.