Overview
- Meta announced Muse Voice Transcribe on Tuesday as its first real‑time audio‑perception model that transcribes speech as it happens.
- The single model performs streaming speech recognition, speaker diarization (assigning words to speakers) and endpointing (detecting when someone stops speaking) without separate post‑processing.
- Muse processes audio in 80 millisecond chunks and uses an adaptive delay trained by reinforcement learning so the model waits for harder words and emits easy words quickly.
- Developers can access Muse through the Meta Model API at $3 per 1,000 audio‑minutes and the feature already powers dictation in Meta AI for Mac and Muse Code.
- Meta cites a 3.1% AA‑WER lead on an English streaming benchmark but the company will not release model weights and independent, cross‑language and long‑form diarization tests are still limited while competitors iterate rapidly.