Meta Launches Muse Voice Transcribe, Undercutting Rival Speech AI on Price
Meta Superintelligence Labs shipped its first real-time audio perception model on September 1, 2026. Muse Voice Transcribe folds streaming transcription, speaker diarization and end-of-speech detection -- normally three separate systems -- into one model, and Meta priced access at a fraction of what specialist voice-AI vendors charge for the same building blocks.
What Meta announced
Meta Superintelligence Labs (MSL) -- the division Mark Zuckerberg formed in June 2025 and put former Scale AI chief executive Alexandr Wang in charge of -- published Muse Voice Transcribe on its research blog on September 1, 2026, describing it as the lab's "first real-time audio perception model." The pitch is architectural, not just about raw accuracy: instead of running separate systems for speech-to-text, telling speakers apart, and detecting when someone has stopped talking, Muse Voice Transcribe does all three inside one streaming model, decision by decision, as audio arrives.
The model is built on Meta's Muse Spark family and processes audio in 80-millisecond chunks -- a rate of 12.5 chunks per second. After each chunk, the model makes a binary call: keep listening, or start emitting text. Special tokens carry the rest of the work inline with the transcript itself -- a <|start_of_turn|> token marks a speaker change, <|speaker_A|> through <|speaker_Z|> style tokens identify who is talking, and dedicated tokens mark where speech starts and ends. That is what lets one pass through the audio replace three separate pipelines, according to the technical writeup from MarkTechPost, which reviewed Meta's release in detail on launch day.
The accuracy numbers, and how it hits them
Meta says the model reaches a 3.1% word-error rate on final transcription, produced about 0.16 seconds after a speaker stops -- with an even faster first-pass partial transcript at 3.6% WER in 0.13 seconds, per MarkTechPost's benchmark summary. On speaker diarization, Muse Voice Transcribe posted a 17.5% average diarization error rate across the AMI-IHM, AMI-SDM and VoxConverse benchmark sets, which MarkTechPost reports as meaningfully better than rival systems scoring between 21.1% and 28.6% on the same benchmarks.
The mechanism behind those numbers is what Meta calls "adaptive delay": a reinforcement-learning policy that decides, word by word, how long to wait before committing to text. Easy, unambiguous words get emitted almost immediately; harder or ambiguous ones get a longer pause while the model gathers more audio context. Meta's framing is that this policy is trained to sit on the Pareto frontier between speed and accuracy, rather than picking one fixed latency for every kind of speech.
Sources: Meta Superintelligence Labs research blog, September 1, 2026; MarkTechPost, September 1, 2026; 9to5Mac, September 1, 2026.
Language coverage and where it ships first
Meta says Muse Voice Transcribe was trained on more than 70 languages, with 25 -- including Chinese, French, Hindi, Japanese, Spanish and Vietnamese -- extensively validated at launch. It handles code-switching within a single sentence as well as between sentences, and can process a continuous audio stream longer than an hour without a separate post-processing step. Optional language, keyword and context biasing lets a developer nudge the model toward expected vocabulary for a given call or session.
The model is API-only -- Meta has not released open weights -- available through the Meta Model API as muse-voice-transcribe-1.0. It is already live in two of Meta's own products: system-wide voice dictation in Meta AI for Mac (triggered with the Fn key) and inside Muse Code, the terminal coding agent MSL shipped on August 5, 2026 as its first commercial product.
The price is the headline for infrastructure buyers
Meta priced API access at $3 per 1,000 audio-minutes, equivalent to $0.18 per hour. MarkTechPost's comparison places that well under Cartesia's Ink-2 at $4 per 1,000 minutes, and less than half of both ElevenLabs Scribe v2 and Deepgram Flux, which it lists at $6.50 per 1,000 minutes apiece. Notably, on raw speed ElevenLabs Scribe v2 edges out Muse Voice Transcribe's first partial transcript (3.6% WER at 0.14 seconds versus Meta's 0.13 seconds is roughly a tie, per the same comparison), so Meta is not claiming a clean sweep on every metric -- the model's advantage is the combination of competitive accuracy, single-pass diarization and endpointing, and aggressive pricing all at once.
Why a coding-and-chatbot lab is shipping speech infrastructure
Muse Voice Transcribe is MSL's second public product in a month, after Muse Code. Wang was brought in specifically to accelerate Meta's AI roadmap following the company's roughly $14 billion investment tied to his hire, and voice has become one of the lab's visible expansion points alongside coding agents -- Meta AI's own Gemini-Live-style voice briefings and system-wide dictation both need exactly the streaming ASR-plus-diarization-plus-endpointing stack Muse Voice Transcribe provides. Selling that stack through the Meta Model API, rather than keeping it purely internal, puts Meta in direct commercial competition with the specialist speech-AI vendors -- Deepgram, AssemblyAI, Cartesia, ElevenLabs -- that voice-agent and contact-center platforms currently build on.
What this means for the voice AI and calling industry
Endpointing is the hard problem for voice agents
Knowing exactly when a caller has finished speaking -- not too early, not with an awkward pause -- is what makes an AI phone agent feel natural instead of robotic. Folding that decision into the transcription model itself, rather than bolting it on afterward, is a direct attack on the turn-taking latency that still trips up many voice agents today.
Big Tech is commoditizing voice-AI infrastructure
At $3 per 1,000 minutes, Meta is pricing a core building block of every voice agent and call-center analytics stack at less than half what specialist vendors charge. That puts pricing pressure on Deepgram, AssemblyAI, Cartesia and ElevenLabs precisely in the layer -- ASR plus diarization -- that most voice-AI products are built on top of.
Multi-speaker, multilingual calls get easier to build for
Native handling of 20-plus speakers and mid-sentence code-switching across 70-plus languages, in a single real-time pass, removes a chunk of the custom engineering that multilingual contact centers and multi-party call platforms currently have to stitch together from separate ASR and diarization tools.
Related reading
Fireflies.ai Launches Voice Agents, Already Handling 40,000+ Calls
Another voice-AI vendor moving from passively processing calls to actively running them -- the layer just above the transcription infrastructure Meta shipped.
Read the story →Do AI Receptionists Sound Robotic?
Why turn-taking, latency and endpointing -- the exact problems Muse Voice Transcribe targets -- are what separate a natural-sounding AI receptionist from a stilted one.
Read the post →AI Receptionist Multi-Language Support
How multilingual call handling works today, and why cheaper, more accurate multilingual transcription infrastructure matters for it.
Read the post →