Microsoft Launches MAI-Transcribe-2, Undercutting Meta and OpenAI on Price
Microsoft AI released MAI-Transcribe-2 on September 3, 2026 -- a speech-to-text model the company says is the fastest, most accurate and cheapest available, priced at $0.10 per audio hour. That's a 72% cut from the model's own predecessor five months earlier, and it undercuts even the aggressively priced transcription model Meta shipped just two days before it.
What Microsoft announced
Microsoft AI -- the division led by Mustafa Suleyman -- published MAI-Transcribe-2 on its own news blog on September 3, 2026, calling it "the fastest, most accurate and cheapest speech recognition model in the world." The model is Microsoft's third release in its transcription line in under a year, following MAI-Transcribe-1 in April 2026 and a 1.5 update in June, and it lands at a moment when every major AI lab is racing to commoditize the speech-to-text layer that voice agents and call-center software are built on top of.
According to Microsoft's own announcement, MAI-Transcribe-2 ranks first on the FLEURS benchmark across 60 languages, with an average word-error rate of 5.2%. On the independent Artificial Analysis leaderboard, Microsoft says the model ranks second on raw word-error rate but defines the Pareto frontier for the combination of accuracy and latency -- meaning no other model available today is simultaneously more accurate and faster. Per that same benchmark, Microsoft claims MAI-Transcribe-2 transcribes roughly 10 times faster than OpenAI's GPT-Transcribe, 7 times faster than ElevenLabs' Scribe v2, and 5 times faster than Google's Gemini 3.5 Transcribe, while still coming out more accurate than all three.
Source: Microsoft AI, September 3, 2026.
The price cut, in real dollars
Microsoft priced MAI-Transcribe-2 at $0.10 per hour of audio as an introductory rate through the end of 2026. That's down from the $0.36 per hour it charged for the original MAI-Transcribe-1 just five months earlier -- a roughly 72% reduction in less than half a year. VentureBeat's coverage of the launch translates that into a concrete enterprise scenario: a large business processing 100,000 call-center hours a year would have paid about $36,000 under the original pricing; at the new rate, the same volume costs roughly $10,000. Language coverage grew alongside the price cuts -- from 25 languages at the April 2026 launch, to 43 with the June 1.5 release, to 60 now.
The pricing also lands below what Meta charged for Muse Voice Transcribe, the real-time transcription model Meta Superintelligence Labs launched on September 1, 2026 and which Botnira covered at the time. Meta priced that model at $3 per 1,000 audio-minutes, which works out to $0.18 per hour -- nearly double Microsoft's $0.10 rate. Coming two days apart, the two launches amount to back-to-back price cuts at the infrastructure layer that voice-agent and contact-center platforms depend on.
Source: VentureBeat, September 3, 2026.
What the model actually does
Beyond raw transcription, MAI-Transcribe-2 bundles features that voice-AI teams typically have to source from separate tools: speaker diarization, word-level timestamps, automatic language identification, keyword biasing for domain-specific terms, and code-switching support for conversations that shift between languages mid-sentence. It also offers configurable output styles -- a "verbatim" mode that preserves filler words and false starts for compliance or legal use, and a "clean" mode that strips them out for readability. Microsoft says the model was built to handle noisy, real-world audio rather than just clean studio recordings, and lists clinical note-taking, legal documentation, accessibility captioning and enterprise call transcription among its target use cases.
The model is currently available in public preview through Microsoft Foundry (formerly Azure AI Foundry), the MAI Playground, and on OpenRouter. Microsoft is explicit that the preview ships without a service-level agreement and is not yet recommended for production workloads -- a caveat enterprise buyers evaluating it for live call-center deployment will need to weigh against the aggressive pricing.
Why Microsoft is targeting frontier labs, not specialists
Speech-to-text has historically been dominated by specialist vendors -- Deepgram, AssemblyAI, Cartesia -- that contact-center and voice-agent platforms build on because general-purpose AI labs treated transcription as a secondary feature. VentureBeat's analysis of the launch notes that Microsoft is instead positioning MAI-Transcribe-2 against the transcription offerings of other frontier labs -- OpenAI, Google and ElevenLabs -- rather than against those established specialists, while its $0.10-per-hour price still undercuts many of the specialist vendors' enterprise contracts. In an April 2026 interview with The Verge, Suleyman described the transcription effort as coming from a "small, focused" roughly-10-person team operating with unusual autonomy inside Microsoft AI, and framed the broader ambition behind the MAI model line as proving these systems can "deliver product value for the millions of enterprises" that already run on Microsoft infrastructure.
What this means for the voice AI and calling industry
The transcription layer is in a price war
Two frontier-lab transcription launches landed two days apart -- Meta at $0.18/hour, then Microsoft at $0.10/hour -- both undercutting specialist vendors. For any voice agent or call-center platform built on third-party ASR, the cost of that core building block is falling fast.
Call-center transcription costs are dropping by two-thirds or more
Microsoft's own math -- $36,000 down to $10,000 for 100,000 call-center hours a year -- shows what this means in practice for any business running large volumes of recorded or live calls through a transcription pipeline.
Multilingual accuracy keeps climbing
Language coverage more than doubled in five months (25 to 60), with built-in code-switching and keyword biasing -- reducing the custom engineering multilingual contact centers previously needed to stitch together from separate ASR tools.
Related reading
Meta Launches Muse Voice Transcribe, Undercutting Rival Speech AI on Price
The transcription-model launch from two days earlier that Microsoft's pricing just undercut.
Read the story →Fireflies.ai Launches Voice Agents, Already Handling 40,000+ Calls
A voice-AI vendor building on top of exactly this kind of transcription infrastructure to run entire calls, not just document them.
Read the story →Do AI Receptionists Sound Robotic?
Why transcription accuracy and latency -- the exact metrics Microsoft is competing on -- shape how natural an AI receptionist sounds on a live call.
Read the post →