Microsoft AI moved to aggressively reshape the landscape of automated speech recognition on Thursday with the launch of MAI-Transcribe-2. The new model, which the company claims is faster, more accurate, and significantly more cost-effective than offerings from established AI giants like OpenAI, Google, and ElevenLabs, marks a pivotal moment in the industry. Perhaps most disruptive is the pricing: Microsoft has pegged the service at just 10 cents per hour of audio, a move that fundamentally changes the economics of large-scale transcription.
This pricing strategy demands attention. When Microsoft first introduced a model in this specific product line a mere five months ago, it was priced at $0.36 per hour. Thursday’s early-bird rate represents a staggering 72% reduction in costs. For an enterprise handling 100,000 hours of call-center audio annually—a volume that is routine for major financial institutions or telecommunications firms—the financial impact is substantial. The annual bill for that organization would plummet from $36,000 to just $10,000. At this price point, transcription shifts from a complex operational line item that requires constant negotiation and budget justification to a commodity that can be deployed at scale without friction.
The release is the latest evidence of a broader, more ambitious strategy within Microsoft. Two years ago, the idea that Microsoft would build its own frontier-class models to replace the technology it had heavily invested in from partners like OpenAI would have seemed implausible. Today, it is the company’s operating reality. Microsoft is methodically building its own models, one modality at a time, and seamlessly integrating them into its product ecosystem. Transcription has become the fastest-moving front in this effort, and MAI-Transcribe-2 serves as the most potent proof point yet of Microsoft’s ability to compete in the AI space while reducing its reliance on the partnership it famously bolstered with $13 billion in investment.
Enhancing Features for the Enterprise Market
MAI-Transcribe-2 is not merely a cheaper version of its predecessors; it is a significantly more capable tool designed for the messy, unpredictable realities of business communications. While competitors often struggle with clean audio, Microsoft claims its new model is engineered to handle real-world challenges, such as heavy background noise, low-quality recordings, and the overlapping speech common in high-pressure call centers. The model supports 60 languages, a substantial increase from the 43 supported by MAI-Transcribe-1.5 in June and the 25 available at the initial launch in April.
Beyond basic transcription, Microsoft has bundled a suite of high-value features into the base product. For enterprise buyers, the inclusion of speaker diarization is critical; by accurately distinguishing between different voices in a multi-party conversation, the model converts raw, unreadable blocks of text into structured, useful meeting transcripts. Furthermore, the addition of word-level timestamps enables sophisticated search and editing capabilities, allowing users to align text precisely with video or audio segments.
The inclusion of keyword biasing is another significant addition, allowing developers to provide the model with a list of specialized terminology—such as pharmaceutical names, product codes, or industry-specific jargon—to prevent transcription errors. Coupled with automatic language identification, which removes the need for users to manually declare the language of an audio file, these tools address the primary pain points that have traditionally forced businesses to pay premium rates to specialty vendors.
Two specific features highlight Microsoft’s focus on the enterprise: a "verbatim" output mode, which captures every filler word, stutter, and false start for compliance and legal documentation, and a "clean" mode that strips these elements for professional captions and summaries. Additionally, the model’s support for code-switching—the ability to move between languages mid-sentence—is a major boon for global businesses. By specifically addressing language mixing like Hinglish or Spanglish, Microsoft is catering to diverse markets where customer service interactions frequently toggle between languages. Historically, vendors charged extra for such capabilities, but Microsoft is now including the full suite in its 10-cent-per-hour rate.
Understanding the Benchmark Claims
Microsoft’s competitive positioning relies on three distinct performance claims, each utilizing a different metric. First, the company asserts that MAI-Transcribe-2 ranks number one on the FLEURS benchmark across 60 languages, with an average word error rate (WER) of 5.2%. FLEURS, a dataset published by Google researchers in 2022, is the industry standard for multilingual evaluation because it uses identical content across over 100 languages. While a 5.2% WER is impressive, it is important to note that the average has risen slightly from the 3.7% reported for MAI-Transcribe-1.5. Experts suggest this is likely due to the inclusion of more challenging, low-resource languages rather than a regression in performance, but it underscores the need for users to look at per-language performance breakdowns.
The second claim concerns the Artificial Analysis word-error-rate leaderboard. Artificial Analysis independently benchmarks models through public APIs, testing them on a mix of simulated agent conversations, earnings calls, and parliamentary speeches. Microsoft reports that MAI-Transcribe-2 has climbed to second place, effectively surpassing ElevenLabs. Perhaps more importantly, the model defines the current "Pareto frontier" for the industry, meaning that no rival currently beats Microsoft on accuracy without sacrificing speed, and none beats it on speed without sacrificing accuracy.
Finally, there is the question of raw speed. According to evaluations by Artificial Analysis, Microsoft’s model is 10 times faster than OpenAI’s GPT-Transcribe, seven times faster than ElevenLabs’ Scribe v2, and five times faster than Google’s Gemini 3.5 Transcribe. In the world of batch transcription, this efficiency is not just about convenience; it is about throughput. A model that runs at 300 times real-time speed requires significantly fewer GPU hours than one running at 30 times, and this efficiency is the structural foundation that allows Microsoft to lower prices while maintaining profitability.
A Rapid Cadence of Innovation
The speed of Microsoft’s release cycle is perhaps the most compelling part of the narrative. Between April and August, the company moved from a 25-language model at $0.36 an hour to a 60-language model at $0.10 an hour, while simultaneously adding complex features that were once considered premium-only. This rapid-fire development cycle suggests that the team behind these models has arrived at a stable, scalable architecture and is now optimizing its data pipeline.
This organizational shift was heavily influenced by Mustafa Suleyman, the CEO of Microsoft AI. Since joining the company, Suleyman has championed a "small, focused" team structure, liberating engineers from traditional corporate bureaucracy to accelerate development. The success of the transcription line serves as the primary test case for whether this flattened, agile approach can consistently produce commercial-grade results, rather than just academic research.
Independence and the Future of AI Strategy
The decision to build these models in-house has profound implications for Microsoft’s relationship with OpenAI. Despite a $13 billion investment and deep integration of OpenAI models across Azure and Office, the relationship has clearly shifted. Following the restructuring of their partnership in late 2025 and subsequent amendments in 2026—which ended Microsoft’s exclusive access to OpenAI’s technology and eliminated revenue-share requirements—Microsoft has gained the freedom to pursue its own path toward artificial general intelligence.
This autonomy is not just about control; it is about margins. Every prompt processed by a third-party model incurs a cost, whereas routing that same prompt through an in-house model running on Microsoft’s own infrastructure is significantly cheaper. The company has already begun replacing third-party models in core products like Word and Excel with its own MAI models. Transcription is the logical starting point for this substitution strategy because the goals are well-defined and the metrics for success are objective. With control over massive volumes of meeting audio via Teams and clinical documentation via Nuance, Microsoft has a clear internal roadmap to shift these workloads to its own infrastructure, effectively ending its reliance on external partners for these services.
While Microsoft is clearly positioning itself against other frontier labs, the long-term impact of these changes will be felt most by specialist transcription firms. By offering a high-performance, feature-rich model at a price point that many incumbents struggle to match, Microsoft is effectively commoditizing the transcription market. The specialists’ only remaining defense is deep domain expertise, such as highly specific medical or legal vocabularies, but even that is under pressure from features like keyword biasing.
As technical decision-makers consider moving their workloads to Microsoft’s new offering, there are still practical questions to be answered. The longevity of the 10-cent price point, the availability of real-time streaming capabilities, and the specifics of data handling remain top priorities for companies in regulated industries. Whether these audio inputs are used to train future iterations of the model is a question that will undoubtedly be at the center of upcoming negotiations.
For now, Microsoft is demonstrating a clear, calculated shift. By treating transcription as a modular, cost-optimized service, the company is proving that it would rather own the manufacturing process than pay for the output. As MAI-Transcribe-2 becomes available on Microsoft Foundry, it stands as a template for how the world’s most valuable software company intends to navigate the next phase of the AI race: by building its own infrastructure, lowering costs, and aggressively claiming the market for itself.

