Microsoft has launched MAI-Transcribe-2-Streaming model, with Microsoft AI CEO Mustafa Suleyman, calling it the most accurate real-time audio transcription model in the world. Unveiled alongside two companion speech generation models – MAI-Voice-2.1 and the high-speed MAI-Voice-2.1-Flash – the release is designed to provide developers with responsive, low-cost building blocks for building autonomous conversational voice agents.
The new voice architecture delivers inference speeds that are 55% faster and operational costs that are 60% cheaper than comparable offerings like ElevenLabs, according to Suleyman. He said in a post on X (formerly Twitter): “We’re launching the most accurate real time transcription model in the world… #1 !!! 55% faster and 60% cheaper than ElevenLabs. The headline addition to the company’s audio lineup, MAI-Transcribe-2-Streaming, provides continuous, low-latency live speech-to-text across 60 languages with built-in automatic language detection. 1 spot for accuracy across both partial and completed transcripts on the independent benchmark platform Artificial Analysis. Instead of pausing until a sentence ends, the model outputs initial text hypotheses within roughly 100 milliseconds of receiving incoming audio. To pair with speech recognition, the team introduced MAI-Voice-2.1, its flagship multilingual text-to-speech engine covering 23 languages and 26 regional locales. MAI-Voice-2.1-Flash was introduced by the company for mission-critical, high-capacity applications that demand immediate response times. The Flash version offers the same 23-language voice consistency as the base model, but it is optimized for high-throughput enterprise workloads, and can produce 45 seconds of synthesized audio with an end-to-end latency of just 150 milliseconds.
Come build agents on our platform! Key performance benchmarks and features include top benchmark ranking as the model has secured the No. It sits on the Pareto frontier for accuracy versus latency, eliminating the historical tradeoff where high precision required sluggish processing times. It refines words as additional context arrives and finalizes stable text immediately. For live captioning and automated dictation workflows, internal tests show text appears on-screen twice as fast as the nearest market competitor. The system’s standout capability is cross-lingual identity preservation. A single synthetic voice persona can switch seamlessly between diverse languages, such as English, Mandarin and German, without carrying a foreign accent over from one language to another. Instead, the synthesized speaker modifies actual local pronunciation and retains a unique vocal identity. That means brands can deploy uniform virtual ambassadors worldwide, educational sites can switch languages while keeping the same instructors, and customer support agents can respond in the caller’s language without missing a beat. You use AI every day. Now get your AI Quotient. Take the AIQ test.
By surfacing partial transcripts as words are spoken, voice bots can trigger API tools and start logical reasoning before the user finishes speaking.

