top of page

Microsoft launches MAI-Transcribe-2-Streaming and MAI Voice 2.1 models for real-time multilingual AI agents

2 hours ago
7 min read
Microsoft MAI-Transcribe-2-Streaming and MAI Voice 2.1 models for real-time multilingual AI agents

Microsoft has launched three new first-party audio models designed to reduce the delay between hearing a user, reasoning about the request and producing a spoken response: MAI-Transcribe-2-Streaming, MAI-Voice-2.1 and MAI-Voice-2.1-Flash.


MAI-Transcribe-2-Streaming continuously transcribes speech across 60 languages instead of waiting for an utterance to finish. Microsoft says the model begins producing partial hypotheses just over 100 milliseconds after receiving audio, allowing an application to start reasoning or prepare a tool call while the user is still speaking.


On the output side, MAI-Voice-2.1 expands Microsoft's multilingual speech generation to 23 languages and 26 locales, while MAI-Voice-2.1-Flash is optimized for high-volume conversational systems where latency and inference cost are more important.


The three models are available through Microsoft Foundry, creating a first-party Microsoft stack that can cover both ends of a real-time voice-agent loop.


··········


THE THREE NEW MAI AUDIO MODELS AT A GLANCE


........


Model

Primary role

Key specification

Price

MAI-Transcribe-2-Streaming

Real-time speech-to-text

60 languages

$0.54/hour introductory

MAI-Voice-2.1

Expressive text-to-speech

23 languages / 26 locales

$22 / 1M characters

MAI-Voice-2.1-Flash

Low-latency text-to-speech

~45 ms model inference

$15 / 1M characters


........


The transcription price is introductory and applies through the end of 2026.


MAI-Voice-2.1 and its Flash variant support the same multilingual voice identity system, allowing one cloned or selected voice to continue across supported languages rather than switching to a different speaker profile.


··········


STREAMING TRANSCRIPTION LETS THE AGENT START WORKING BEFORE THE USER STOPS TALKING


Traditional speech-to-text pipelines often treat an utterance as a completed object: audio arrives, the speaker stops, transcription finishes and only then does the language model receive usable text.


MAI-Transcribe-2-Streaming changes that sequence by returning partial transcripts continuously.


The first hypotheses appear just over 100 milliseconds after audio begins arriving. They are then revised as additional context becomes available before the system commits a stable final transcript.


That architecture allows downstream computation to overlap with speech.


If a caller says, “I need to move my reservation from Friday to…”, an agent does not necessarily need to wait for the complete sentence before recognizing that the task concerns a reservation change. It can begin loading reservation data or preparing the relevant tool while transcription continues.


Microsoft reports that words appear in the live transcript in most cases roughly 320 milliseconds after they are spoken, compared with more than 500 milliseconds for the nearest competitor in its published evaluation.


The practical gain is not simply faster subtitles. It shortens the time between human speech and the point at which an agent can begin useful computation.


··········


MAI-TRANSCRIBE-2-STREAMING SUPPORTS 60 LANGUAGES WITH CONTINUOUS LANGUAGE DETECTION


The model supports 60 languages and automatic continuous language detection.


That is particularly relevant for voice agents operating in multilingual environments because language identification does not need to be treated as a separate preliminary stage.


Microsoft distinguishes the streaming model from the existing MAI-Transcribe-2.


MAI-Transcribe-2 remains the more appropriate option when a completed recording needs features such as speaker diarization and word-level timestamps.


MAI-Transcribe-2-Streaming is optimized for situations where an application needs to process speech as it arrives.


The distinction separates two different latency requirements: transcription quality for completed audio versus incremental understanding for live interaction.


··········


MICROSOFT REPORTS NUMBER-ONE ACCURACY FOR BOTH PARTIAL AND FINAL TRANSCRIPTS


Microsoft says MAI-Transcribe-2-Streaming ranks first on Artificial Analysis for the accuracy of both its partial and final transcripts.


Partial-transcript accuracy is especially relevant to agent systems.


A streaming model can produce text extremely quickly by making aggressive early guesses, but those guesses are less useful if they are frequently revised. An agent that begins acting on incorrect partial text can waste the latency advantage by starting the wrong operation.


The engineering target is therefore a combination of low latency and useful early accuracy, rather than minimum latency in isolation.


Microsoft says the model sits on the accuracy-versus-latency Pareto frontier in the published evaluation, meaning its latency improvement is not achieved simply by accepting a large accuracy penalty.


··········


MAI-VOICE-2.1 EXPANDS THE SAME VOICE ACROSS 23 LANGUAGES


MAI-Voice-2.1 is Microsoft's new higher-fidelity speech-generation model.


It supports 23 languages across 26 locales, up from the 15-plus-language coverage of the previous MAI-Voice-2 generation.


The model is designed to preserve the identity of a speaker while changing languages.


A single assistant can therefore speak English, Mandarin or German while retaining the same recognizable voice and adapting pronunciation to the target language rather than carrying one accent across every output.


MAI-Voice-2.1 also supports granular emotion control and zero-shot voice prompting from a short reference recording.


Microsoft lists model inference latency at approximately 550 milliseconds for generating 45 seconds of audio, with pricing of $22 per million characters.


The target workloads are applications where speech quality and expressive fidelity take priority over minimum latency, including narration, content creation, education and branded voice experiences.


··········


MAI-VOICE-2.1-FLASH CUTS MODEL INFERENCE LATENCY TO ABOUT 45 MILLISECONDS


MAI-Voice-2.1-Flash keeps the same 23-language coverage and cross-language voice capability but changes the optimization target.


Microsoft lists its model inference latency at approximately 45 milliseconds for a 45-second audio generation workload, compared with approximately 550 milliseconds for MAI-Voice-2.1.


The company separately reports around 150 milliseconds of end-to-end latency for producing 45 seconds of audio once the surrounding serving pipeline is included.


The distinction between model inference and end-to-end latency is important: the first measures model execution, while the second also includes system overhead around the request.


Flash costs $15 per million characters, 31.8% less than the $22 rate for MAI-Voice-2.1.


It is therefore the more natural deployment option for call-center agents, IVR systems and conversational assistants where every turn can generate new speech.


··········


VOICE 2.1 ALSO REPRESENTS A LARGE LATENCY IMPROVEMENT OVER MICROSOFT'S PREVIOUS GENERATION


Microsoft's published model specifications allow a direct comparison with MAI-Voice-2.


........


Model

Previous generation

New generation

Change

Standard MAI Voice inference latency

~1,000 ms

~550 ms

45% lower

MAI Voice Flash inference latency

~225 ms

~45 ms

80% lower

Language coverage

15+

23

Expanded

Standard price

$22 / 1M chars

$22 / 1M chars

Unchanged

Flash price

$15 / 1M chars

$15 / 1M chars

Unchanged


........


Data Studios calculation: based on Microsoft's published inference figures, MAI-Voice-2.1 reduces standard-model inference latency by approximately 45%, while MAI-Voice-2.1-Flash reduces Flash inference latency by approximately 80% relative to the previous generation.


The API prices remain unchanged.


Microsoft has therefore used the 2.1 refresh primarily to move the latency and language-coverage frontier, rather than to lower the nominal price per generated character.


··········


THE VALUE OF THE THREE-MODEL STACK IS IN OVERLAPPING LISTENING, REASONING AND SPEAKING


A conversational agent can be simplified into four stages:


hear → understand → decide → speak


If every stage waits for the previous stage to finish completely, latency accumulates sequentially.


Streaming transcription allows the understanding and reasoning stages to begin while speech is still arriving. Low-latency synthesis then reduces the delay between the agent reaching an answer and the user hearing it.


The resulting system can perform useful work during time that would otherwise be spent waiting.


For example, MAI-Transcribe-2-Streaming can recognize an emerging request, the agent can query a booking system before the speaker finishes, and MAI-Voice-2.1-Flash can begin producing the response shortly after the tool result arrives.


For agent developers, this means voice latency is increasingly a pipeline problem rather than a single-model problem.


Reducing 200 milliseconds at one end of the interaction can create additional time for reasoning, retrieval or tool execution elsewhere without increasing the total conversational pause perceived by the user.


··········


THE COST DIFFERENCE BECOMES MATERIAL AT CALL-CENTER SCALE


MAI-Transcribe-2-Streaming costs $0.54 per hour of audio under its introductory pricing through the end of 2026.


A simple Data Studios calculation shows the scale:


100 hours of streamed audio = $54


10,000 hours = $5,400


For speech generation, 10 million generated characters would cost:


MAI-Voice-2.1: $220


MAI-Voice-2.1-Flash: $150


The $70 difference is modest for a small deployment but compounds as conversational volume grows.


Actual voice-agent economics also include language-model inference, retrieval, tool calls, telephony infrastructure and periods where audio is transmitted without requiring synthesized output, so these figures should be treated as component-level costs rather than total cost per conversation.


··········


VOICE CLONING REMAINS AVAILABLE WITH CONSENT GUARDRAILS


Both MAI-Voice-2.1 variants support voice matching from a short reference clip without requiring model fine-tuning.


The capability extends across the supported languages, so a cloned speaker identity can be reused in multilingual applications.


Microsoft says consent protections are built into the system to prevent unauthorized voice cloning.


For enterprise deployments, that makes the technology suitable for controlled brand voices and authorized speakers while placing identity verification and consent inside the product architecture rather than treating them purely as application-level controls.


··········


MICROSOFT IS MAKING THE MODELS AVAILABLE BEYOND A SINGLE API SURFACE


All three models are available through Microsoft Foundry, including Azure Speech integration and direct Foundry API access.


Microsoft is also distributing the new MAI voice models through additional developer platforms. MAI-Voice-2.1 and MAI-Voice-2.1-Flash are available through OpenRouter, while Microsoft also lists Vercel and LiveKit among the supported or planned access channels.


The company has additionally built a demonstration called Chatter inside the MAI Playground to show the transcription and speech models operating together in a live conversational agent.


This distribution strategy makes the models usable as independent speech components rather than limiting them to Microsoft Copilot products.


··········


MICROSOFT IS OPTIMIZING VOICE AGENTS AROUND THE COMPLETE CONVERSATIONAL LOOP


The new MAI release is less about a single speech benchmark than about reducing latency across an entire agent interaction.


MAI-Transcribe-2-Streaming moves speech recognition into the period while the user is still talking. MAI-Voice-2.1 expands expressive multilingual generation, while MAI-Voice-2.1-Flash pushes model inference for latency-sensitive applications down to roughly 45 milliseconds.


Microsoft has simultaneously preserved the previous Voice-2 pricing while materially increasing language coverage and lowering inference latency.


For real-time AI agents, the resulting architecture creates more room for the expensive middle of the interaction — reasoning, retrieval and tool execution — without forcing the user to spend that time waiting in silence.


··········


FOLLOW US FOR MORE.


DATA STUDIOS


datastudios.org

bottom of page