Microsoft now owns both sides of the conversation
On October 1, Microsoft AI announced MAI-Transcribe-2-Streaming, its first real-time speech-to-text model, alongside MAI-Voice-2.1 and MAI-Voice-2.1-Flash. The three are built to work as a set: one listens, two speak. They are available through Microsoft Foundry, Azure AI Speech, the MAI Playground, and third-party gateways like Vercel's AI Gateway and OpenRouter. Microsoft demoed them together as a voice agent called Chatter.
The threshold between walkie-talkie and phone call
The pitch is latency. Rather than waiting for a speaker to finish, the streaming model accepts continuous audio and emits partial transcripts that it revises as more context arrives, letting an application act on speech before the sentence is done. One analysis of the numbers put it plainly: roughly 130 milliseconds to hear, 350 to think up a reply, 150 to start speaking. A voice agent built entirely on Microsoft's own stack can complete a turn in under a second, which is where voice stops sounding like a walkie-talkie and starts sounding like a phone call.
For developers, the real significance is ownership. Microsoft is assembling a complete multimodal stack on its own models, reducing its dependence on OpenAI and other suppliers for the voice interface of its agent platform. Whether that yields better products or just better margins is a separate question.
Read the footnotes
Microsoft says the transcription model ranks no. 1 on Artificial Analysis for both partial and final accuracy. Those are vendor-reported benchmark positions, not independently reproduced results. The 100-millisecond figure describes how fast the model produces a hypothesis, not how fast your deployed agent completes a turn, which also depends on the network, the application, and the model doing the reasoning.
Then there is the status. Everything here is public preview, which means no service-level agreement and a label that says not recommended for production. The pricing is introductory: $0.54 per audio hour runs through December 31, and Microsoft has not said what it costs on January 1. Streaming transcription is also priced substantially above batch, putting a clear price on low latency.
The voice models carry the sharpest footnote. They can clone a voice from seconds of reference audio, with what Microsoft calls built-in safeguards against misuse. In one of its own tests, about half of 4,000 participants believed the generated voices belonged to a real person. A safeguard described in one sentence deserves more than one sentence of scrutiny.
What this means in practice
If you build voice agents on Azure, you now have a priced, first-party option for both hearing and speaking, which simplifies procurement even if it does not settle the quality question. If you do not, this is still worth watching as a vendor play: voice is becoming the interface layer for agents, and the companies that own the stack get to set its economics. Watch what happens to the transcription pricing in January. That will say more about Microsoft's strategy than any benchmark.
Sources
- [1] Microsoft AI official announcementRead source
- [2] The Decoder coverageRead source
- [3] Let's Data Science coverage — preview caveat and benchmark detailRead source
- [4] Runtime Wire coverageRead source