
Microsoft AI has released MAI-Transcribe-2-Streaming, its first streaming speech-to-text (STT) model. It launched on October 1, 2026, alongside 2 text-to-speech models, MAI-Voice-2.1 and MAI-Voice-2.1-Flash. Artificial Analysis ranks it #1 of 38 models for final and first partial transcript accuracy. The model targets voice agents, live captions and dictation, where latency decides the experience. What Microsoft Shipped MAI-Transcribe-2-Streaming is the real-time sibling of the batch MAI-
Microsoft AI released a real-time speech-to-text model that converts spoken audio into text while a speaker is still talking, rather than requiring them to finish first. According to independent testing, it ranks first among 38 competing models for accuracy, delivering both initial partial transcripts and final versions with equal 2.5% error rates at extremely low latency. The model supports 60 languages with automatic detection and costs $0.54 per hour under an introductory pricing structure. This capability matters for applications like voice agents and live captioning where speed and accuracy determine user experience.

What are Decision AI Models? Decision AI models are a new class of model that returns a decision, not a paragraph. You send text with typed questions. The model returns choices, scores or yes/no probabilities your code can branch on directly. The category went mainstream when TypeSafe AI launched Jev after 2 years in stealth. TypeSafe calls it a ‘System One model,’ after Daniel Kahneman’s fast, intuitive System 1 thinking. Within 3 weeks, Fastino Labs shipped 2 rival

Making enterprise reporting tools openly available could democratize access to automation that typically requires expensive proprietary software licenses.

Mobile Systems
Want to go deeper than the news? Explore live, cohort-based AI courses taught by practitioners.
Browse AI courses on Maven