Microsoft Targets Contact Center Market With Real-Time AI Voice Models
By PYMNTS

AI summary of the source article
Microsoft has released three new artificial intelligence voice models aimed at powering customer service agents, multilingual assistants, and live captions. The MAI-Transcribe-2-Streaming model continuously converts live speech into text across 60 languages with low-latency hypotheses. The MAI-Voice-2.1 model offers expressive text-to-speech generation across 23 languages, while MAI-Voice-2.1-Flash is optimized for high-volume, responsive voice applications supporting the same languages. The launches follow Microsoft's July announcement of a $2.5 billion investment in the Microsoft Frontier Company to embed 6,000 engineering and industry experts with customers to assist with enterprise artificial intelligence implementation.
Why it matters
The real-time models allow enterprise systems, such as customer-service agents and assistants, to begin reasoning and identifying requests before a speaker finishes their sentence. They also offer developers flexibility to balance speech fidelity against responsiveness and operating costs at scale.
Key facts
- Microsoft released three new AI voice models: MAI-Transcribe-2-Streaming, MAI-Voice-2.1, and MAI-Voice-2.1-Flash.
- MAI-Transcribe-2-Streaming transcribes live speech across 60 languages, while the text-to-speech models support 23 languages.
- Microsoft previously announced a $2.5 billion investment in its Microsoft Frontier Company unit to embed 6,000 experts with customers implementing AI.