⚡
Brief IA
›

Microsoft AI Launches MAI-Transcribe-2 and MAI-Voice-2.1

🎨 Creative AI·Tom Levy·

Microsoft AI Launches MAI-Transcribe-2 and MAI-Voice-2.1

Microsoft AI Launches MAI-Transcribe-2 and MAI-Voice-2.1
⚡
Key Takeaways
1Microsoft AI launches MAI-Transcribe-2-Streaming for real-time transcription
2Two new text-to-speech models, including MAI-Voice-2.1, cover 23 languages
3The models are offered at introductory prices and are accessible on multiple platforms
💡Why it matters — These innovations strengthen Microsoft AI's offerings for conversational agents by combining speed, multilingualism, and security measures for voice cloning.
⚡Le brief IA que lisent les pros

Le brief IA que les pros lisent chaque soir

Les 7 actus IA du jour, décryptées en 5 min. Gratuit.

Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.

Choisis ton rythme

Gratuit · Pas de spam · Désabonnement en 1 clic

📄
Full Analysis

Microsoft AI adds real-time transcription and two synthetic voices to its catalog for conversational agents. The company claims very low response times, broad language support, and security measures to regulate voice cloning, along with promotional pricing and wide distribution.

Regulated voice cloning and perceived naturalness in testing

Microsoft AI's two voice synthesis models are capable of reproducing a voice from just a few seconds of reference audio. Microsoft AI specifies that it has integrated security measures to limit abusive uses. In an experiment, about half of the 4,000 participants estimated that the generated voices belonged to a human.

Announced performance in transcription and synthesis

Microsoft AI introduced MAI-Transcribe-2-Streaming, an instant transcription system that the company claims is at the top of the accuracy ranking established by Artificial Analysis. This system covers 60 languages and displays its first intermediate results in just over 100 milliseconds, which would allow a voice agent to intervene while the user is still speaking. Two additional voice synthesis models are available: MAI-Voice-2.1, designed to express itself in 23 languages with the same voice and a native accent for each, and its variant MAI-Voice-2.1-Flash, announced with a latency of 150 milliseconds.

Availability platforms and announced pricing

MAI-Transcribe-2-Streaming is offered at $0.54 per hour of audio until the end of the year, as an introductory rate. For voice synthesis, MAI-Voice-2.1-Flash is priced at $15 per million characters, compared to $22 for the other pricing tier. The models are accessible via Microsoft Foundry and MAI Playground, and both voice models are also available on OpenRouter.

⚡

Brief IA — L'actualité IA en français

L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.