Alibaba Qwen Audio 3.0 TTS Plus Dominates Voice Ranking

Le brief IA que les pros lisent chaque soir
Les 7 actus IA du jour, décryptées en 5 min. Gratuit.
Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.
Choisis ton rythme
Gratuit · Pas de spam · Désabonnement en 1 clic
Alibaba Qwen Audio 3.0 TTS Plus Dominates Voice Ranking
Alibaba's new text-to-speech model, Qwen-Audio-3.0-TTS-Plus, is leading the supplier voice rankings in the Artificial Analysis dashboard. With an Elo score of 1,236, it narrowly surpasses Simba 3.2 (1,234). Gemini 3.1 Flash TTS (1,214) and Sonic 3.5 (1,207) follow behind.
The model comes in two versions. Flash is designed for real-time interaction with around 300 milliseconds of latency, while Plus aims for high-quality voice output. It supports 16 languages, including less commonly covered languages such as Tagalog, Malay, Thai, and Vietnamese, as well as several Chinese dialects. Users can steer the speech style with natural language or add non-verbal cues using tags like "[angry]" or "[laughing]." Alibaba also claims that the model handles noisy or reverberant reference recordings better than previous versions when cloning voices.
Speed is a weak point: at 16 characters per second, it is significantly outpaced by Sonic 3.5 (120) and Simba 3.2 (30.2). The price is $27.60 per million characters via Alibaba Cloud Model Studio. A collection of audio samples is available.
Brief IA — L'actualité IA en français
L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.