Open TTS Leaderboard Assesses Speed, Clarity, and Streaming

Le brief IA que les pros lisent chaque soir
Les 7 actus IA du jour, décryptées en 5 min. Gratuit.
Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.
Choisis ton rythme
Gratuit · Pas de spam · Désabonnement en 1 clic
A new open-source ranking of text-to-speech synthesis focuses on objective measures and an explicit streaming protocol. It covers multiple languages, compares voice cloning, and opens the door to community feedback without replacing human preferences.
The Streaming Ranked by TTFA on Shared Prompts and H200
The ranking compares streaming capabilities using Time to First Audio (TTFA), which measures the wait time before a sound is playable, a criterion highlighted for voice agents and other interactive uses. For models with a streaming API, TTFA is the delay until the first audio chunk arrives, while for models without streaming, it corresponds to the complete generation of the utterance. All models are evaluated in batch size 1 on the same 50 English prompts from CV3-Eval, in their default voice and on identical hardware; the first three runs serve as warm-up and are excluded, after which the median TTFA is reported. The default view presents results on an H200 GPU, and CPU measurements are available for a limited but growing set. kyutai/pocket-tts is noted for its strong performance in streaming on both GPU and CPU. Evaluation scripts are expected to be opened soon, with contributions via GitHub.
Objective Metrics for Intelligibility, Speed, and Timbre
The dashboard relies on objective metrics covering various facets. Intelligibility is measured via WER and CER between the text and the transcription of the generated audio, using Qwen3 ASR presented as the best open-source model in an ASR ranking. Speed combines RTFx for offline batch inference on H200 GPU and TTFA for latency in streaming on both GPU and CPU H200. Speaker similarity is based on cosine similarity between WavLM embeddings of a generated sample and a reference clip. According to the project promoters, these metrics reduce evaluation time from weeks to hours. The system does not claim to replace human judgments: WER and speaker similarity do not assess naturalness, expressiveness, or preference. However, the numerical results can assist voting-based rankings in selecting models to include.
Hexgrad, Supertone, and fishaudio Lead According to Macro-Averaged WER
By default, models are ranked based on a macro-averaged WER on English, combining Seed TTS Eval and CV3 Eval in zero-shot. On this criterion, hexgrad/Kokoro-82M, Supertone/supertonic-3, and fishaudio/s2-pro come out on top, and Pareto charts show trade-offs between WER, RTFx, and size. English performance does not always generalize, allowing for multiple languages to be activated in the ranking. Seed TTS Eval only covers English and Chinese; for other languages, scores come from CV3 Eval in zero-shot. For Chinese, Japanese, and Korean, the reported measure is CER, and the intercultural average is a macro-average. On the multilingual side, k2-fsa/OmniVoice, fishaudio/s2-pro, and FunAudioLLM/Fun-CosyVoice3-0.5B-2512 are noted as strong performers. By enabling voice cloning, the table adds a SIM column and two additional Pareto charts to visualize trade-offs with batch inference and size. In this mode, some models like bosonai/higgs-tts-3-4b and openbmb/VoxCPM2 see their average WER improve when a reference audio is provided.
A Tab to Listen and Compare Voice Syntheses
A Listen tab allows users to hear the outputs that underpin the metrics to concretely compare models. The user selects language or dataset, decides whether to include voice cloning, and can choose models or let a random selection guide them. This space is presented as filling a gap in existing rankings by offering exploration of a wide range of outputs. Feedback can be left on the audios; if enough community votes are collected, this data could be integrated into the ranking. To reduce spam and bots, users are asked to log in with a Hugging Face account. The organizers claim they want a system shaped by the community and are seeking suggestions on datasets, models, and metrics to include.
Why a New Ranking and What It Focuses On
TTS evaluation is described as fragmented, even as more than 8000 models are listed as of September 30, 2026, on the Hub. Comparison arenas, based on user-evaluated duels then rated in Elo via models like Bradley–Terry, are accused of failing to keep pace with releases. As of September 30, 2026, among the 92 models listed by Artificial Analysis, only 16 have open weights, with a similar bias noted on Voice Arena. Several practical reasons are put forward: adding via an API key proves easy, while an open model requires hosting and serving by the operator, and commercial providers are more motivated to achieve a ranking. There is no assurance regarding the stability of voters over time, and their preferences are likely to change. In response, the Open TTS Leaderboard has been built around open-source models that are less visible in the arenas and emphasizes multilingual evaluation, as English performance is not considered a good substitute for other languages.
Brief IA — L'actualité IA en français
L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.