Real World VoiceEQ: Voice AI Facing Humanity's Challenge

Le brief IA que les pros lisent chaque soir
Les 7 actus IA du jour, décryptées en 5 min. Gratuit.
Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.
Choisis ton rythme
Gratuit · Pas de spam · Désabonnement en 1 clic
Voice, the New Interface of AI
Current evaluations of voice artificial intelligence performance often place it at a level close to that of humans. However, when examining real interactions, a different reality emerges. Voice is becoming the preferred interface for interacting with AI. Whether in customer support, healthcare, education, entertainment, or personal assistants, speech is gradually replacing text as the primary means of communication.
The advancements made in recent years in voice models are undeniable. Word error rates have significantly decreased, latency has reached levels allowing for smooth conversations, and many benchmarks now seem saturated. Yet, for those who regularly use voice AI, it is evident that something is still missing.
Current voice models can give the impression of changing personality during a conversation, lack natural pauses or hesitations, and struggle to handle accents, background noise, or emotions in the voice. These aspects are often overlooked in benchmarks that focus on latency and word error rate. Users are looking for systems that can truly listen, respond appropriately, and maintain a natural and reliable interaction.
Towards a More Comprehensive Benchmark
To evaluate these essential qualities, the Real World VoiceEQ benchmark has been created. It aims to measure the human quality of vocal interactions. This benchmark assesses the ability of voice systems to recognize, produce, and respond to acoustic information that simple transcriptions cannot capture, such as tone, emotion, speaker identity, and surrounding context.
Real World VoiceEQ analyzes over 40 voice models, whether proprietary or open-source, across more than 15 key evaluation dimensions and over 60 metrics. These evaluations cover areas such as Automatic Speech Recognition (ASR), Text-to-Speech (TTS), Speech-to-Speech (S2S), and Speech Understanding.
The development of Real World VoiceEQ is based on over 1 million individual human evaluations, collected from diverse demographics, speech styles, and acoustic environments. The current benchmark includes 785,000 TTS evaluations and 48,000 STS evaluations, making it one of the largest human evaluations of voice AI to date.
Each evaluation was conducted using Kairos, a flexible evaluation platform specifically designed for voice. This infrastructure allows AI labs and companies to conduct customized assessments tailored to specific use cases, identify precise failures in voice systems in production, generate human preference data, and continuously improve models through reinforcement learning and user feedback.
The Revelatory Results of Real World VoiceEQ
-
Increasing Specialization of Voice Systems. Advances in voice AI are moving towards greater specialization. Rather than seeking a universally superior voice model, current systems optimize specific capabilities such as technical accuracy, emotional understanding, conversational intelligence, expressiveness, and robustness. A high-performing model for repeating reservation numbers or complex names may struggle to express emotions convincingly.
-
Listening, a Challenge for Voice Models. Speech-to-Speech models show the greatest variation among all evaluated categories. Some systems excel in recognizing emotions but struggle to respond naturally. Access to audio does not guarantee that agents utilize the available paralinguistic information. Some systems remain focused on transcription, neglecting cues such as tone, rhythm, hesitation, emphasis, and volume.
-
Traditional Benchmarks Outpaced by Reality. Many established benchmarks are reaching their limits and do not reflect real-world conditions. Models continue to face difficulties with accented speech, overlapping speakers, emotion, background noise, and prolonged conversations. In our evaluation, performance varied significantly between leading open-source and proprietary models, contrary to what traditional benchmarks suggest.
-
The Importance of Human Evaluation. Preliminary research shows that some models can be optimized for established public benchmarks. Several reproduce known errors in reference transcriptions, follow arbitrary spelling conventions, and even reconstruct masked words absent from the audio.
The Urgency for a New Measure for Voice AI
As voice establishes itself as one of the key interfaces of AI, speed and technical accuracy will no longer suffice to determine the success of systems. The models that will be adopted by the public will be those capable of understanding, expressing, and responding like humans, not only under ideal benchmark conditions but also through the complexity of real conversations.
For decades, voice AI has progressed by optimizing according to quantitative metrics on standardized benchmarks, from WER for transcription accuracy to objective perceptual metrics like PESQ and DNSMOS for speech quality. Real World VoiceEQ hopes to expand this paradigm by providing a human-centered metric to evaluate the components of synthetic vocal interactions.
Brief IA — L'actualité IA en français
L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.