ElevenLabs v4: Long-Lasting Consistency and Turbo at 150 ms on Sale

Le brief IA que les pros lisent chaque soir
Les 7 actus IA du jour, décryptées en 5 min. Gratuit.
Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.
Choisis ton rythme
Gratuit · Pas de spam · Désabonnement en 1 clic
ElevenLabs launches v4 and its Turbo variant for voice synthesis. The publisher highlights improvements in expressiveness, pronunciation, and long-duration coherence, while aiming for real-time latency. Promotional pricing and data localization options accompany the general release.
v4 surpasses Cartesia and Gemini in a third-party ranking
In the Provider Voice Arena ranking by Artificial Analysis, Eleven v4 ranks ahead of Cartesia Sonic 3.6 and Google’s Gemini 3.8 Flash TTS. It achieves a score of 91.7% on a pronunciation benchmark, compared to 85.6% for v3. In blind tests conducted by ElevenLabs, about three-quarters of listeners preferred v4 over models from Cartesia, Inworld, and Google. In these same tests, v4 was deemed more expressive in 65 to 81% of comparisons depending on the competitor. Regarding latency, ElevenLabs reports that Turbo starts speaking in 150 milliseconds, compared to 262 milliseconds for Cartesia Sonic 3.6, and 814 milliseconds for OpenAI's GPT-4o mini TTS. According to ElevenLabs' internal benchmarks, v4 Turbo initiates speech much faster than the other tested models.
Capabilities: 10,000 characters, 90+ languages, and revived cloning
With version 4, ElevenLabs announces a more faithful execution of instructions as well as a more stable voice during the generation of long content. This iteration now covers over 90 languages and offers an improvement in voice cloning. It allows processing of up to 10,000 characters per request and provides more precise management of laughter, whispers, and sound effects compared to v3, which already included these tags but with less accuracy. The company describes an architecture analyzing tone, rhythm, and context, with possible instructions via tags, simple phrases, and phonetic spelling, and pronunciation controls touted as more reliable. The voices of narrators and characters are expected to remain consistent even after several regenerations of lines. According to ElevenLabs, 10,000 characters correspond to about ten minutes of audio, and long works are assembled in segments that must maintain a constant rhythm. In dialogue, AI speakers take into account the scene context. Compared to the approximately 70 languages in v3, v4 now exceeds 90. Clones are expected to speak other languages with native accents without temporal drift. Professional clones are once again available, and the Instant Voice Clone requires only ten seconds of audio.
Turbo targets live use with a new architecture
The v4 Turbo variant targets voice agents and responds, according to ElevenLabs, in 150 milliseconds, a value claimed to be superior to the competition. The publisher indicates that it has designed a new architecture to enable real-time use, observing a speech start around 150 milliseconds. This variant is intended for customer service calls or game characters, with the goal of no longer having to choose between speed and expressiveness. ElevenLabs specifies that it has optimized Turbo in connection with its ElevenAgents platform.
Licensed voices and monetization for actors
ElevenLabs presents high-quality voice clones as suitable for dubbing, including the use of the same actor's voice across all supported languages. The company holds licenses for the voices of the individuals who record them. Voice actors can offer a well-trained clone of their voice in the ElevenLabs library and receive revenue when it is used by paying clients. A dedicated marketplace already provides access to celebrity voices, such as that of Michael Caine.
Prices drop until October 12 and immediate access
The API is priced at $80 per million characters for v4 and $40 for Turbo, with a temporary discount to $22 and $11 until October 12. Subscribers to the Creator plan at $22 per month or more can use v4 in ElevenCreative at no additional cost for two weeks, up to double their monthly credits, according to ElevenLabs. For comparison, Artificial Analysis lists Sonic 3.6 at $49 per million characters and Gemini 3.8 Flash TTS at $16.49. The v4 and v4 Turbo models are now available in ElevenAgents, ElevenCreative, and via the API.
Data residency and another launch in generative audio
By default, ElevenLabs stores customer data in the United States. Companies can opt for isolated environments in the EU, India, or Singapore, with the possibility that some processing occurs outside the chosen region. In the EU, a data retention-free mode allows API processing to remain within the region. Meanwhile, the publisher launched Music 2.5 in mid-September, a model designed to produce denser songs with a more natural rendering.
Brief IA — L'actualité IA en français
L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.