Gemini 3.1 Flash TTS: A Voice Revolution with Innovative Audio Tags
Le brief IA que les pros lisent chaque soir
Les 7 actus IA du jour, décryptées en 5 min. Gratuit.
Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.
Choisis ton rythme
Gratuit · Pas de spam · Désabonnement en 1 clic
Gemini 3.1 Flash TTS: A Step Forward in Voice Synthesis
Google today unveiled Gemini 3.1 Flash TTS, a major advancement in the field of voice synthesis. This new model promises enhanced quality and expressiveness, providing developers, businesses, and everyday users the ability to create next-generation voice applications. Starting today, Gemini 3.1 Flash TTS is available in preview via the Gemini API and Google AI Studio for developers, as well as on Vertex AI for businesses. Workspace users can also benefit from it through Google Vids.
Enhanced Performance and Control
The Gemini 3.1 Flash TTS model stands out with a notable improvement in speech quality, making it more natural and expressive than its predecessors. It achieved an Elo score of 1,211 on the Artificial Analysis TTS ranking, a benchmark that evaluates human preferences. This model is positioned in the "most attractive quadrant" due to its balance between high speech quality and reduced cost. It supports over 70 languages and offers detailed creative control through natural language. Additionally, it features native multi-speaker dialogue, allowing for smoother and more natural interactions between multiple characters.
Introduction of Audio Tags for Enhanced Expressiveness
One of the major innovations of Gemini 3.1 Flash TTS is the introduction of audio tags. These tools allow for intuitive control over vocal style, pacing, and delivery of speech. By integrating natural language commands into the text, users can refine the voice output with unprecedented precision. Developers can experiment with these tags and other features in Google AI Studio, with configurable controls that place them in the role of director.
Advanced Features for Developers
- Scene Direction: Allows for defining the environment and providing specific dialogue instructions, helping characters interact naturally.
- Speaker-Specificity: With unique Audio Profiles and Director's Notes, developers can adjust the pacing, tone, and accent of characters.
- Seamless Export: Settings can be exported as Gemini API code, ensuring vocal consistency across various projects.
Global Reach with Over 70 Languages
Gemini 3.1 Flash TTS is designed to operate on a global scale, offering high-fidelity voice synthesis in over 70 languages. The optimizations allow for advanced control over style, pacing, and accent, facilitating the creation of localized and expressive voice experiences. Early users, whether developers or businesses, are already highlighting the positive impact of these innovations, particularly thanks to the audio tags that transform text into a rich vocal performance.
Enhanced Security with SynthID
All audio content generated by Gemini 3.1 Flash TTS is embedded with a SynthID watermark. This imperceptible watermark, integrated directly into the audio, reliably detects AI-generated content, thus contributing to the prevention of misinformation.
Brief IA — L'actualité IA en français
L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.