ChatGPT Revolutionizes Voice: Listen and Speak Simultaneously

Le brief IA que les pros lisent chaque soir
Les 7 actus IA du jour, décryptées en 5 min. Gratuit.
Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.
Choisis ton rythme
Gratuit · Pas de spam · Désabonnement en 1 clic
OpenAI Unveils Advanced Voice Models for ChatGPT
OpenAI has recently updated its artificial intelligence voice technology to better align with its text capabilities. Two new models, named GPT-Live-1 and its smaller version, are now available to all ChatGPT users. These models leverage the latest technological advancements from OpenAI, allowing users to choose from three levels of intelligence based on the complexity of the desired responses. Atty Eleti, product lead for voice at ChatGPT, emphasized that the AI's voice is now more natural and conversational.
Continuous Interaction: A Major Advancement
One of the most significant developments is the introduction of a process called continuous interaction. This mechanism allows the artificial intelligence to receive information and produce responses simultaneously. In the context of voice, this means that the AI can "listen" and "speak" at the same time, unlike previous versions that waited for the end of your input to respond. This feature is particularly useful for live translations, where you can speak in one language and receive an instant translation in another, with minimal delay.
Task Delegation for Increased Efficiency
The new architecture also allows for the delegation of certain tasks to OpenAI's advanced models. For example, when GPT-Live encounters a question requiring deep thought, it can transfer that task to GPT-5.5, which is capable of handling multiple tasks in parallel. Kundan Kumar, research lead for GPT-Live, explained that this capability allows the AI to remain in conversation with the user while performing complex calculations in the background. Once the task is completed, the response is seamlessly integrated into the conversation, mimicking human interactions.
Visual Responses and Broader Availability
OpenAI has also integrated newer AI design technology to generate visual responses when relevant. This includes elements such as weather reports or sports scores, which are sometimes better presented visually than verbally. These new voice models are available to all ChatGPT users, whether on mobile or web, and regardless of whether they are paid subscribers or free users. OpenAI notes that the full deployment may take a few days. Additionally, the launch of the GPT-5.6 series is scheduled for Thursday, marking another step in the evolution of AI capabilities.
A More Human Voice, but Potential Risks
A notable feature of this update is the addition of filler words, such as "uh" or "um," to make the AI's voice closer to that of humans. Furthermore, the AI can be interrupted with less delay and can remain silent until activated by a wake word. Unlike other voice assistants like Siri or Alexa, ChatGPT stays attentive until it is manually turned off.
Enhanced Safety Measures
OpenAI has implemented measures to protect user privacy. Audio clips are retained for 30 days to provide context for conversations but can be deleted upon request. Additionally, OpenAI automatically opts you out of AI training with voice mode, ensuring a more secure and privacy-respecting use. However, the increased humanization of the AI raises concerns. Anthropomorphism can have negative consequences, particularly for individuals suffering from mental health issues. OpenAI assures that the new models come with expanded safety measures and outperform previous versions in critical areas such as self-harm, psychosis, violence, and sexual content.
Brief IA — L'actualité IA en français
L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.