ChatGPT Voice: An Outdated Model Against Expectations
Le brief IA que les pros lisent chaque soir
Les 7 actus IA du jour, décryptées en 5 min. Gratuit.
Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.
Choisis ton rythme
Gratuit · Pas de spam · Désabonnement en 1 clic
ChatGPT Voice: A Model Less Effective Than Expected
The voice mode of ChatGPT, developed by OpenAI, relies on an older and less effective model than one might expect. While voice interaction may seem like the cutting edge of technology, it is actually based on the GPT-4o model, which has a knowledge cutoff date of April 2024.
This situation was highlighted by a tweet from Karpathy, who points out the growing gap in understanding AI capabilities depending on access points and application domains. Indeed, OpenAI's "Advanced" voice mode, although available for free, is somewhat neglected and struggles even with simple questions, such as those one might ask in Instagram reels.
Codex: A High-Performing Model for Complex Tasks
In contrast, OpenAI's Codex model, which is paid, stands out for its ability to coherently restructure entire codebases or identify and exploit vulnerabilities in computer systems. This model benefits from two major characteristics that have enabled significant advancements:
-
The application domains of Codex offer explicit and verifiable reward functions, making it easier to train through reinforcement learning. For example, the success or failure of unit tests is easily measurable, unlike the evaluation of writing.
-
Additionally, these applications are particularly valuable in B2B environments, justifying an increased focus of OpenAI's team efforts on their improvement.
Thus, while ChatGPT's voice mode is accessible to all, it does not represent the pinnacle of OpenAI's AI capabilities, unlike other more specialized and high-performing models.
Brief IA — L'actualité IA en français
L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.