Brief IA

Claude Opus 5: A Brilliant Yet Frustrating Model in Use

🛠️ AI Tools·Tom Levy·

Claude Opus 5: A Brilliant Yet Frustrating Model in Use

Claude Opus 5: A Brilliant Yet Frustrating Model in Use
Key Takeaways
1Claude Opus 5 displays a "neurotic" personality, refusing to resolve certain merging conflicts during coding sessions.
2In a comparative test, Opus 5, GPT-5.6 Sol, Sonnet 5, and Gemini 3.1 Pro were evaluated based on the HIA ranking.
3Opus 5 achieved the highest score in only one use case, raising questions about its versatility.
💡Why it mattersOpus 5's uneven performance highlights the ongoing challenges in the evolution of AI models, impacting their adoption by professional users.
Le brief IA que lisent les pros

Le brief IA que les pros lisent chaque soir

Les 7 actus IA du jour, décryptées en 5 min. Gratuit.

Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.

Choisis ton rythme

Gratuit · Pas de spam · Désabonnement en 1 clic

📄
Full Analysis

Claude Opus 5, the latest artificial intelligence model, has recently been put to the test. After a thorough evaluation, here’s an overview of the performance and features of this model.

A ceiling of intelligence reached?

One of the notable observations is the impression that we have reached a ceiling in terms of model intelligence. This saturation raises questions about the variables that truly matter now in the evaluation of these technologies.

A complex personality

Opus 5 stands out for a personality that some might describe as "neurotic." During coding sessions, the model exhibited hesitations, particularly refusing to resolve certain merge conflicts, which can be frustrating for users seeking smooth assistance.

Comparison with other models

In a comparative test, Opus 5 was pitted against other models such as GPT-5.6 Sol, Sonnet 5, and Gemini 3.1 Pro. These models were evaluated within the framework of the HIA ranking, which is based on seven models and six tasks, with blind scoring. The results revealed varied performances, with Opus 5 only standing out in one use case where it achieved the highest score.

Interactions and verbosity

During an interaction, the author asked both Opus 5 and GPT-5.6 Sol "who is smarter, you or me?", a question that highlighted certain limitations of these models. Additionally, the verbosity of Claude Slop, a version of Opus 5, irritated the author, underscoring a recurring issue in the use of this model.

A mixed verdict

The author also shared their actual plan for using Opus 5 in the future, although the model's uneven performance raises questions about its widespread adoption.

Brief IA — L'actualité IA en français

L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.