Claude Opus 5: A Brilliant Yet Frustrating AI Model to Master

Le brief IA que les pros lisent chaque soir
Les 7 actus IA du jour, décryptées en 5 min. Gratuit.
Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.
Choisis ton rythme
Gratuit · Pas de spam · Désabonnement en 1 clic
Use of AI: Review of Claude Opus 5 and Browser Use in Codex
Use of AI in Codex (5 Concrete Examples)
The use of artificial intelligence in software testing has demonstrated a capacity to surpass traditional human methods. Claire, for instance, observed that when she tests her own integration flow, she instinctively follows the most direct path. In contrast, Codex, the AI tool, explored the workflow from both a team and an individual perspective, scrutinizing edge cases of mandatory fields. This approach uncovered a critical bug that had gone unnoticed for months.
Boundary models, when used with a certain freedom of thought, can offer superior performance. Claire initially provided a list of 25 items to test but later simplified her request by asking the AI to "QA the integration flow." This simplification allowed the model to determine the best approach to take.
Persona testing, a method where AI simulates a product's use by a specific user, can reveal frictions that traditional user research does not detect. The idea from her husband, EJ, to have the AI use the product as a particular user highlighted a structural issue in the reference flow between threads.
Claire also found that she was initially using far more resources than necessary on LinkedIn. By adopting a medium-effort model, she was able to manage unread messages and draft contextual responses without requiring an official API connection.
Finally, Codex enabled Claire to perform complex tasks like opening firewall ports and configuring SSH on her Mac Mini remotely, which would have been nearly impossible to do manually from a hotel room.
Model and Level of Effort
It is crucial to tailor the model and level of effort to the task at hand. For example, sorting LinkedIn messages requires a very different type of reasoning than that needed for writing production code. Choosing the appropriate amount of resources can accelerate these workflows and reduce costs, especially when browser use extends throughout the day.
Human Intervention
In certain situations, human intervention remains necessary. For instance, during an online shopping session, Claire had to step in to complete a CAPTCHA before she could regain control. This division of labor, where AI handles navigation and filtering while humans manage moments requiring identity verification, proves effective.
From Zero to Coding to Hardware Hacker: How Cursor and a Raspberry Pi Make AI Fun
Transition to Hardware
Maddie Reese transitioned from having no coding experience to creating functional hardware projects such as a Twitter pager and an AI-powered receipt printer. She shares how using Cursor and a Raspberry Pi facilitated the transformation of her ideas into tangible projects.
Key Points
Maddie emphasizes that mastering coding is not necessary to create interesting projects. She compares her coding skills to knowing just enough Spanish to get by, allowing her to read parts of the code and spot errors without having to write complete functions.
Her creation process follows a simple sequence: she starts with an idea, conducts an interview, creates a shopping list, makes purchases, and then moves on to building. Cursor helps her define technical requirements before she spends any money.
For Maddie, the architecture of a project does not need to be elegant. AI handles the coding, allowing her to focus on the overall functionality of the project.
She has also developed a personal API containing useful information, such as her coffee orders and favorite restaurants. This could enable an agent to automatically book a restaurant when she is in San Francisco.
Maddie prefers to work in a clean Cursor chat environment rather than a full terminal during the ideation phase, which helps her concentrate on planning.
Review of Claude Opus 5: This Model is Brilliant (but Frustrating)
Model Evaluation
Claire tested Claude Opus 5 against six other leading AI models, and it took first place. She explains that while it produces some of the best work she has seen, it remains one of the most frustrating models to use.
Key Points
The AI industry may be entering a phase where new models emerge every week, with benchmark scores continually rising. However, Opus 5 stands out for its sometimes timid behavior and reliance on human approval, often requiring Claire to tell it to "just do it."
A notable issue is the "Claude slop," where although the model produces writing intended for humans, it is often too long and laden with unnecessary commentary.
Despite these flaws, Opus 5 finished first in Claire's blind benchmark, with an overall score of 78, just ahead of Claude Sonnet 5 at 77 and GPT-5.6 Sol at 76.
Brief IA — L'actualité IA en français
L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.