Brief IA

Claude Fable 5 and ChatGPT: AI in the Face of UX Evaluation

🛠️ AI Tools·Tom Levy·

Claude Fable 5 and ChatGPT: AI in the Face of UX Evaluation

Claude Fable 5 and ChatGPT: AI in the Face of UX Evaluation
Key Takeaways
1AI can assist in UX evaluations, but it does not replace human expertise, especially for complex tasks.
2Claude Fable 5 demonstrated better accuracy than ChatGPT-5.6 Sol, but both produced analyses with errors.
3A 2025 study revealed that GPT-4o identifies 21% of usability issues, highlighting the importance of human expertise.
💡Why it mattersAI is a powerful tool for accelerating UX work, but human expertise remains essential for accurate and contextual results.
Le brief IA que lisent les pros

Le brief IA que les pros lisent chaque soir

Les 7 actus IA du jour, décryptées en 5 min. Gratuit.

Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.

Choisis ton rythme

Gratuit · Pas de spam · Désabonnement en 1 clic

📄
Full Analysis

AI and UX Evaluation: A Promising but Imperfect Duo

Artificial intelligence (AI) is increasingly establishing itself as a powerful tool in the field of user experience (UX) evaluation. However, it cannot yet fully replace human expertise. The accuracy of the results obtained largely depends on the quality of the information provided to the AI, rather than the intrinsic capabilities of the model. In a recent project, I was tasked with conducting a UX evaluation for the search, results, and product detail pages of a portal. To do this, I had access to a set of questions about the target audience, objectives, and USP, as well as a market research document and access to Google Analytics and Clarity.

The widely held initial hypothesis is that AI is already sufficiently capable of performing this type of task. The idea is that the designer only needs to verify the result, and on topics like this, AI could almost replace humans. Faced with this opportunity, I decided to test this hypothesis to see if it really held up.

A Rigorous Test with Two AI Models

To conduct this experiment, I used two AI models in parallel: Claude Fable 5 and ChatGPT-5.6 Sol, both accessible via paid subscriptions. I provided each of them with the entire set of available data: the brief, questions and answers, research material, exports from Google Analytics, heatmaps from Clarity, attention and scrolling maps, as well as the necessary URLs. The task was clearly defined: to produce an expert UX evaluation document with a specific objective.

It is crucial to clarify that I did not use a single prompt for this task. I structured the work in several stages, asking the models to justify their conclusions with data from Google Analytics and Clarity whenever possible. The goal was not to test an "all-in-one" approach, but to see what a structured workflow could produce.

At the same time, I also conducted the analysis manually to have a basis for comparison.

Results of the Experiment: Analyzes in Need of Improvement

It is important to emphasize that this report is based on a personal experience and not on a controlled study. The results obtained by both AI models were filled with irrelevant or even non-existent issues. However, some useful observations emerged, although they were only actionable at the level of formulation and structure. Claude Fable 5 proved to be more precise than ChatGPT-5.6 Sol, at least in terms of formulation and structure. This remains a subjective impression and not an objective measure.

A notable difference between the two models is that Claude could produce the result in a single document, while ChatGPT could only do so partially. This seems more related to the output format than to the quality of the analysis.

It is true that some problems I had not noticed were highlighted by the AI, which proved to be useful. However, the AI also failed to detect several other issues.

These observations align with a 2025 study that compared GPT-4o to Nielsen's usability heuristics. The study found that GPT-4o identified about 21% of the usability problems spotted by experts while raising several new issues, sometimes incorrectly. AI is thus an excellent tool for a first pass, but it does not yet replace expert evaluation.

Final Assembly and the Role of AI

In the end, I chose to compile the analysis with the help of Claude. I integrated my own observations into an ongoing conversation with the model, which allowed us to determine what to keep or discard. Claude revised the text for grammar, coherence, and other similar aspects, suggesting modifications that proved to be very helpful.

Although I did not measure the time precisely, I estimate that if I had done all the work myself, it would have taken about 15% more time.

Thus, AI proves to be an excellent companion and a helpful partner that allows me to accelerate and refine my work. However, it cannot yet replace me. Together, we form a more effective duo than ever.

Future Perspectives for AI in UX

One might wonder if a guided browser visit would have improved the results. Today, there are tools that allow models to operate on the live site in a browser, rather than relying solely on static data and URLs. I did not test this approach, but I suspect it would not have led to better results, as the main limitation lies in judgment, that is, understanding context and objectives, weighing severity, and distinguishing between real problems and noise. A live guided visit does not fundamentally change these aspects.

For the future, I plan to develop my own agent capability for this task, integrating established usability heuristics, severity weighting shaped by professional experience, and the expert perspective that a UX practitioner brings to an interface. The goal is not to let AI decide for me, but to transform my own expert judgment into a reusable and coherent form.

This aligns with what is observed in research on the subject: AI heuristic evaluations at 95% accuracy from Baymard were not achieved through better prompts, but by anchoring the model in their own sought-after UX knowledge base. Certainly, in their case, it involves usability heuristics, which is simpler than a complex and contextual evaluation like this one, but the principle remains the same. Accuracy stems from the expert framework you provide to the model. That is why it is wise to encode your own standard into a capability.

The Future of AI in UX Design

I am aware that AI will eventually replace our work; that is a certainty. The only question is when this will happen.

I have read articles suggesting a certain slowdown or plateau in AI development. According to a late 2024 Reuters article, several leading AI researchers, including Ilya Sutskever, believe that the scaling of models has reached its limits. Future advancements will come from new approaches, such as "o1" type models that reason step by step, rather than from sheer size. In other words, development continues, but the focus is on smarter methods rather than raw size.

Of course, other developments and research continue in the background, so there is no doubt that AI will continue to progress in the long term.

I recently discovered a page suggesting that we still have a few good years ahead of us, particularly in complex research and design, where using various AI models can accelerate, refine, and support work.

Tips for Tomorrow

  • Start: Use AI as a partner for a first pass, feed it with your real data, work in stages, and encode your own heuristics and severity into a reusable prompt or capability.

  • Stop: Do not deliver the AI output as is, do not expect a single prompt to replace a workflow, and do not delegate your judgment. Severity, prioritization, and business context remain your responsibility.

And do not panic at the thought that AI will take your job tomorrow. It will not happen, but designers who learn to guide it will surpass those who do not.

Brief IA — L'actualité IA en français

L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.