Brief IA

MIT: AI in Healthcare Must Adapt to Users

🔬 Research·Tom Levy·

MIT: AI in Healthcare Must Adapt to Users

MIT: AI in Healthcare Must Adapt to Users
Key Takeaways
1A study from MIT reveals that the explainability of AI in healthcare varies by user, influencing diagnostics.
2Non-experts improve their accuracy with AI but risk blindly following incorrect recommendations.
3Clinicians prefer predictions without explanations, demonstrating resilience to AI errors.
💡Why it mattersAdapting AI interfaces to user expertise could reduce diagnostic errors and enhance trust in these systems.
Le brief IA que lisent les pros

Le brief IA que les pros lisent chaque soir

Les 7 actus IA du jour, décryptées en 5 min. Gratuit.

Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.

Choisis ton rythme

Gratuit · Pas de spam · Désabonnement en 1 clic

📄
Full Analysis

The Impact of AI on Medical Diagnostics According to User Expertise

Researchers from MIT, in collaboration with other institutions, conducted an in-depth study on the use of explainable artificial intelligence (AI) tools in the healthcare sector. This research highlighted that the results produced by these tools can vary significantly depending on the user, whether they are non-experts or healthcare professionals. The study revealed that non-experts, when assisted by AI, improved their diagnostic accuracy. However, this improvement was primarily due to their tendency to rely blindly on AI models. In contrast, primary care providers showed a preference for AI predictions without explanations, which led to better outcomes.

The MIT Study and Its Implications

Published in the journal Nature Medicine, the study focused on dermatological diagnosis, a field where AI is already being used to support some clinicians and is beginning to be accessible to patients through AI-based research products. Marzyeh Ghassemi, an associate professor in the Department of Electrical Engineering and Computer Science at MIT, emphasized the importance of these findings for the design of AI interfaces in healthcare. She stated that while AI systems can enhance performance in certain contexts, it is crucial to find a balance with algorithmic deference, which can lead to additional errors. Ghassemi also noted that AI and explainability methods can induce automation bias in users, an effect that must be considered when designing these systems.

The Influence of Interfaces on Diagnosis

Explainable AI aims to provide users with elements to evaluate a model's output. For example, a system may highlight areas of a medical image that influenced its diagnosis or show similar images to support a prediction. Large language models (LLMs) offer another approach by producing a clear language account of a model's reasoning, presenting a diagnosis in terms accessible to a general audience. The MIT study tested several of these approaches by presenting participants with medical images accompanied by predictions of skin diseases.

Non-experts assessed whether images of moles showed cancer, while clinicians faced a broader task: they had to provide a differential diagnosis for dermatological diseases.

Non-Experts and Their Dependence on Linguistic Explanations

The results showed that each explainability approach tested improved non-experts' accuracy, particularly in identifying non-cancerous moles. However, the researchers also tested a model designed to address bias against darker skin tones, which improved accuracy while reducing diagnostic disparities based on skin tone. Despite these improvements, non-experts exhibited a strong dependence on the model's recommendations, leading to decreased performance when the model was incorrect. The explanations provided by LLMs had the strongest effect on this deference, with participants trusting these explanations, whether they were correct or not. Users who received LLM assistance reported greater confidence in incorrect responses.

Roxana Daneshjou, an assistant professor of biomedical data science and dermatology at Stanford University, stated that patients with limited medical knowledge are the most vulnerable to erroneous outputs from explainable AI.

Primary Care Providers and Their Use of AI

In contrast to non-experts, clinicians demonstrated resilience against erroneous recommendations or explanations from AI. Their best performance was observed with a more limited interface, where the system provided only the model's prediction without explanation. LLM explanations had a lesser impact on improving clinicians' accuracy. This does not mean that explanations play no role in clinical practice, but rather that an explanation format suitable for a patient or novice may not be appropriate for a trained user.

Orson Xu, the lead author and an assistant professor in the Department of Biomedical Informatics at Columbia University, stated, “It all depends on how each group uses the explanation. A clinician already has a diagnosis in mind and checks the AI against their own training, so a poor explanation is detected.”

Timing and Automation Bias

The study also explored when users received AI assistance. The results showed that users became more deferential when the system presented an explanation before they had the opportunity to make their own diagnosis. This finding underscores the importance of interface design, which could include a request for the user's initial diagnostic hypothesis, followed by an AI recommendation. The users most dependent on AI were also those who had the most difficulty completing the task without assistance, highlighting the need for support tailored to their level of expertise. The research compared human and AI performance across different presentations of diseases. AI systems outperformed humans when symptoms appeared subtly. Humans performed better when the image contained atypical symptoms or unrelated features. Tools for clinicians may require a direct model output that supports review against professional judgment. Tools intended for patients require particular attention regarding LLM explanations, especially when the system presents a confident narrative for an incorrect recommendation.

Brief IA — L'actualité IA en français

L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.