Brief IA

AI Chatbots: Over Half of Medical Diagnoses Fail

🔬 Research·Tom Levy·

AI Chatbots: Over Half of Medical Diagnoses Fail

AI Chatbots: Over Half of Medical Diagnoses Fail
Key Takeaways
1A study from Nature Medicine reveals that AI chatbots correctly identify medical conditions in less than 34.5% of cases.
2Large language models, such as ChatGPT and Llama 3, often fail to provide accurate diagnoses due to insufficient initial information.
3Despite their popularity, chatbots should not be used as a primary source of medical advice, especially in serious cases.
💡Why it mattersThe growing reliance on chatbots for medical advice could lead to potentially dangerous diagnostic errors.
Le brief IA que lisent les pros

Le brief IA que les pros lisent chaque soir

Les 7 actus IA du jour, décryptées en 5 min. Gratuit.

Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.

Choisis ton rythme

Gratuit · Pas de spam · Désabonnement en 1 clic

📄
Full Analysis

AI Chatbots and Their Limitations in Medical Diagnostics

A recent study published in the scientific journal Nature Medicine highlights the limitations of chatbots and large language models (LLMs) when it comes to providing medical advice. This research was conducted in the UK with 1,298 participants who used LLMs such as ChatGPT and Llama 3 from Meta to seek medical guidance. The results are concerning: these tools correctly identified medical conditions in only 34.5% of cases. The study concludes that chatbots and large language models should not be the first source of medical advice.

LLM Performance in the Study

The study emphasizes that, although LLMs now achieve scores on medical knowledge benchmarks comparable to those required to pass the United States Medical Licensing Examination, their practical effectiveness remains limited. Clinical documents generated by these models are deemed equivalent, if not superior, to those written by physicians. However, a major issue was identified: users often do not provide enough initial information, compromising the accuracy of diagnoses. In 16 out of 30 sampled interactions, the initial messages contained partial information. In two cases, the LLMs provided initially correct responses but then added incorrect answers after users supplied additional details. This indicates that engaging further with chatbots does not necessarily improve the likelihood of obtaining a correct diagnosis. After the initial diagnosis, the LLMs provided the correct follow-up steps only 44.2% of the time.

Use of Chatbots for Medical Advice

According to a survey conducted by OpenAI, the owner of ChatGPT, 3 out of 5 American adults use AI for health-related questions. They utilize these tools to gather information when feeling unwell, to prepare for their clinician visits, and to better understand doctors' instructions and recommendations. Although there is a warning on the ChatGPT site stating that "ChatGPT may make mistakes. Verify important information," many people nonetheless take the chatbot's responses as facts. The study reminds us that ChatGPT and similar chatbots should not be considered reliable sources for medical advice, especially in serious situations.

Brief IA — L'actualité IA en français

L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.