Brief IA

Google DeepMind: Medical AI Surpasses GPT-5.4

🤖 Models & LLM·Tom Levy·

Google DeepMind: Medical AI Surpasses GPT-5.4

Google DeepMind: Medical AI Surpasses GPT-5.4
Key Takeaways
1Google DeepMind's co-clinical AI has outperformed GPT-5.4 in blind tests but still lags behind experienced doctors.
2In tests on 98 medical queries, the AI was preferred over other tools with a score of 67 to 26 against an existing tool.
3In the RxQA benchmark, the AI achieved 73.3% correct answers, surpassing GPT-5.4, but doctors remain better for critical tasks.
💡Why it mattersThe co-clinical AI promises to assist doctors, but it cannot yet replace their expertise in critical situations.
Le brief IA que lisent les pros

Le brief IA que les pros lisent chaque soir

Les 7 actus IA du jour, décryptées en 5 min. Gratuit.

Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.

Choisis ton rythme

Gratuit · Pas de spam · Désabonnement en 1 clic

📄
Full Analysis

Google Deepmind Develops Promising Co-Clinical AI

Google Deepmind has developed a co-clinical AI designed to assist doctors in diagnosing and treating patients during daily medical care. This AI has been tested in blind evaluations using realistic scenarios from general medicine. Physicians rated the responses provided by this system as superior to those from other AI tools, including GPT-5.4. However, despite these promising results, a simulation study revealed that experienced doctors still outperformed the AI, particularly in identifying critical warning signs and conducting physical examinations.

The co-clinical AI is based on the concept of "triadic care," where AI agents assist patients while leaving clinical authority and oversight to physicians. The goal is to create an AI system that functions as a member of the medical team, supporting patients under the supervision of a clinician. To evaluate this system from a clinician's perspective, the Google Deepmind team collaborated with academic physicians to adapt the NOHARM framework, checking for two types of errors: commission errors and omission errors.

Performance Comparison with Other Tools

In a blind comparison using 98 realistic primary care queries, physicians consistently preferred the responses of the co-clinical AI over those of an existing clinical AI tool, scoring 67 to 26 against that tool and 63 to 30 against GPT-5.4. However, the system recorded a critical error in one of the 98 cases studied. The gap was even more pronounced on medication-related questions. The RxQA benchmark covers 600 questions on active ingredients, interactions, and dosages, drawn from national drug directories in two countries and verified by licensed pharmacists. These questions are challenging for primary care physicians: with reference books, they achieved 61.3% correct answers, and only 48.3% without.

The co-clinical AI scored 73.3%, just ahead of GPT-5.4 at 72.7%. The gap widened when questions were posed openly rather than in multiple-choice format, as physicians typically do in their work. In this case, the co-clinical AI achieved a quality score of 95.0%, compared to 90.9% for OpenAI's model.

Multimodal Telemedicine: A New Dimension

Beyond text-based support, Google Deepmind is testing how the co-clinical AI handles audio and video in real-time for telemedicine. In collaboration with physicians from Harvard and Stanford, the team conducted a randomized simulation study with 20 synthetic clinical scenarios, 10 physicians acting as patient actors, and a total of 120 telemedicine visits.

The co-clinical AI demonstrated capabilities that go beyond what purely text-based systems can achieve. It corrected a patient's inhalation technique and guided patients through shoulder examinations to detect a rotator cuff injury. For conversations with patients, the co-clinical AI operates on a two-agent system: a "Planner" module monitors the conversation to ensure that the "Speaker" agent remains within safe clinical boundaries. When physicians use the system, it prioritizes strong clinical evidence and performs checks and citations during searches.

Experienced Doctors Maintain Their Superiority

The study evaluated over 140 aspects of consultation quality across seven domains: triage, history taking, clinical reasoning, communication and counseling, treatment steps, detection of warning signs, and physical examinations. The conclusion is discouraging for anyone hoping that AI could replace a doctor: experienced physicians outperform the AI overall, particularly in detecting "warning signs" and guiding critical physical examinations.

However, the co-clinical AI matched or surpassed primary care physicians in 68 of the 140 areas assessed. OpenAI's GPT-realtime lagged behind in all seven domains. Researchers conclude that systems like this work best as support tools for physicians, rather than as replacements for clinical judgment.

It remains uncertain whether this research project will transform into a real product. The results show progress in AI-driven evidence synthesis and telemedicine consultations, but they also highlight that there is still a gap to bridge with experienced physicians, particularly on critical safety tasks like detecting warning signs. "While it is still early, the promise is clear," says Deepmind researcher Alan Karthikesalingam.

Brief IA — L'actualité IA en français

L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.