OpenAI restricts ChatGPT's health advice to premium subscribers

Le brief IA que les pros lisent chaque soir
Les 7 actus IA du jour, décryptées en 5 min. Gratuit.
Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.
Choisis ton rythme
Gratuit · Pas de spam · Désabonnement en 1 clic
OpenAI Limits Health Advice from ChatGPT to Premium Subscribers
ChatGPT will provide lower-quality health advice if you do not pay.
OpenAI is rolling out its "Health in ChatGPT" feature for U.S. users aged 18 and older, allowing them to connect Apple Health, medical records, and wellness apps to review lab results, prepare for appointments, and analyze health data.
Free users receive lower-quality health advice powered by the GPT-5.5 Instant model, while paying subscribers have access to the more advanced GPT-5.6 Sol model.
Despite over 300 million people asking health questions to ChatGPT each week, the risks remain significant: AI chatbots have shown that they provide incorrect medical results with high confidence rather than admitting their uncertainty.
OpenAI is launching "Health in ChatGPT" after unveiling and testing the feature since January. Paying for a subscription could literally save your life.
Users can connect Apple Health, medical records, and wellness apps to review lab results, prepare for medical appointments, and analyze sleep or activity data. OpenAI claims it will not use connected health data for model training or advertising.
Free Users Get Inferior Health Advice
Free users of ChatGPT receive lower-quality health advice. OpenAI powers this feature with GPT-5.5 Instant, which performs worse on health benchmarks than the new flagship model, GPT-5.6 Sol, reserved for paying users. OpenAI will likely defend this two-tier system on ethical grounds by emphasizing that both models outperform doctors' responses on the HealthBench Professional test.
On OpenAI's HealthBench Professional, GPT-5.6 Sol surpasses responses written by doctors and the older models GPT-4o and GPT-5.5 Instant across all categories. The largest gaps are in completeness, at 88.0% versus 53.2%, and the usefulness of health decisions, at 83.0% versus 50.8%.
Even when benchmark results seem decisive, they come from artificial testing environments designed to measure knowledge, and doctors may score lower for several reasons. They may be under time pressure, suffering from fatigue, or taking the test without tools such as patient records or colleagues' opinions.
Benchmarks also cannot capture much of what happens during a real medical examination, where doctors can assess patients in person, pick up on non-verbal cues, and rely on years of experience to evaluate their overall condition. OpenAI itself repeatedly states in the announcement that ChatGPT can still make mistakes and cannot replace medical advice. The company claims that over 260 doctors helped develop the health features.
OpenAI's early tests also revealed that more than 70% of participants asked health questions outside the dedicated health section because transitioning to it was too cumbersome. The company has since made health available in any conversation while maintaining the separate health section to manage data and access previous health exchanges.
A Better Dr. Google, Perhaps, but Not a Substitute for a Doctor
OpenAI indicates that over 300 million people are now asking health questions to ChatGPT each week, up from 230 million in January. The company has not specified if or when the Health feature will be available in Europe. When OpenAI announced it in January, it specifically excluded the European Economic Area, Switzerland, and the United Kingdom. Stricter EU data privacy rules and the possibility that the feature could be classified as high-risk under the EU AI Act are likely reasons.
These risks are not hypothetical. In the latest radiology benchmark, RadLE 2.0, none of the 16 AI models tested performed as well as human radiologists. The main issue was that chatbots provided incorrect results with high confidence instead of admitting when they had reached their limits, while human radiologists were much better at recognizing uncertainty. This mix of overconfidence, persuasion, and flattery, where chatbots validate users rather than challenge false assumptions, has also contributed to serious mental health harm.
At the same time, some reports show that AI detects patterns in health data that medical professionals miss, sometimes with striking results. MIRA, a system for electronic health records, and AMIE both performed roughly as well as primary care physicians during simulated consultations. One researcher compared these AI agents to an airplane's autopilot: "These systems can support and relieve medical professionals by taking on routine tasks, but ultimate responsibility will always remain with the doctors."
Brief IA — L'actualité IA en français
L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.