AI and Human Judgment: A Delicate Balance
Le brief IA que les pros lisent chaque soir
Les 7 actus IA du jour, décryptées en 5 min. Gratuit.
Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.
Choisis ton rythme
Gratuit · Pas de spam · Désabonnement en 1 clic
AI and Human Judgment: A Precarious Balance
The Often Overlooked Fundamental Principle
In the development of artificial intelligence technologies, a fundamental principle is often overlooked: AI should serve as a support for decision-making, leaving the human as the final decision-maker. This principle is based on the idea that AI can provide valuable insights, identify risks, or patterns that a human alone might miss. However, it is up to the human to take these elements into account to make an informed decision. Unfortunately, in critical sectors like law, healthcare, or criminal justice, this principle is often violated from the product design stage.
AI systems are often designed to be fast and appealing, with a human approval step added superficially. This creates an illusion of decision support, while in reality, it shifts the responsibility to the human without equipping them to fully assume it. AI gets the credit for the outcomes, while the human bears the risk of potential errors.
The gap between these two elements is where people get hurt, are sanctioned, and sued.
The Requirements for True Decision Support
For AI to genuinely play its role as a decision support tool, it must enhance human judgment by highlighting elements that a human alone could not see. This implies that the human must be able to add their own context, judgment, and experience to make a decision they can defend. This requires the user to be genuinely able to evaluate the information provided by AI, rather than simply passively approving a recommendation.
However, many AI products are not designed to facilitate this critical evaluation. Often, the design is limited to a high-performing model and a clear user interface, with a simple review step that does not allow for true critical interaction with the data provided by AI.
This is not what most products are designed to produce. I have observed how AI-assisted decision-making workflows are constructed, and the pattern is consistent: the product team builds a good model, the UX team creates a clear interface for the output, somewhere in the specifications, there is a review step, usually a button, sometimes a confirmation window. The legal department approves because a human technically approves each action. The product is launched.
What is not designed: what the human actually needs to evaluate what they are looking at.
The Dangers of Automation Bias
Research has shown that when AI recommendations appear on screen, users tend to trust them, especially under pressure. This phenomenon, known as automation bias, leads users to accept the outputs of a system without questioning them. A 2023 study tested this bias with 457 clinicians across 13 states. The results showed that standard AI predictions improved diagnostic accuracy by 4.4 percentage points. However, when AI made biased predictions, diagnostic accuracy decreased.
Explanations, such as image-based saliency maps, did not help. When the model was incorrect, showing clinicians the reasoning behind that erroneous response did not protect them from agreeing with it. The explanation failed because the human was no longer engaged in evaluating the case. They were assessing whether they should trust the system. These are different cognitive tasks, and the design had already addressed the latter by the time the explanation appeared.
A 2023 study titled "Putting a Human in the Loop: Increased Adoption but Decreased Accuracy" found something worse: introducing a human reviewer actually increased the frequency with which people followed the AI's recommendation, as the presence of a human led participants to believe that the decision had already been validated. The human was not correcting errors. The human was providing cover.
Zana Buçinca and Krzysztof Gajos at Harvard tested whether forcing users to form their own opinion before seeing the AI's response would change this. It worked, reducing the frequency with which people followed the AI even when it was incorrect. However, participants did not appreciate these designs, as they forced them to think more, which is often perceived as undesirable friction. Yet, in high-stakes decisions, the absence of friction is often the mode of failure.
Concrete Examples of Breaching the Contract
In the Legal Field
In 2023, a New York lawyer, Steven Schwartz, cited fictitious cases generated by ChatGPT in a brief. The court imposed a $5,000 fine and required Schwartz to inform each judge mentioned in those fictitious cases. This case is not isolated: by the end of 2025, over 1,356 incidents of AI hallucinations in legal documents have been recorded, with fines reaching up to $30,000.
In California, at least one court has begun suggesting that the opposing lawyer might have a duty to detect the AI-generated fakes from the other side. The lawyer was the decision-maker. The design provided them with no means to verify the input they were relying on.
In the Healthcare Sector
IBM Watson for Oncology spent over $62 million without successfully integrating into hospital systems. An audit revealed that its recommendations were not based on current evidence. Similarly, Epic's sepsis prediction tool failed to correctly detect the majority of cases, missing 67% of sepsis cases at the recommended threshold, generating unnecessary and ineffective alerts.
Hundreds of clinicians were nominally the decision-makers on sepsis while relying on a tool whose real-world performance was unchallengeable.
In Autonomous Systems
In 2025, Tesla was found partially responsible for a fatal accident involving its Autopilot system. The jury determined that the way the system was marketed encouraged drivers to place excessive trust in it, contributing to the accident.
For years, the standard response to "who is responsible when Autopilot is engaged?" was "the driver; the system requires constant supervision." The jury reacted. When the design of a product leads a human to relinquish a judgment they were supposed to maintain, the product shares the outcome.
In Criminal Justice
Eric Loomis was convicted partly based on a risk score generated by COMPAS, a tool whose algorithms are opaque. The Wisconsin Supreme Court allowed the use of COMPAS, despite documented biases in its predictions, including a tendency to overestimate the risk of recidivism among Black defendants. ProPublica's 2016 investigation revealed that COMPAS was nearly twice as likely to falsely flag Black defendants as future offenders compared to white defendants at the same risk level.
The judge was the decision-maker. The design placed an unexamined number on the page and then assumed that the human could reason around it.
The Reasons for the Persistence of the Problem
Automation bias is an immediate explanation for the failures, but another more insidious issue is the deterioration of human skills. In aviation, excessive reliance on autopilot has led to a decline in manual piloting skills, as demonstrated by the Asiana Flight 214 accident in 2013 at San Francisco International Airport. This incident prompted the FAA to issue guidelines requiring pilots to manually fly more often during low-workload phases to maintain their skills.
In the medical field, a similar experiment is underway. Trials show that using AI to detect polyps during colonoscopies can affect doctors' ability to make accurate diagnoses without assistance.
In conclusion, the design of AI systems must evolve to truly support human decision-making by providing the necessary tools to critically evaluate and interpret AI recommendations.
Brief IA — L'actualité IA en français
L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.