Generative AI: Instilling Doubt Rather Than Certainty

Le brief IA que les pros lisent chaque soir
Les 7 actus IA du jour, décryptées en 5 min. Gratuit.
Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.
Choisis ton rythme
Gratuit · Pas de spam · Désabonnement en 1 clic
Studies highlight a widespread flaw in language models: responding with confidence that is disconnected from actual reliability. Researchers describe concrete levers to correct this: making reasoning explicit, allowing for doubt, and, above all, permitting abstention. This discipline does not come alone; it is enforced through a software "harness," execution rules, and selective prediction.
Saying When We Don't Know and Allowing Silence
Models must choose an answer before knowing if it is correct, and a confident tone does not indicate the direction of their assumption. The default behavior of off-the-shelf systems is to respond with confidence. To address this, teams impose a real confidence threshold: below this, no guessing; the AI stops, clarifies its weaknesses, and asks for more details. This framework aligns with the selective prediction described by Ran El-Yaniv and Yair Wiener, where a model remains silent in cases of uncertainty instead of responding at all costs. When the AI answers everything, users fall into an automation bias, validating falsely confident outputs. Research by Sunnie S. Y. Kim, Q. Vera Liao, Mihaela Vorvoreanu, Stephanie Ballard, and Jennifer Vaughan shows that a simple "I'm not sure, but..." leads users to accept less blindly and to spot more errors. However, Haochen Guo and Petr Polak emphasize that transparency only guides trust with an explicit admission of uncertainty; otherwise, it adds reading without guidance. The proposed approach is to make the operation visible and to express doubt when it exists.
What Tests Show: Caution Evaporates When AI Becomes Useful
Hiroshi Okumura observed that models hold back their judgments in controlled academic settings but almost never when asked for practical advice. In one test, one out of every two hundred responses retained some form of caution; the sought-after utility prevails, and the reserve disappears when the stakes rise. Miao Xiong and his colleagues find that when asked about their own confidence, these models almost always convey excessive confidence, closer to a tone of certainty than a measure. Left to their own devices, machines do not implement transparency and doubt; cutting-edge models produce assured responses because they are designed to provide plausible answers and move the user forward. An excerpt shared by Casey Newton and Tommy Vietor, from the Instagram account husk.irl, illustrates this automatism: in response to the question about the month containing an X, a chatbot successively delivered "December," then "October"—before spelling it out while noting the absence of X—then "February"; three confident answers, all incorrect. The intention is not to lie but to respond: having an answer and being right are different. In practice, a testimony reports that these models almost always return "something": sometimes correct, sometimes close, sometimes incomprehensible in their reasoning; the output arrives anyway, and the question of its legitimacy is deferred to the user. The cultural example of Cliff Clavin in Cheers, confidently asserting that humans only use 17% of their brains and not a full 64%, reminds us how confidence can mask error.
The "Harness": Execution Rules and Safeguards Above the Model
Rather than modifying the model, engineers interpose a "harness": a structure that frames execution through rules, tools, checkpoints, and safeguards. Traian Rebedea and his team at NVIDIA have developed a toolkit that sits between the user and the model, imposing in real-time what the system is allowed to say or do, without touching the underlying model. The model continues to reason; the harness defines the standard of good behavior and maintains it, aiming to establish accountability, preserve power, and prevent a façade of certainty from misleading the user. The analogy of a parachutist trusting their equipment at ten thousand feet illustrates this logic of restraint. In this framework, showing the work and admitting doubt are not spontaneous impulses; they are imposed behaviors. Forcing the expression of gray areas through rules prevents the machine from smoothing them over; simply asking to "be honest" costs more tokens and results in the same confident assumption delivered in a more humble tone. Honesty must therefore be integrated and enforced by the structure to truly exist.
Making Reasoning Visible Without Overwhelming the User
The "confidence calibration" described by Sunnie S. Y. Kim relies on what the AI shows of its process: showing nothing only provides a veneer of confidence, poorly directed. Hence approaches where reasoning is constructed in view, to be challenged and interrupted if it strays, in the spirit of what we expect from a human who lays out their work and distinguishes the certain from the uncertain. But "more" does not equal "better": Herman Saksono, Vivien Morris, Andrea G. Parker, and Krzysztof Z. Gajos have shown that explanations that are difficult to understand reinforce dependence on AI. Design must therefore target the elements that are truly useful for decision-making and clarify the assumptions. Conversely, some products intentionally obscure the process in the name of an appearance of intelligence, with generic screens like "Analyzing your request" and responses from a black box. The user sees neither the reasoning nor the potential errors; trusting what one cannot see hinders correction. In this context, reminding that the best thing an AI can learn to say is "I'm not sure" reflects the importance of uncertainty as a usage signal, in the face of systems sometimes designed to appear self-assured.
Step-by-Step Reasoning: Correcting Better Than a Global Verdict
Making each step of reasoning visible serves not only transparency but also allows for teaching the system by correcting precisely where logic falters. Jonathan Uesato and his colleagues have shown that evaluating and correcting steps, rather than just the final answer, makes reasoning much more reliable. A black box offers only a thumbs up or down, while a visible process provides the red pen for fine-tuning.
Brief IA — L'actualité IA en français
L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.