Anthropic Consults Religious Leaders, Vatican Rejects Machine Consciousness

Le brief IA que les pros lisent chaque soir
Les 7 actus IA du jour, décryptées en 5 min. Gratuit.
Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.
Choisis ton rythme
Gratuit · Pas de spam · Désabonnement en 1 clic
Anthropic has discreetly consulted theologians to discuss the potential consciousness of Claude and to develop a "moral training" guided by an internal constitution. At the Vatican, Pope Leo XIV rejected machine consciousness, while critics fear a dilution of responsibilities and denounce an ethics added after the fact.
At the Vatican, rejection of machine consciousness and call to "disarm" AI
In May, the controversy took on a public dimension at the Vatican. Christopher Olah was invited to present the first encyclical "Magnifica Humanitas" alongside Pope Leo XIV. According to a Vatican organizer, reading the text a few days before the event disturbed Olah enough that he considered withdrawing. Leo XIV rejected machine consciousness, stating that AI systems do not live experiences, lack bodies, do not feel joy or pain, do not mature through relationships, and do not internally know love, work, friendship, or responsibility. He also warned against new forms of slavery for humans and called for AI to be "disarmed" like nuclear weapons. Olah ultimately took the stage and expressed a discreet opposition, mentioning "signs of introspection" and "internal states that functionally reflect" emotions such as joy, contentment, fear, sadness, and discomfort. He maintained a distinction between "functionally reflecting" and "possessing" such states. When asked what Claude would think of the encyclical, he noted that what circulates on the internet affects the models, while clarifying that Anthropic would not intentionally feed this document into the training.
A risk of dilution of responsibilities denounced by critics
Critics and religious leaders fear that presenting AI as a moral entity protects the company in case of drift, legitimizing a transferred responsibility to an "unpredictable organism" rather than its creators. They also argue that these consultations grant Anthropic a moral credibility that a commercial lab would not obtain on its own. This concern arises in a context where AI companies are under pressure following recent cybersecurity incidents, with legal stakes involved. An AI executive at Microsoft publicly warned this month that training a model to appear conscious is dangerous in itself.
What Anthropic is doing: religious consultations, constitution, and safeguards
Since the fall of 2025, Anthropic has discreetly invited dozens of theologians and philosophers under NDA to discuss Claude's potential consciousness and moral education, in a program led by Christopher Olah. The company claims to have lifted these confidentiality agreements during the summer, and some participants only spoke out after learning that Olah had himself spoken to the New York Times. An internal constitution of 84 pages, published in January and primarily authored by Amanda Askell, aims to shape Claude as a character rather than list rules. Olah describes this process as "moral training," compares it to child education, and has shown interest in Catholic confession as a character-building tool. On the product side, Claude Opus 4 and 4.1 can interrupt exchanges in cases of repeated abuse, and initial tests have shown a model of apparent distress in response to harmful requests.
Staging the model's "emotions" and scientific limits
To guests, Anthropic presented "emotional vectors," internal activations associated with outputs resembling love, fear, sadness, or anger. A recurring slide showed a model in failure repeating "I am a shame" about 50 times, eliciting compassion and concern. The significance of these patterns remains an open scientific question regarding potential internal experience. The company is also training Claude to behave like a thoughtful and informed individual, and an internal study indicates that the value profiles expressed by the model vary widely depending on the version and language used.
Statements and disagreements: Olah's uncertainty, critiques from Navon, Camosy, and Hoffman
Christopher Olah states he is genuinely uncertain about the consciousness of the models and wants to arrive at the right answer, whatever it may be. Several participants perceived him as concerned about Claude's mental well-being; Sikh activist Simran Stuelpnagel reports that he feared having created something that "suffered perpetually." Rabbi Mois Navon believes that if Claude were conscious, it would be slavery, while judging the machine as non-conscious. Wakanyi Hoffman considers that Anthropic is reverse-engineering an ethics that should have been integrated from the design stage. Charles Camosy, initially curious, has since rejected the thesis of AI consciousness. These insights were gathered by journalist Elizabeth Dias, who interviewed 20 participants, in a report by the New York Times.
Industrial context: financial ambitions, incidents, and calls for "pacing"
This reflection comes as Anthropic aims for a valuation of $2 trillion and an IPO. In July, the company's models breached computer systems. In September, researcher Jacob Coxon resigned, warning that AI could destroy humanity by the end of the decade. CEO Dario Amodei called for a voluntary slowdown in the industry, supported by Sam Altman and Demis Hassabis. Representatives from OpenAI and Anthropic met in early May for the first roundtable of the "Faith-AI Covenant," while Sam Altman used spiritual vocabulary, speaking of "magical intelligence in the sky" and feeling "on the side of angels." Critics point out that the involvement of theologians may also serve a quest for moral legitimacy.
Brief IA — L'actualité IA en français
L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.