Anthropic Faces Internal Tensions: Its AI Models Suspended
Le brief IA que les pros lisent chaque soir
Les 7 actus IA du jour, décryptées en 5 min. Gratuit.
Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.
Choisis ton rythme
Gratuit · Pas de spam · Désabonnement en 1 clic
Anthropic's artificial intelligence models have been taken offline following internal conflicts, marked by the statement "They betrayed us," according to information shared by sources close to the company and the U.S. administration. These tensions have emerged in the context of export controls imposed by the U.S. government on the Mythos and Fable models developed by Anthropic.
Logan Graham, head of the Frontier Red team at Anthropic, Dave Orr, responsible for security measures, and Nicholas Carlini, an influential blogger, are expected in Washington D.C. for a meeting with the Department of Commerce. This meeting could be crucial for the future of Anthropic's models as the company seeks to navigate a complex regulatory landscape.
Discussions surrounding the security of Anthropic's models focus on resistance to circumvention. While the company claims that no universal circumvention has been identified for Claude Mythos, the U.S. government has reacted to a potential circumvention, described as "non-universal" by Anthropic. A source close to the administration suggested that an adjustment in attitude may be necessary for everyone to feel safe, protected, and happy.
Anthropic has also been working on "Constitutional Classifiers" to enhance the security of its models. However, whether these efforts will be sufficient to meet U.S. requirements remains an open question. Additionally, the 2023 document on Universal and Transferable Adversarial Attacks on Aligned Language Models raises concerns about Anthropic's ability to manage these types of attacks. The company must now ensure that its models cannot be circumvented, although achieving perfection in security is challenging.
Brief IA — L'actualité IA en français
L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.