Anthropic Restricts Fable 5: AI Facing Security Challenges

Le brief IA que les pros lisent chaque soir
Les 7 actus IA du jour, décryptées en 5 min. Gratuit.
Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.
Choisis ton rythme
Gratuit · Pas de spam · Désabonnement en 1 clic
Anthropic recently introduced an artificial intelligence model, Claude Fable 5, which belongs to the Mythos class and incorporates extensive safety measures. This model is designed to be accessible to the general public, but it comes with advanced protections that can be triggered by seemingly innocuous requests. When a request is flagged by the security classifiers, the Mythos-class model is either blocked or downgraded to Opus 4.8, an earlier version, to provide a response.
Anthropic explained that these safety measures are essential for making the model accessible to the general public. If you ask a simple question about cybersecurity or biology to this new model, you might find that it falls short. This is because the Mythos-class model is so powerful that it requires broad protections that can, by mistake, flag harmless requests.
After some users reported triggering the security response with basic questions about cancer or safety, Business Insider decided to conduct a test. By asking Fable 5 simple questions about cancer, such as how misinformation about cancer spreads online and breaking down the different types of cancer, the model quickly switched from Fable 5 to Opus 4.8, informing the user of the change before responding.
The displayed message stated: "Fable 5 has safety measures that flag messages on most cybersecurity or biology topics. They can also flag safe and normal content. These measures allow us to bring you Mythos-level capabilities in other areas more quickly, and we are working to refine them."
Anthropic launched Fable 5 on Tuesday, claiming it is as powerful as its Mythos 5 model but with additional protections. This launch came two months after the company stated that Mythos was too powerful for broad release due to cybersecurity concerns. Instead of being made public, Mythos was only accessible to a small group as part of a cybersecurity project.
Anthropic clarified that safety measures were necessary to make the model accessible to the general public. "With the launch of Claude Fable 5, our first model in the Mythos class, we believe that models now have a greater capacity to perform real scientific tasks and that malicious actors could potentially use our models for very risky biological research," a spokesperson for Anthropic said in a statement to Business Insider. "We have always used classifiers to prevent our models from assisting with requests related to biological weapons. To deploy Fable 5 safely, we believe it was necessary to be excessively conservative with our safety measures so that they block most requests related to biological work."
The company indicated that there are three categories of requests that could be flagged by its security classifiers: cybersecurity, biology and chemistry, and distillation of Fable 5's capabilities. When the safety measure is triggered, Fable 5 will either be blocked from responding or the model will revert to Opus 4.8 before responding, depending on the user's preference.
Anthropic stated that it has been conservative with the safety measures and plans to improve them. In its announcement, Anthropic mentioned that the safety measures could result in the flagging of safe and normal content, but their initial data showed that over 95% of Fable sessions did not revert to Opus. "To make the model both safe and fast, we have set these safety measures conservatively," Anthropic said, adding that it is working to improve the measures to reduce false positives.
"We intend to make Mythos-class models available without these safety measures to the broader community of biology and life sciences so that these capabilities can be used to accelerate biomedical research and drug discovery," the spokesperson for Anthropic stated.
The launch came about a week after researchers at Anthropic stated that AI is progressing so rapidly that leading laboratories may need to slow down or take a temporary pause for the company to keep up. David Kasten, head of policy at Palisade Research, said it is "very clear" from Anthropic's public statements that the company is concerned about the risks posed by increasingly powerful models.
While he views these safety measures as a good faith attempt by Anthropic to mitigate risks, he emphasized that historically, "people eventually find a way to circumvent security restrictions." "It's always a bit of a cat-and-mouse game between the attacker and the defender," he added, noting that there is always some risk associated with the release of the more powerful model.
He also mentioned that the fact that Anthropic's most powerful model frequently reverts to a less capable model could create a gap in public understanding of the increasing power of AI models. "This gap in understanding could be really dangerous in leading policymakers, or the public in general, to not fully grasp the risks these models pose in terms of the capabilities they offer."
Brief IA — L'actualité IA en français
L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.