Brief IA

Anthropic and OpenAI: AI Safeguards Hinder Cybersecurity

🤖 Models & LLM·Tom Levy·

Anthropic and OpenAI: AI Safeguards Hinder Cybersecurity

Anthropic and OpenAI: AI Safeguards Hinder Cybersecurity
Key Takeaways
1Restrictions on AI models from Anthropic, such as Mythos, limit cybersecurity researchers.
2The verification programs from OpenAI and Anthropic aim to control access but are criticized by experts.
3Researchers are turning to open-source models to bypass the limitations imposed by AI giants.
💡Why it mattersCurrent safeguards risk slowing down cybersecurity efforts, exposing systems to increased threats.
Le brief IA que lisent les pros

Le brief IA que les pros lisent chaque soir

Les 7 actus IA du jour, décryptées en 5 min. Gratuit.

Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.

Choisis ton rythme

Gratuit · Pas de spam · Désabonnement en 1 clic

📄
Full Analysis

AI Giants and Their Restrictive Safeguards

For several months, major companies in the artificial intelligence sector have implemented strict security measures to prevent the malicious use of their models by hackers. However, these same restrictions are now causing problems for cybersecurity researchers, who find their ability to defend networks limited. AI giants like Anthropic have established safeguards to protect their models from malicious uses, but these measures also affect those working to enhance system security.

In June, the U.S. government imposed export restrictions on Anthropic's AI models, including Mythos and Fable. This decision was motivated by a report indicating that the safeguards of these models could be circumvented, allowing their use for malicious cyberattacks. This measure was taken to prevent potentially dangerous uses of these advanced technologies, but it has also had the effect of restricting access for legitimate researchers who seek to use these tools to improve security.

Anthropic has often presented Mythos as a powerful cyber machine, requiring use by carefully vetted users under strict security conditions. While export controls on Fable 5 and Mythos 5 have been lifted, Fable 5 became publicly accessible again on July 1, while Mythos 5 remained limited to verified U.S. organizations. This selective approach aims to ensure that only those who have been thoroughly evaluated can access these powerful AI tools, while maintaining a high level of security.

Verification Programs and Criticism

Anthropic and OpenAI have established verification programs for cybersecurity researchers, allowing them to access models with fewer restrictions. OpenAI offers the "Trusted Access for Cyber" program, while Anthropic provides the "Cyber Verification Program." These initiatives aim to provide controlled access to researchers who can prove their legitimacy and intent to work for the benefit of cybersecurity.

These initiatives have faced significant criticism, particularly from researchers whose work involves identifying vulnerabilities in systems before cybercriminals can exploit them. In a recent podcast, Mark Dowd, a recognized security researcher, expressed his discomfort with the arbitrary decisions made by large companies regarding what is considered safe in terms of security. Dowd, who has spent decades discovering and selling zero days—previously unknown software flaws and the exploits that take advantage of them—to Western governments, rather than reporting them to software manufacturers for correction, emphasized that these restrictions can hinder the crucial work of researchers seeking to uncover flaws before they are exploited by malicious actors.

Challenges for Offensive Cybersecurity Researchers

Chris Anley, chief scientist at NCC Group, explained the importance of an AI model testing the exploitation of a bug to confirm its severity. However, if a safeguard prevents the model from responding, it undermines the efforts of defenders. Anley highlighted that the ability to use AI to simulate attacks is essential for identifying and fixing vulnerabilities before they are exploited by hackers. He stated that "fix this code" as a prompt is both an essential mechanism for defense and a roadmap for finding critical vulnerabilities in the codebase.

Anley compared this tool to a hammer, indispensable for building a house but potentially dangerous. When researchers encounter obstacles, they turn to open-source models without safeguards. These models offer a flexibility that restricted commercial models cannot provide, allowing researchers to freely test attack and defense scenarios.

Paolo Stagno, CTO of CrowdFense, also criticized AI companies for treating customers like children needing supervision. He specified that his team uses cutting-edge models solely for reverse engineering, avoiding the use of AI to identify vulnerabilities to prevent the risk of disclosing sensitive data. Stagno emphasized that using locally run open-source models allows for complete control over sensitive data, thus avoiding the risks associated with sharing data with cloud-based models.

The Impact of Safeguards on Research

Giuseppe Cali, a security researcher, claimed that safeguards do not hinder his work, as he does not use AI for offensive tasks. He prefers to rely on AI for reverse engineering and tool creation while maintaining control over the discovery and exploitation of bugs. Cali explained that AI can accelerate the code analysis process, but he prefers to keep control over identifying and exploiting vulnerabilities.

An anonymous researcher from a smartphone component manufacturer stated that since his company is not included in Anthropic's CVP program, the available tools are of little use for identifying vulnerabilities due to overly strict restrictions. This researcher expressed frustration over the inability to fully utilize AI tools to enhance the security of his company's products.

Open-Source Models as an Alternative

Chris Thompson, CEO of RemoteThreat and founder of Offensive AI Con, pointed out that the safeguards of AI models are often inconsistent, even within verified programs. This situation forces researchers to spend more time circumventing these limitations than focusing on security. Thompson explained that this inconsistency can be frustrating for researchers looking to use AI to improve security. He stated that safeguards can be inconsistent and operate differently each day, even within the looser confines of Anthropic and OpenAI's verified programs.

As a result, many researchers are turning to Chinese open-source models, such as GLM, which can be used locally without restrictions. Thompson warned that these safeguards are pushing responsible researchers away from U.S.-regulated systems toward foreign systems, which could be more harmful than beneficial. He called for a revision of safeguard policies to allow for more responsible and effective access to AI tools.

Thompson urged for greater openness in AI lab programs, allowing for responsible access while holding accountable those who abuse the tools. Without this, defenders risk losing the AI race against an unprecedented wave of attacks. He emphasized the importance of finding a balance between security and accessibility to enable researchers to continue protecting systems against emerging threats.

Brief IA — L'actualité IA en français

L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.