Brief IA

1Password: AI Fails to Fix 74% of Software Vulnerabilities

🤖 Models & LLM·Tom Levy·

1Password: AI Fails to Fix 74% of Software Vulnerabilities

1Password: AI Fails to Fix 74% of Software Vulnerabilities
Key Takeaways
1A study by 1Password reveals that AI is only able to fix 26% of software vulnerabilities.
2Large language models generated 6,080 patch attempts, often ineffective.
3AI-generated patches sometimes introduce new bugs or alter the behavior of applications.
💡Why it mattersAI, while useful for identifying vulnerabilities, remains inadequate for effectively fixing them, necessitating increased human oversight.
Le brief IA que lisent les pros

Le brief IA que les pros lisent chaque soir

Les 7 actus IA du jour, décryptées en 5 min. Gratuit.

Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.

Choisis ton rythme

Gratuit · Pas de spam · Désabonnement en 1 clic

📄
Full Analysis

AI and Software Vulnerability Patching: An Unresolved Challenge

A recent study conducted by 1Password highlights the current limitations of artificial intelligence in the field of cybersecurity. Large language models (LLMs), often praised for their advanced capabilities, prove to be ineffective when it comes to creating patches for security vulnerabilities. According to this research, only 26% of the patches generated by AI are usable, underscoring that these technologies are not yet ready for large-scale use in vulnerability remediation.

Key Findings of the Study

The study, conducted by 1Password's security research team, Off-By-1-Labs, tested the capabilities of LLMs by asking them to generate patches for recently discovered complex vulnerabilities. The researchers initially hoped that AI could achieve a success rate of 67% in creating patches. However, the results fell far short of expectations, with performance being "significantly lower and uneven."

Hypothesis and Results

To test the capabilities of AI models, the team selected six vulnerabilities from open-source software. The research paper, titled "Vulnerability Patches from Advanced Models are Often F.L.A.W.E.D. Artifacts with Built-in Flaws," aimed to provide insight into the current capabilities of LLMs rather than a direct comparison among them.

Vulnerabilities Studied

The vulnerabilities analyzed include:

  • CVE-2026-31431: Linux Privilege Escalation (Copy Fail)
  • CVE-2026-34197: Remote Code Execution ActiveMQ
  • CVE-2026-8512: Use After Free in Chrome's File System Access API on macOS
  • CVE-2026-45185: Unauthenticated Remote Code Execution EXIM
  • CVE-2026-22738: Remote Code Execution SpringAI SpEL
  • GHSA-wpqr-6v78-jr5g: Remote Code Execution Gemini CLI

The models were tasked with generating patches for each of these vulnerabilities, resulting in a total of 6,080 attempts, or about 3,040 per vulnerability, under varying environmental conditions and with nine different prompts per bug.

Results of AI-Generated Patches

The results show that AI successfully produced appropriate patches in only 26% of cases. Additionally, 21% of the patches, while correcting the bug, altered the application's behavior. In 53.9% of cases, the LLMs failed to create an effective patch, sometimes introducing new bugs.

Reasons for AI Model Failures

The failure of AI models to generate effective patches is primarily due to the production of "Fix-Like" artifacts with built-in flaws. These patches, while appearing functional on the surface, do not fully resolve the vulnerabilities and may introduce fragile security mechanisms or new bugs, sometimes altering the normal behavior of applications.

Future Perspectives

To address these shortcomings, 1Password has made its FLAWED tools available on GitHub, allowing researchers to continue their own studies. Although AI is not yet ready to handle patching autonomously, it can still play a crucial role in identifying and triaging vulnerabilities. Keith Hoodlet, head of Off-by-1 Labs, emphasizes the importance of collaboration between human defenders and AI tools to determine the most critical bugs to fix. Human oversight remains essential for making informed decisions about which patches to apply and assessing the associated business risks.

Brief IA — L'actualité IA en français

L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.