1Password: AI Fails to Fix 74% of Software Vulnerabilities

Le brief IA que les pros lisent chaque soir
Les 7 actus IA du jour, décryptées en 5 min. Gratuit.
Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.
Choisis ton rythme
Gratuit · Pas de spam · Désabonnement en 1 clic
AI and Software Vulnerability Patching: An Unresolved Challenge
A recent study conducted by 1Password highlights the current limitations of artificial intelligence in the field of cybersecurity. Large language models (LLMs), often praised for their advanced capabilities, prove to be ineffective when it comes to creating patches for security vulnerabilities. According to this research, only 26% of the patches generated by AI are usable, underscoring that these technologies are not yet ready for large-scale use in vulnerability remediation.
Key Findings of the Study
The study, conducted by 1Password's security research team, Off-By-1-Labs, tested the capabilities of LLMs by asking them to generate patches for recently discovered complex vulnerabilities. The researchers initially hoped that AI could achieve a success rate of 67% in creating patches. However, the results fell far short of expectations, with performance being "significantly lower and uneven."
Hypothesis and Results
To test the capabilities of AI models, the team selected six vulnerabilities from open-source software. The research paper, titled "Vulnerability Patches from Advanced Models are Often F.L.A.W.E.D. Artifacts with Built-in Flaws," aimed to provide insight into the current capabilities of LLMs rather than a direct comparison among them.
Vulnerabilities Studied
The vulnerabilities analyzed include:
- CVE-2026-31431: Linux Privilege Escalation (Copy Fail)
- CVE-2026-34197: Remote Code Execution ActiveMQ
- CVE-2026-8512: Use After Free in Chrome's File System Access API on macOS
- CVE-2026-45185: Unauthenticated Remote Code Execution EXIM
- CVE-2026-22738: Remote Code Execution SpringAI SpEL
- GHSA-wpqr-6v78-jr5g: Remote Code Execution Gemini CLI
The models were tasked with generating patches for each of these vulnerabilities, resulting in a total of 6,080 attempts, or about 3,040 per vulnerability, under varying environmental conditions and with nine different prompts per bug.
Results of AI-Generated Patches
The results show that AI successfully produced appropriate patches in only 26% of cases. Additionally, 21% of the patches, while correcting the bug, altered the application's behavior. In 53.9% of cases, the LLMs failed to create an effective patch, sometimes introducing new bugs.
Reasons for AI Model Failures
The failure of AI models to generate effective patches is primarily due to the production of "Fix-Like" artifacts with built-in flaws. These patches, while appearing functional on the surface, do not fully resolve the vulnerabilities and may introduce fragile security mechanisms or new bugs, sometimes altering the normal behavior of applications.
Future Perspectives
To address these shortcomings, 1Password has made its FLAWED tools available on GitHub, allowing researchers to continue their own studies. Although AI is not yet ready to handle patching autonomously, it can still play a crucial role in identifying and triaging vulnerabilities. Keith Hoodlet, head of Off-by-1 Labs, emphasizes the importance of collaboration between human defenders and AI tools to determine the most critical bugs to fix. Human oversight remains essential for making informed decisions about which patches to apply and assessing the associated business risks.
Brief IA — L'actualité IA en français
L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.