Claude Trapped: A Flaw Exposes Your Sensitive Personal Data

Le brief IA que les pros lisent chaque soir
Les 7 actus IA du jour, décryptées en 5 min. Gratuit.
Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.
Choisis ton rythme
Gratuit · Pas de spam · Désabonnement en 1 clic
A Flaw in Claude's Security System
Claude's web_fetch tool, designed to withstand data exfiltration attacks, was put to the test by Ayush Paul, who discovered a flaw in its architecture. This system, which was initially meant to protect users, turned out to be vulnerable to certain manipulations.
Risks Associated with Accessing Private Data
Claude's chat, due to its access to private data and its ability to interact with online content, is exposed to potential attacks. These attacks can exploit the memories of users' past interactions and use URLs to extract sensitive data.
Protective Measures by Anthropic
Anthropic had implemented restrictions for web_fetch, limiting its use to URLs entered by the user or generated by the web_search tool. This measure aimed to prevent malicious operations, such as concatenating responses to suspicious URLs.
The Flaw Exploited by Ayush Paul
Ayush Paul discovered that web_fetch could also access URLs embedded in previously retrieved pages. This meant that a malicious site could be created to entice the agent to follow a series of nested links, thereby facilitating data exfiltration.
Example of a Successful Attack
The attack exploited a trick where the AI agent had to navigate through a site, letter by letter, to access user profiles. The URLs were designed to be visited sequentially, such as:
- https://coffee.evil.com/a
- https://coffee.evil.com/b
This method allowed for the extraction of personal information such as the user's name, city of residence, and employer.
Anthropic's Response and Fixing the Flaw
Anthropic did not offer a reward for this discovery, claiming to have identified the flaw internally. They have since corrected the issue by preventing web_fetch from following additional links in the retrieved content. This fix aims to enhance security and prevent similar future attacks.
Brief IA — L'actualité IA en français
L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.