Brief IA

Anthropic: Claude Labels All Its Generated Content

⚖️ Regulation & Ethics·Tom Levy·

Anthropic: Claude Labels All Its Generated Content

Anthropic: Claude Labels All Its Generated Content
Key Takeaways
1Anthropic marks all content generated by Claude, including text, code, and images
2The marking applies worldwide, motivated by the European AI Act
3The mechanisms differ depending on the medium: semantic watermarking for text, name selection for code, C2PA manifest for images
4Detection only proves a passage from Claude, and removing the marking is often simple
💡Why it mattersThis system generalizes the traceability of AI-generated content, but its technical limitations leave uncertainties regarding attribution.
Le brief IA que lisent les pros

Le brief IA que les pros lisent chaque soir

Les 7 actus IA du jour, décryptées en 5 min. Gratuit.

Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.

Choisis ton rythme

Gratuit · Pas de spam · Désabonnement en 1 clic

📄
Full Analysis

Anthropic indicates that all content generated by Claude now carries a mark, whether it is text, code, or images. This feature, activated in the context of compliance with the European AI Act, is applied globally and employs different techniques depending on the medium. The publisher and a specialist also describe the limitations: detection only proves that the assistant had a role at some point, and removal can sometimes be straightforward.

Detection has limits and varies by text, code, and image

Text marking is considered the most resilient, disappearing only after significant rewriting. In contrast, code retains the signal less effectively, and simply renaming variables or comments is enough to erase it. Images, on the other hand, lose the mark as soon as a screenshot is taken. A positive detection only indicates that content was generated or modified by Claude at some point, including during a simple proofreading. Conversely, the absence of a mark does not allow one to conclude that there was no intervention from Claude. The marking system is real, but less robust than some concerns suggest.

Global scope and product variation of marking

Marking applies to all Claude products: the assistant, Claude Code, Cowork, Tag, as well as access via the API and the three main cloud platforms where the service is available. Anthropic specifies that the application of marking is global. Although the approach meets a European requirement, all of the company's models are affected, regardless of the deployment location.

European framework and timeline: implementation of Article 50

Anthropic has adhered to the transparency code related to the EU's AI Act. Article 50 will come into effect on August 2, 2026, and the company commits as a model and system provider. Models launched in the European Union from this date will incorporate marking from their launch, while earlier models will remain in a transition period. Anthropic expects that all its offered models will be subject to this change.

How text is marked: equivalent semantic variations

For text, marking relies on the selection of different but semantically equivalent formulations, allowing for the encoding of a detectable pattern. This principle is akin to steganography: just as a slight modification of pixels can hide a message in an image, subtle word choices can carry a signal without altering the understanding of the text. The exact mechanism is not detailed, but the illustration emphasizes variants that are indifferent to the reader.

In code: names and comments as signal vectors

In code, marking relies on elements where choice is free, such as the names of index variables, helper identifiers, or comments. These elements, interchangeable for execution and often overlooked during proofreading, serve as the basis for the marking pattern. However, very short snippets carry this signal less effectively.

In images and files: C2PA manifest in the header

Generated files contain a header distinct from the pixel content, capable of storing structured data. Claude uses a C2PA manifest in this area, which includes the tool used, a timestamp, and a cryptographic signature calculated on the pixels. This signature distinguishes C2PA from an EXIF tag: any modification of the image invalidates the proof and can be detected during verification. The display of the image remains unchanged, with the header not affecting the visible pixels. The use of C2PA is based on an open standard that is publicly documented.

What removes the mark: rewriting, renaming, conversion

For images, converting the format, re-saving the file, or taking a screenshot removes the header and thus the mark; automatic processes like resizing by a CMS or recompression by messaging can have the same effect. For text, only significant transformations, such as extensive paraphrasing or translation that significantly alters word choices, eliminate the signal. In code, renaming variables and helpers and rewriting comments is often sufficient; automatic naming rules applied by a linter can remove the marking without voluntary intervention. The description of the text watermark implementation relies on subtle variations comparable to steganography, and the application of marking is not limited to European users.

Brief IA — L'actualité IA en français

L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.