Claude: Anthropic Details Watermarking and Its Deployment

Le brief IA que les pros lisent chaque soir
Les 7 actus IA du jour, décryptées en 5 min. Gratuit.
Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.
Choisis ton rythme
Gratuit · Pas de spam · Désabonnement en 1 clic
Anthropic explains how Claude marks its texts using a secret key, without adding words or compromising quality, according to the company. Marking varies in its application across different uses, such as translations, proofreading, and code, and is more reliable on longer texts. An API for detection has been announced, while the European framework and model-level activation render geographical avoidance or VPN use ineffective. Older, unmarked models only provide a temporary escape.
Anthropic Applies Marking Everywhere and Prepares Coverage for Older Models
The European Code of Good Practices on Transparency of AI-generated Content was adopted in July 2026 and has nearly 190 signatories, including most major model providers. Each player is implementing its own marking system, with a specific key and, where applicable, methods that may differ. The corresponding obligations have been in effect since August 2.
For Claude, models prior to August 2 are not yet subject to watermarking. Anthropic specifies that it benefits from a transition phase and plans to extend this marking to those models in the coming months. Therefore, it is still possible, for now, to avoid watermarking, but according to the company, this option is not expected to last.
Marking is applied directly at the model level and affects all users from the moment of implementation, regardless of the country. Using a VPN does not allow for circumventing this marking. Anthropic claims it does not have a sustainable solution to limit marking to a geographical area and indicates it is continuing to explore other alternatives. Connecting from outside the European Union has no impact on the generated text.
Fully Marked Translations, Much Less for Proofreading and Code
When it comes to translation, the entire text is marked, as Claude chooses each word; detection is therefore straightforward, and it is indeed a generated text. In contrast, for proofreading, since the majority of terms come from the original author, the watermark integrates poorly, and depending on the extent of the changes, the model's action may go unnoticed.
For code, the presence of marking remains very limited. Arbitrary choices must be made between different possible formulations, while code requires precise output, and any variation could compromise its functionality. Marking may appear in comments, which has only a minimal impact on the code itself.
A superficial modification is not enough to remove the watermark, while a complete rewrite eliminates it—provided that every word has been replaced. The reliability of detection also decreases on shorter texts due to an insufficient number of lexical choices.
A Secret Key Guides Word Choice Without Additional Cost or Measured Quality Loss
The watermark adds nothing to the text: it replaces random selection among equivalent candidates with the use of a secret key combined with context. The model is not pushed toward words it would not have chosen, and the resulting sequence becomes verifiable by the key holders. Anthropic has adopted the SynthID-Text method, published by Google DeepMind in Nature in 2024 and inspired by a proposal from Scott Aaronson in 2022.
According to Anthropic, no effect has been measured on content, creativity, or readability. A marked model serving part of the Gemini traffic did not produce a significant difference in user votes, and human evaluators did not note any difference in quality. Furthermore, the marking does not generate any additional tokens, hence the absence of extra costs or slowdowns.
The system does not allow tracing back to a user or organization: neither the watermark nor the key contains identifying information. The key only responds to the probability of Claude's intervention, and the marking pertains to the model's outputs.
Image Files: A Readable C2PA Signature, Dedicated Tool Planned
Images generated in PNG, JPG, or SVG format include a cryptographically signed annotation in their metadata, compliant with the open standard C2PA, already accessible via compatible tools. This is not a watermark, as the file is neither modified nor altered in a hidden manner. Anthropic plans to provide its own tool to read this metadata.
Detection API Announced, Distinct Approach from Keyless Detectors
Anthropic announces a detection API, the specifics of which are yet to be clarified. It will not operate like market AI text detectors that lack the key and rely on writing quirks, such as the construction "it's not X, it's Y."
The result of the detection will remain probabilistic and will not allow distinguishing a text entirely written by Claude from a text revised by Claude.
Brief IA — L'actualité IA en français
L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.