Anthropic has provided more details on how it plans to watermark text generated by its Claude chatbot, as the company moves to comply with transparency requirements under the European Union’s AI Act.
In a blog post published Friday, Anthropic addressed questions about how the watermarking system will work, whether it can be removed through editing and how the technology will affect computer code.
The move has drawn attention from Claude users since Anthropic announced earlier this week that it would introduce watermarking to meet the EU AI Act’s Transparency Code, which requires AI companies to implement systems capable of identifying AI-generated content.
Anthropic said its approach will rely on patterns embedded in Claude’s responses that are not visible to readers but can be detected using a corresponding key.
The company explained that when Claude makes what it described as “low-stakes choices” — such as selecting between the words “overcast” and “grey” when describing the weather, it can use those choices to create a detectable pattern within the generated text.
According to Anthropic, the pattern would remain “undetectable to the reader” while being identifiable to anyone who has the key needed to decode it.
“Watermarking does not impact the quality of Claude’s output,” the company said. “To a reader, a watermarked response is indistinguishable from an unwatermarked one.”
Anthropic said it will adopt the SynthID-Text watermarking method developed by Google DeepMind in 2024 and plans to launch an API that will allow users to detect the watermark.
The company also distinguished its watermarking system from AI detection tools offered by firms such as Pangram, which analyse writing patterns or stylistic “tells” to determine whether content was generated by AI.
Anthropic said identifying such patterns is fundamentally different from detecting a watermark embedded in AI-generated text.
