Overview
- Anthropic began adding invisible, statistical watermarks to text from Claude models launched on or after August 2, 2026, and says older models will be retrofitted by the December 2, 2026 deadline.
- The watermark is implemented at the model level by nudging word choices so a statistical detector can flag text that the model processed rather than by adding visible tags.
- Anthropic applies the marks globally and records provenance for images and other files using C2PA-signed metadata, but the company has not yet released the detection tool or full technical details for outside verification.
- Researchers and open-source developers have already published methods to test, degrade, or try to remove the marks, and experts warn the signal can be weakened by paraphrasing, heavy editing, screenshots, or metadata edits.
- Critics say the system risks misattributing lightly edited human work, centralizing detection authority in Anthropic, and prompting an arms race between marking and evasion while the EU law enforces fines for noncompliance.