Overview
- Anthropic has begun embedding a probabilistic, machine‑readable watermark into new Claude models and says older models will be retrofitted by December to comply with the EU’s AI Act.
- The watermark operates by biasing token and word choices to create a statistical signal that detectors can spot, not by inserting visible tags or metadata into text.
- Anthropic has not yet published a public detector, detection thresholds, or empirical error rates, leaving uncertainty about false positives, false negatives and who will get access to detection tools.
- Experts warn the watermark can persist through copy‑pasting, translation and some editing but is probabilistic: it is more reliable on longer samples and cannot show how much a human used Claude or whether the model only edited text.
- The rollout prompted user backlash and quick emergence of open‑source tools that claim to weaken or remove the mark, and other major AI firms are pushing their own provenance approaches so methods and policy responses are likely to evolve rapidly.