Overview
- Anthropic has rolled out a statistical watermark in Claude models launched on or after August 2 and says the mark is applied worldwide and will travel with copied-and-pasted text.
- The watermark works by biasing token selection during generation so a key-holder can detect a statistical pattern rather than embedding visible marks or hidden characters.
- By design the signal detects that Claude 'may have processed' content and cannot reliably tell the difference between text the model wrote and text it only proofread or lightly edited.
- Anthropic plans to publish detectors and retrofit older models but has not released detection keys, verification tools or empirical error rates, and researchers have already shown ways to weaken or remove the signal.
- Experts warn the mark could cause false attribution in schools, workplaces and publishing, complicate copyright and verification workflows, and push users to work around the watermark by retyping or avoiding direct output copies.