Particle.news

Open-Source Tool Quickly Defeats Claude’s Hidden Text Watermark

Anthropic’s refusal to let its assistant run the remover exposes a fast-moving technical arms race over provenance and user control.

Overview

  • Anthropic has applied a covert, statistical fingerprint to Claude’s text outputs that embeds choices between trivial tokens so a holder of the detection key can flag AI‑generated text.
  • An open-source project called watermarks-remover combines three layers—deterministic cleaning of invisible characters, statistical rewriting to break token patterns, and metadata stripping across file formats—and gained rapid community uptake on GitHub.
  • Users who asked Claude to install the remover as a skill were refused and received explanations defending the company’s watermarking policy and user preferences for labeled outputs.
  • A different model, GLM 5.2, was reported to accept and run the remover skill after Claude declined, showing how other providers and community tools can bypass platform limits.
  • The breach and refusal to cooperate create a fast escalation that raises practical questions for provenance standards, platform controls, user rights, and how regulators can keep pace with open‑source countermeasures.