Claude's new text watermark barely lasted a week before 'removers' flooded GitHub

Anthropic started marking text written by Claude. Within days, free tools and paid services popped up claiming to strip the marks off. None of them can prove it works.

ThreatVectr Newsdesk· 3 min read
Full-frame edge-to-edge photoreal news-editorial image of two identical translucent glass server modules side by side on a dark brushed-steel surface, one glowi
Share

Key points

  • Anthropic began quietly watermarking text produced by its Claude chatbot in early November 2024, a way of hiding a signal in the writing that says "a machine wrote this."
  • Within days, more than a dozen "watermark remover" tools appeared online, including one open source project on GitHub that has already collected over 4,500 stars.
  • Paid services are also charging users to "clean" AI writing so it slips past detectors used by schools and employers.
  • None of the removers can actually prove they work, because Anthropic has not released the matching detector that would let anyone check.
  • The situation shows how quickly a defensive feature can turn into a marketing opportunity for the other side, even when nobody knows if the defence is real.

Anthropic, the company behind the Claude chatbot, has started watermarking the text its model writes. A watermark here is not something you can see. It is a subtle pattern in word choice that a special detector can spot later, telling a teacher or an editor: this was written by a machine.

The feature went live quietly. Almost as quickly, the counter-industry showed up.

As first reported by BleepingComputer, at least a dozen tools now claim to strip Claude's watermark out of any text you paste in. One of them, hosted on GitHub, has already picked up more than 4,500 stars from developers. Others are paid services aimed squarely at students and freelancers who want AI writing to pass as human.

How does a text watermark actually work?

Think of it as a fingerprint hidden in the choice of words. When Claude writes a sentence, the model has many equally good options for the next word. Anthropic nudges it toward a specific pattern of choices that, across a paragraph or two, forms a signal only its detector can pick up.

A reader will not notice. A statistical tool checking thousands of words will.

This is not a new idea. Google's DeepMind published a similar system called SynthID, and OpenAI has admitted for over a year that it built one for ChatGPT but has not shipped it. The appeal for the companies is obvious. Schools, publishers and hiring managers all want a reliable way to tell human writing from machine writing.

Can the 'removers' really remove anything?

Honestly, nobody knows. That is the awkward part.

Anthropic has not released the detector that would go with its watermark. Without that detector, there is no way to paste text in, run a remover, and check whether the mark is gone. The removers are selling a cure for an illness nobody can currently diagnose.

In practice, most of these tools appear to do what older "AI humanisers" already did: reword sentences, swap synonyms, shuffle clauses. That might disrupt a watermark. It might not. It might also just make the writing worse. Anthropic has not commented on whether any of the tools defeat its system.

Should ordinary people care?

If you are a student, a job applicant or a small business owner using AI to draft emails, the practical takeaway is small but real. Detection tools are getting more confident, and the arms race between watermarking and removing is going to keep running. Any tool promising to make AI text "undetectable forever" is selling something it cannot deliver.

The failure mode here is a familiar one. A vendor ships a security or trust feature, keeps the verification piece private for good reasons, and a market immediately forms around claims nobody can test. It happened with AI image detectors last year. It is happening again now, faster.

One thing the post-mortem on this whole watermarking push will probably say: if you do not release the detector, you do not really have a watermark. You have a press release.

Operational takeaway: treat any "undetectable AI" service the same way you would treat a lock-picking kit that refuses to name the lock.

© 2026 Threat Vectr