Claude's new text watermark barely lasted a week before 'removers' flooded GitHub

Anthropic started marking text written by Claude. Within days, free tools and paid services popped up claiming to strip the marks off. None of them can prove it works.

ThreatVectr NewsdeskUpdated · Editor: Lee Brown· 3 min read
A GitHub repository page on multiple monitors showing watermark removal tool repositories and modifications uploaded rapidly, code commit timestamps visible, re
Share

Key points

  • Anthropic began quietly watermarking text produced by its Claude chatbot in early November 2024, hiding a signal in the writing that says a machine wrote it.
  • Within days, more than a dozen watermark remover tools appeared online, including one open source project on GitHub with over 4,500 stars.
  • Paid services are charging users to clean AI writing so it slips past detectors used by schools and employers.
  • None of the removers can prove they work, because Anthropic has not released the matching detector.
  • The situation shows how quickly a defensive feature becomes a marketing opportunity for the other side, even when nobody knows if the defence is real.

Anthropic, the company behind the Claude chatbot, has started watermarking the text its model writes. A watermark here isn't visible. It's a subtle pattern in word choice that a special detector can spot later, signalling to a teacher or an editor that a machine produced the text.

The feature went live quietly. The counter-industry showed up almost as fast.

As first reported by BleepingComputer, at least a dozen tools now claim to strip Claude's watermark from any text you paste in. One, hosted on GitHub, has already picked up over 4,500 stars from developers. Others are paid services aimed squarely at students and freelancers who want AI writing to pass as human.

How does a text watermark actually work?

Think of it as a fingerprint hidden in word choice. When Claude writes a sentence, the model has many equally good options for the next word. Anthropic nudges it toward a specific pattern that, across a paragraph or two, forms a signal only its detector can read.

A reader won't notice. A statistical tool scanning thousands of words will.

This isn't a new idea. Google's DeepMind published a similar system called SynthID. OpenAI has acknowledged for over a year that it built one for ChatGPT but hasn't shipped it. The appeal for these companies is obvious: schools, publishers and hiring managers all want a reliable way to tell human writing from machine writing. We covered the parallel challenge of tracing AI-generated video to its source model back in August, and the verification problem there is the same one Anthropic is now running into with text.

Can the 'removers' really remove anything?

Honestly, nobody knows. That's the awkward part.

Anthropic hasn't released the detector that would accompany its watermark. Without it, there's no way to paste text in, run a remover, and confirm whether the mark is gone. The removers are selling a cure for an illness nobody can currently diagnose.

In practice, most of these tools appear to do what older AI humanisers already did: reword sentences, swap synonyms, shuffle clauses. That might disrupt a watermark. It might not. The writing quality often suffers regardless. Anthropic hasn't commented on whether any of the tools defeat its system.

Should ordinary people care?

If you're a student, a job applicant or a small business owner using AI to draft emails, the practical implication is small but real. Detection tools are growing more confident, and the arms race between watermarking and removing will keep running. Any service promising to make AI text undetectable forever is selling something it cannot deliver.

The failure mode here is familiar. A vendor ships a trust feature, keeps the verification piece private, and a market immediately forms around claims nobody can test. It happened with AI image detectors in 2025. It's happening again now, faster.

What the post-mortem on this watermarking push will probably conclude: if you don't release the detector, you don't really have a watermark. You have a press release.

Operational takeaway: treat any undetectable AI service the same way you'd treat a lock-picking kit that refuses to name the lock.

© 2026 Threat Vectr