AI Watermarks Are Here. The Danger Is They're Not That Good.

I use Claude every day. I’m writing this in Claude Code, the tool I use most, so treat this as a biased-but-informed opinion.
Everything it writes carries an invisible mark that can’t be turned off or switched off.
Most of the reaction I’ve read is about being caught.
That’s the wrong thing to worry about.
What Actually Changed On August 2, 2026
Article 50 of the EU AI Act became applicable on August 2, 2026. Providers of generative AI systems must now ensure their outputs are “marked in a machine-readable format and detectable as artificially generated or manipulated.”
Anthropic complied. Claude models launched on or after that date mark their output everywhere you use Claude: the API, the app, Claude Code, Cowork and Tag.
Which AI Providers Are Actually Watermarking Text
Every company below except two signed the EU’s General-Purpose AI Code of Practice, the voluntary pledge that says we will meet these obligations. Article 50 is the law; the Code is the promise. The gap between the two columns is the entire story.
| Company | Signed the Code | Marks text | Public detector |
|---|---|---|---|
| Anthropic | Yes | Yes, models from August 2, 2026 | Announced, not shipped |
| Yes | Yes, SynthID since 2024 | Portal built, waitlist only | |
| OpenAI | Yes | No, built one and shelved it | None for text |
| Meta | No, publicly declined | No text mark published | None published |
| Microsoft | Yes | No text mark published | None published |
| Mistral | Yes | No text mark documented | None |
| xAI | Safety chapter only | No text mark documented | None |
Status as published by each company as of 17 August 2026, not as implied by its marketing.
More on who signed what
Anthropic, Google and OpenAI also attach C2PA metadata to generated files, which is a separate obligation from marking text. Meta is absent from the signatory list entirely; xAI signed the Safety and Security chapter only, which leaves it obliged to demonstrate transparency compliance "via alternative adequate means." Mistral is the interesting case: it signed every chapter, but much of its catalogue ships as open weights, where marking cannot be enforced on hardware the provider does not control.
How Do AI Watermarks Work?
Two marks, applied in two different places, neither of them visible. What each one is attached to decides what destroys it.
Text
A version of SynthID-Text, the scheme Google DeepMind published in Nature in 2024. It lives in which words Claude picks, not in anything bolted on afterwards.
- Survives a copy-paste
- Dies to a paraphrase
Files
Signed C2PA metadata, travelling with the file rather than inside what you see or hear. The same standard the camera industry uses for provenance.
- Survives a straight copy
- Dies to a screenshot
No opt-out. It applies everywhere you use Claude.
Google has marked Gemini's text since 2024. On August 14, 2026 it went the other way on the visible mark, making it optional while keeping SynthID underneath.
-
Claude writes one word at a time
The night was ▮
- cold
- freezing
- chilly
- icy
Four words. All four finish the sentence perfectly well.
-
The mark tips the scales
- cold
- freezing
- chilly
- icy
Two of them are quietly favoured, so one of those two usually wins. The sentence reads exactly the same either way, which is the whole point.
-
One choice proves nothing. Thousands add up.
Claude62%A person50%The gap is the signal. It only exists across a lot of writing, which is why a short passage or a light edit gives a detector almost nothing to go on.
The night was ▮
- cold
- freezing
- chilly
- icy
- and
- then
- so
- yet
- the
- its
- that
- this
- streets
- roads
- blocks
- lanes
- emptied
- cleared
- quieted
- thinned
Two of the four are quietly favoured. One of those usually wins, and the sentence reads the same either way.
I Simplified This: Google DeepMind's own description: models generate "one word (token) at a time", each with a probability score, and SynthID "adjusts these probability scores to generate a watermark" without affecting quality. Percentages here are illustrative, not measured figures from any published detector.
Icons from the European Commission, reproduced unaltered in their white variant: "These icons are made publicly available for everyone to use freely, without the need for attribution to the Commission or the AI Office." Shown here to describe them, not to label this page. Retrieved 17 August 2026.
The developer reaction arrived within days, and Theo’s is a fair sample of where most of it went: what does this mean for the code I ship?
That is the reasonable question. It is not the dangerous one.
The Watermark Doesn’t Say What You Think It Says
This is Anthropic’s own documentation, not my editorializing. The company files all of this under a heading called “Limitations”:
https://support.claude.com/en/articles/16266773-how-claude-marks-ai-generated-content
A hit doesn’t prove Claude wrote it, and a miss doesn’t prove a human did. It can’t see other models at all, and it works poorly on short samples.
A signal that can be wrong in both directions is not evidence. It’s a hint.
All of this is on the record, published by Anthropic itself. Almost nobody acting on a result will read that page.
The Thing Everyone Got Wrong About Proofreading
The claim that spread fastest was that editing your own writing with Claude gets your writing stamped.
That’s not quite right. Anthropic says “the watermark only applies to words Claude chooses,” so after a light proofread there is “very little (if anything) for the watermark to attach to.” Article 50 agrees, exempting systems that perform “an assistive function for standard editing.”
The real problem is subtler, and worse.
The mark sits only on Claude’s words, but the detector can’t tell you which words those were. A heavy rewrite of your draft and a passage Claude wrote from nothing produce the same signal and the same one-word verdict.
Easy To Remove, Hard To Fake
As Claude wrote it
The night was cold and the streets emptied.
4 of 8 words sit on the favoured side
After a free paraphraser
The evening turned bitter, so the roads stood empty.
Every favoured word replaced
Same meaning. Same quality. Nothing left for a detector to count.
Illustrative sentences, not output from any real detector. The underlying finding is ETH Zurich's: an ordinary paraphraser stripped SynthID-Text more than 90% of the time.
Forging a mark onto human writing is harder. The same lab’s earlier watermark-stealing work broke an older green-list scheme for under $50, but SynthID held up far better against it. Treat the cheap-forgery headlines with suspicion: they’re about a different system.
Removing a real markHow often a rewrite both reads well and comes back clean
Faking a mark onto human writingHow often a fake both reads well and fools the detector
Figures from ETH Zurich's SRI Lab, which probed SynthID-Text directly. Both metrics are strict: an attempt only counts if it clears a quality bar and beats the detector, scored against a detector tuned to a false-positive rate of 1 in 1,000. Tripling the query budget from 30,000 to 90,000 took spoofing from 4% to 15% — still hard, and getting less hard with money. The under-$50 figure is that lab's earlier ICML 2024 attack on KGW2-SelfHash, a different and older scheme: it is the number most of the coverage has attached to SynthID by mistake. Retrieved 17 August 2026.
Meanwhile the detection API isn’t out. Anthropic says it is “in the process of working out the details of its implementation.”
Right now, anybody can remove the mark and nobody can check it.
The People Who Get Marked Are The Ones Who Complied
Anthropic marks its text. Google marks its text. OpenAI built a text watermarking tool in 2024 and never shipped it, saying it could “stigmatize the use of AI as a useful writing tool for non-native English speakers,” and that close to 30% of surveyed users said they’d use ChatGPT less if it existed.
That objection was the correct one, and it is about to come true anyway.
“AI detected” does not mean this person 100% used AI.
It will mean this person used the vendor that complied.
Publishers are working through their own version of that. Caleb Ulku’s Claude’s Watermarks Just Broke SEO argues the mark turns into a ranking liability for anyone writing with the tools that carry it. Whether search engines ever read it or not, the exposure lands the same way: on the people who used the vendor that complied.
What I’d Actually Do About It
Disclose your own process, in your own words, before a tool does it for you badly.
Do not build a hiring, grading or client policy on a hint. The company that built the detector says it isn’t conclusive. If they won’t go that far, you shouldn’t either.
Keep your drafts and your version history. That is real evidence, and it’s what you’ll want if you ever have to argue about this.
If you’re accused, ask what the detector returned and what its own documentation says that means. The honest answer is usually weaker than the accusation.
The Part That Should Worry You
Nobody is going to lose a contract because a mark correctly identified AI-generated filler.
They’ll lose it because a signal its own maker calls “not fully conclusive” got read as a verdict by somebody with no way to interrogate it, and no incentive to try.
This was never the surveillance problem people took it for. It’s an evidence problem, and the weakness is the danger.
A mark that a paraphrase removes, that nobody can check yet, and that only covers the companies which followed the rules is going to be treated as proof anyway.
Plan for that, not for the version in the headlines.
If you’re working out what any of this means for how your business uses AI, that’s what I do. You might also want the less comfortable piece about what all this means for jobs, or the guide to the tools worth using.
Every claim here was verified against the primary source in August 2026: Anthropic’s own documentation, the text of Article 50, the Nature paper, and the ETH Zurich research. This moves weekly. Check before you rely on it.