Anthropic's new content marking: an invisible watermark weaves through everything Claude writes
- Starting August 2, text from Claude carries an invisible watermark, and generated image files come with signed provenance info. This applies across all products.
- This isn't just Anthropic. The EU law goes into effect on August 2, and 82 providers—including Google, Meta, OpenAI, and Microsoft—have signed the same Code of Practice.
- The key takeaway, straight from Anthropic: finding the marker doesn't prove the content was AI-generated, and the reverse is also true.
Everything Claude Writes Now Carries an Invisible Mark
Anthropic has published an official help doc detailing how Claude will mark its output. From now on, every piece of text you generate will contain a watermark you can't see, but machines can check.
The official document lays out four promises:
Claude models released in the EU on or after August 2, 2026, support machine-readable markers from day one. Text gets an embedded watermark, and files get signed provenance metadata.
This covers Claude Platform (API), Claude, Claude Code, Claude Cowork, and Claude Tag, in all regions where Claude is available. Some platforms or features might not support all marker types.
The Code of Practice requires providers to let users and third parties detect their markers. Anthropic says more technical details will be published later.
Models released before August 2 are in a transition period. Marker support is being added and this document will be updated.
One important note: this document is about commitments and design. The watermark algorithm is not public, and the detection tool isn't available yet. The original text says "future technical documentation" in both cases.
Required by EU Law—and Fines Start August 2
Why the sudden change? Article 50 of the EU AI Act, which covers "transparency obligations," is the answer. Specifically, Article 50(2) sets a clear requirement for model providers: outputs from generative AI systems must carry "effective, reliable, robust, and interoperable" machine-readable markers so that AI-generated or manipulated content can be identified.
This obligation becomes enforceable from August 2, 2026, overseen by market surveillance authorities in each member state. That's exactly where Anthropic's "August 2, 2026, and later" line comes from.
€15 million / 3%
Maximum fine for violating Article 50: €15 million or 3% of annual global turnover, whichever is higher. SMEs face the lower amount.
~190 orgs
Organizations that signed the related "AI-Generated Content Transparency Code of Practice" by end of July, including 82 providers.
Signing the Code is voluntary, but Article 50 obligations are mandatory. Signing has a direct benefit: the Commission and the AI Office have already deemed the Code "appropriate," letting signatories use it to demonstrate compliance. Non-signatories must prove their alternative measures suffice, case by case, to national regulators.
The 82 signatories in the provider section include Anthropic, Google, Meta, Microsoft, OpenAI, Mistral, Cohere, Aleph Alpha, Black Forest Labs, and Synthesia. Essentially all major model vendors are under the same Code. About half the signatories are small or new companies.
So viewing Claude's watermark as just "Anthropic added a new feature" misses the point. This is a whole industry moving in step, pushed by the same law at the same time. Anthropic just happens to have documented its approach publicly.
Two Types of Markers: Watermarks in Text, Signatures on Files
Since it's required by law, what exactly is Anthropic doing? Two technologies, each handling a different medium.
When Claude generates text, it weaves an imperceptible watermark directly into the text itself. This doesn't change the meaning, quality, or readability of the answer. The watermark is applied at the model level, so it's present regardless of which Claude product generated the text.
When Claude generates supported file types (.svg, .png, .jpg), it attaches a digital signature with provenance information, using the C2PA open standard. This can verify that a file was processed by Claude and also detect if it was later altered.
A watermark is like a pattern pressed into the paper pulp itself—it's part of the paper, only visible when held up to the light. Provenance metadata is like the tracking history on a shipping label: the package itself is unchanged, but the label records where it's been. The key distinction: you can't peel off a watermark, but you can peel off a label.
This distinction determines how much each marker can survive. Try the demo below to see what happens to each under common scenarios.
How an Invisible Watermark Gets Embedded in Text
An invisible mark that doesn't change the meaning sounds contradictory: if the words are the same, where can the mark hide?
Anthropic has not published its watermark algorithm. The official text only says they "weave an imperceptible watermark into the text itself." The following explanation is based on the public approach of similar technology, from Google's SynthID-Text paper in Nature (2024). This is not Anthropic's implementation, but it illustrates the core principle behind this type of watermark.
The key insight: when a model writes a word, it doesn't have just one choice. It calculates probabilities for a set of candidate words and picks one. Ask the model to write the same sentence twice, and you might get different word choices. This "selection freedom" is where the watermark hides.
The method works by subtly biasing this choice: a key, combined with the last few words written, generates a random number. This number scores each candidate word, and the higher-scoring candidate wins the selection process. The chosen words are still perfectly reasonable, the sentence still reads naturally, but collectively, the word choices carry a statistical bias.
With this principle in mind, the official list of limitations makes much more sense. They aren't just dry disclaimers—they're direct consequences of how the mechanism works.
Finding the Mark Doesn't Mean AI Wrote It
So can this be used to determine if a piece of writing is AI-generated? Here's the official statement: detecting the marker means the content may have been processed by Claude. It does not, by itself, confirm the complete origin of the content.
Two situations distort the signal. First, Claude might not have been involved in the original creation: people routinely use it to proofread, translate, summarize, or reformat, so text and ideas that originated elsewhere still carry Claude's marker on the output. Second, content might have been edited, excerpted, or mixed with other material after Claude processed it.
You can click through these scenarios to see when the marker is present and how to interpret it.
The Absence of the Mark Doesn't Mean It Wasn't AI
The reverse is also false. The official doc lists five cases where content was genuinely generated by Claude but won't have a detectable marker:
Content generated by models released before marker support was added won't have the marker. Support for older models is still being rolled out.
Text that has been significantly edited, rewritten, translated, or broken up and mixed into other writing.
Passages that are too short to support a reliable statistical signal.
Files that have gone through format conversion, re-saving, or screenshots, which can remove the provenance record.
The fifth case is platform differences: content generated on a platform, feature, or file type that doesn't support that marker type will also lack the marker. The official doc applies the same caveat to cloud platforms: when accessing Claude via AWS, Google Cloud, or Microsoft Foundry, the text watermark is applied, but the signed provenance metadata "may not be supported on all platforms."
So the Watermark is a Source Signal, Not a Plagiarism Detector
Since you can't infer reliably in either direction, how should you use it? This mechanism provides a source signal, not an AI determination.
If you're hoping to check student essays, submissions, or vendor deliverables, you'll be disappointed: finding a marker isn't proof of guilt—someone might have just used Claude to translate their own work. Not finding one isn't proof of innocence either—a rewrite, a screenshot, or just a few words of text can erase the signal.
These limitations come straight from Anthropic's own document. They're not external criticisms. And they were written with unusual candor. One section is even titled "Limitations," with the cases listed in both directions, including the admission that "detecting a marker doesn't confirm the full origin of content."
"Detection of a Claude marker indicates the content may have been processed by Claude. It does not, by itself, confirm the complete origin of the content."
Claude Help Center, "How Claude marks AI-generated content"
Looking at the bigger picture, this reflects the design intent of the EU law: Article 50 demands detectability, not condemnation. It adds a layer of source information to the information ecosystem, giving people an extra data point when consuming content—not a hammer for judgment.
If You Build Products on Claude, You Have a Duty Too
One final note for those integrating Claude into their own products. The document is unambiguous here: if you deploy Claude into your product, you are responsible for evaluating what Article 50 requires of your product and service.
This reflects the liability structure under EU law. It divides the generative AI chain into two ends with different obligations:
Per its commitments under the Code, Anthropic aims to support you in fulfilling your own transparency obligations, with technical guidance to be released gradually. Until then, a watermarked model does not equal a compliant product.
AI watermarks are a source clue, not a tool for AI determination. Finding the mark might just mean someone used Claude to translate or proofread. Not finding it might just mean the text was rewritten, was too short, or was screenshotted. You can't infer reliably in either direction—this is in Anthropic's official documentation.
Claude's Invisible Watermark: What It Can and Can't Prove
Anthropic's new content marking explained in one page: how the mark works, what it proves, and what it doesn't.
↓ Everything on one page · Includes an animated figure
New models ship with it; older models are getting it
Starting August 2, 2026, Claude models released in the EU embed an invisible watermark in the text they generate. Image files get signed provenance info. This covers the API, Claude, Claude Code, Cowork, and Tag across all regions. The source is the official Claude Help Center doc, which describes commitments and design—the algorithm isn't public, and the doc says detection tools are coming later.
2026-08-02
New Claude models ship with machine-readable markers
5 products
API · Claude · Claude Code · Cowork · Tag, globally
EU law mandates it—and Anthropic isn't alone
This comes from Article 50 of the EU AI Act, which requires model providers to add machine-readable markers to generated content. It's enforceable from August 2, 2026. Signing the accompanying transparency Code of Practice is voluntary, but the Article 50 obligation is mandatory. Signatories can use the Code to demonstrate compliance.
€15M / 3%
Max fine for violating Art. 50 (EU official data, whichever is higher)
82 providers
Signed the transparency Code by end of July, incl. Google, Meta, Microsoft, OpenAI
This is an entire industry moving in step under one law, not a single company’s choice.
How a sentence gets subtly marked
There are two separate layers. The text watermark is part of the text itself and survives copy-paste. File signatures use the C2PA standard and are attached to files like .svg, .png, and .jpg. They can be stripped by a format conversion or a screenshot. Anthropic hasn't disclosed its algorithm. To illustrate the general principle, we use a similar public technology (Google SynthID-Text, Nature 2024). This is not Anthropic's implementation.
At each step, the model picks a word from several candidates. The watermark uses a secret key to score these candidates, letting higher-scoring words win—while the sentence stays completely natural. Illustrative figure. Based on the public paper for Google SynthID-Text (Nature, 2024), not Anthropic's implementation; candidates and scores are illustrative.
Finding it isn't proof. Not finding it isn't proof.
Anthropic states this directly in the doc: detecting the marker only indicates content may have been processed by Claude.
Claude might have only proofread, translated, summarized, or converted the format. The original ideas came from elsewhere. Content might also have been edited, excerpted, or mixed with other material after processing.
If you build on Claude, you have obligations too
The marker at the model layer is Anthropic's duty. But if you deploy Claude in your product, you might fall under Art. 50(4) (deployer obligations). Deepfakes, and AI-generated text on public interest topics published without human review, require your own labeling.
smooth…
in every word
/ 3%
- × Meta
- × OpenAI
- × Microsoft
