Claude Now Watermarks Its Output. That's Not What Will Cost You a Placement.

Claude Now Watermarks Its Output. That's Not What Will Cost You a Placement.

Maurizio Petrone 19 min read Updated

A client recently asked us a question you may be about to ask your own content provider: Anthropic has started watermarking Claude’s output — do you monitor this, and does it change what you produce for our placements?

It’s the right question, and most providers will fumble it. Here is the public version of our answer, written for anyone about to put the same question to their own vendor.

The short version: the watermark is real, it matters, and it still isn’t the thing that will cost you a placement.

The detail most people are missing sits in one phrase of Anthropic’s policy on how Claude marks AI-generated content1: as of August 2026, the trigger for marking is “processed by Claude” — not “generated by Claude”. The question everyone is asking — “is this AI-written?” — is already the wrong question.

But the trigger and the signal are two different things, and the second is the one that decides anything. A mark is not stamped on the text; it is spread across the word choices Claude itself makes, so the more of the text Claude wrote, the more there is to find. That distinction is what most coverage of this collapses, and it is where the practical answers live. The rest of this post is about the right questions.

As of August 2026: what changed, and where detection stands

Everything in this section is dated on purpose: the policy is new, and these facts will move.

What Anthropic announced

As of August 2026, Anthropic’s page describes two complementary techniques: an imperceptible watermark embedded in generated text, and C2PA signed provenance metadata attached to generated files such as images. The same page states the text watermark is applied at the model level — present regardless of which Claude product or surface produced the text — and that marking applies wherever Claude is offered, worldwide, not only in the EU. Anthropic gives a reason for the global scope that is worth reading as provisional rather than settled: it is applying watermarking everywhere at launch “because we don’t yet have a durable way to scope it by region”, and says it will keep evaluating approaches.

A day later, Anthropic published a second and far more specific document explaining how the text watermark works2, and it changes what can honestly be said here. The method is named: it is a version of SynthID-Text, the technique Google DeepMind published in Nature3 in 2024, in a family of approaches going back to a 2022 proposal by Scott Aaronson. Nothing is inserted into the text and there are no hidden characters. Instead, where the next word is a genuinely open choice — Anthropic’s example is “the weather today was cold and…”, where “overcast” and “grey” are equally good — the watermark changes only the source of the randomness used to settle it, deriving the pick from a secret key and the preceding words. Read the sequence back with the key, and you can assign a probability that Claude produced it. Two consequences follow that clients ask about immediately: the mark carries no identifying information and, per Anthropic, cannot be traced to a person, organisation or chat; and it costs nothing, since it produces no extra tokens.

The trigger, as of the same date, is content processed by Claude, not only content authored by it. The page names proofreading, translation, summarization and file conversion as examples. But the second document is candid that these are not equivalent, and the difference matters more than the trigger does — the caveats near the end of this post pick that up, including a conclusion we got wrong the first time.

The regulatory background explains the date. Under the EU AI Act4, providers of generative AI systems must ensure outputs are marked in a machine-readable format and detectable as artificially generated; Anthropic says it has signed the Act’s Article 50(2) Code of Practice. The European Commission5 confirms those transparency rules apply from 2 August 2026 — the date Anthropic uses as its marking cutover — and that the recent “AI Omnibus” simplification moved the high-risk deadlines, not these.

One detail worth getting right: models launched before 2 August 2026 are not exempt. As of August 2026, Anthropic states the law provides those models a transition period and that it is working to add marking support to them as well — not yet marked, not a safe harbor.

Where detection stands

As of 17 August 2026, the technical description exists and the tool does not. Anthropic’s second document says a watermark detection API is coming — “We will soon be offering a watermark detection API. We’re in the process of working out the details of its implementation” — which is a commitment, not a shipped product, and it is the only thing that will be able to check a Claude mark.

Anthropic is direct about what such a check would prove, and it is less than people assume. Using the key, “one can only answer the question ‘What is the likelihood this was partly written by Claude?’” It cannot confirm text was human-written, it cannot tell whether some other AI wrote it, and it “cannot distinguish ‘Claude wrote this’ from ‘Claude heavily edited this’”. Absence of a mark proves nothing either: detection “doesn’t work well on small samples”, and a complete rewrite in which every word is replaced removes the watermark outright.

The regulation sets the floor those limits sit on. Under the final Code of Practice, per Bird & Bird’s measure-by-measure reading6, imperceptible watermarking is required for any free-form text longer than 200 tokens — roughly 150 words — with the only carve-out being “very short text” below that threshold. Free-form text is also the one case allowed a single marking layer rather than two, because it cannot carry metadata at all. And the detection duty is weaker here than elsewhere: providers must generally offer detection free of charge, but for free-form text watermarks — “because they are inherently less reliable” — they may restrict detection to verified expert users. The obligation to make watermark detection interoperable across providers does not bite until 2 February 2027.

The nearest comparable public tool is no help for text either: as of the same date, Google DeepMind’s SynthID Detector portal7 accepts images, video and audio — not text — and is still in limited testing.

And the question a placement buyer actually cares about: as of 16 August 2026, we could find no statement from any search engine or AI platform — in either direction — that provenance signals feed ranking, indexing or filtering. That is a bounded “none found”, dated for a reason, not a promise about next year.

Provenance and quality are different axes

Here is the reframe the whole topic needs: provenance — which tool touched the text — and quality — whether the text deserves to exist — are different axes. Search engines have only ever told us they act on the second.

Google has never described provenance as a ranking signal. That is not the same as Google saying it is not one; no public statement exists in either direction. But the uses Google has described are exactly two. In Google’s blog post on its C2PA integration8, from 2024, provenance metadata powers the “About this image” label in Search and, in Ads, Google’s goal to “use C2PA signals to inform how we enforce key policies”. Labeling and ads enforcement — and the Search integration is image-only. As of this writing, there is no text-provenance surface in Search at all.

What Google does enforce sits on the quality axis, and it is written to be tool-agnostic. Google’s spam policies9 define scaled content abuse as many pages generated “for the primary purpose of manipulating search rankings and not helping users” — adding that this holds “no matter how it’s created”. The first listed example is “Using generative AI tools or other similar tools to generate many pages without adding value for users”, and Google’s guidance on using generative AI content10 repeats that sentence as its operative line.

The Google Search Central Blog’s 2023 framing11 — “Rewarding high-quality content, however it is produced” — already drew its line at intent (“Using automation—including AI—to generate content with the primary purpose of manipulating ranking in search results is a violation of our spam policies”), and it should be read as the 2023 position; the current guidance is the narrower one. And in Google’s March 2024 spam update announcement12, the company wrote the tool out of the policy entirely: scaled content abuse applies “whether automation, humans or a combination are involved”.

The scale of AI-assisted publishing makes the same point from the other side. Graphite’s web-scale study13 found that AI-generated articles overtook human-written ones in November 2024 and have hovered around parity since — and, per the same page, “these articles largely do not appear in Google and ChatGPT”. Roughly half of new articles are AI-generated, and the low-value ones still fail to surface. That is quality selection at work, not provenance detection — and quality is the axis you can actually manage. We have written up how we produce content in the gen-AI era — the production side of exactly this argument.

What actually costs you a placement

If provenance isn’t the operative risk, what is? Two things, and both predate watermarking.

Editorial rejection came first

Publishers didn’t wait for watermarks to police AI content — they wrote it into editorial policy. WIRED’s generative-AI policy14, updated in May 2023, states: “We do not publish stories with text generated by AI, except when the fact that it’s AI-generated is the whole point of the story” — and treats undisclosed AI text as “tantamount to plagiarism”. Clarkesworld Magazine’s live submission guidelines15 refuse work at the point of entry: it “will not consider any submissions translated, written, developed, or assisted by these tools”, and warns that attempting to submit such work “may result in being banned from submitting works in the future”.

Where AI content slipped through anyway, the consequences were editorial, not algorithmic. Futurism’s investigation16 found Sports Illustrated publishing product reviews under fake authors with AI-generated headshots — a source involved said “The content is absolutely AI-generated” — after which publisher The Arena Group said it was “removing the content while our internal investigation continues and have since ended the partnership” with the contractor behind it. CNET’s own editor’s note on its AI experiment17, in early 2023, records 77 AI-drafted stories, a full audit after factual errors surfaced, appended corrections, and a halt — “We’ve paused”, in the publisher’s own words. Removal, correction, pause, terminated contract: that machinery ran on editorial judgment alone, and it still does.

The judgment is not anti-tool, either. Muck Rack’s 2026 State of Journalism report18 found 82% of journalists use at least one AI tool — while 88% “immediately disregard pitches that miss their beat”. The people who would run your copy through a detector are AI users themselves. Their filter is beat, relevance and quality; it always was.

The false-positive trap

The detectors people actually run today aren’t watermark checks. They are post-hoc classifiers — statistical guesses made from the text alone, with no key and no mark to find — and they misfire on real humans. Anthropic draws the same distinction, naming Pangram as an example: such services work by spotting the “tells” in AI phrasing, which is “fundamentally different from checking for a watermark”. Whatever a client’s detector reports, it is not reading the thing this post is about. A Stanford-affiliated study on arXiv19, from 2023, tested seven widely used GPT detectors on essays by non-native English speakers and found they “misclassified over half of the TOEFL essays as AI-generated (average false positive rate: 61.22%)”, with 97.80% of the essays flagged by at least one detector.

The vendors know. OpenAI withdrew its own classifier within six months; the shutdown note, as quoted by Ars Technica20, read: “As of July 20, 2023, the AI classifier is no longer available due to its low rate of accuracy.” At launch, OpenAI had disclosed it caught only 26% of AI-written text while “incorrectly labeling human-written works 9 percent of the time”. Vanderbilt University21 disabled Turnitin’s AI detector in 2023 and did the arithmetic in public: at the 1% false-positive rate Turnitin claimed, the 75,000 papers Vanderbilt had submitted in 2022 would mean “around 750 student papers could have been incorrectly labeled”.

Graphite’s study, cited above, validated its chosen detector on articles published before ChatGPT existed — and found it “classifies 4.2% of these articles as primarily AI-generated”. Even Google’s SynthID paper, published in Nature3 in 2024, notes that “post hoc detection systems can themselves be computationally expensive to run, and their practical usage is limited by their inconsistent performance.”

For placements, the false-positive cost lands hardest where a real expert’s voice is the product — ghostwritten expert commentary, where a classifier flagging a genuine expert’s own words kills a genuine opportunity. A provider who promises to “pass detection” is promising to satisfy tools whose documented failure mode is accusing humans.

Why there will be no universal AI detector

The absences in the dated section aren’t a temporary gap waiting for tooling to catch up. They’re structural.

Verifying a statistical watermark takes the generator’s private key. Google’s SynthID documentation22 instructs developers that each watermarking configuration “should be stored securely and privately, otherwise your watermark may be trivially replicable by others”, and offers detector deployment options from fully private to public. The barrier is not cost — the Nature paper is explicit that detection “does not require performing computationally expensive operations or even access to the underlying LLM”. The barrier is keys. Each vendor can detect, at most, its own marks. A universal detector would need every generator’s cooperation, and the incentives run the other way.

Coverage has holes no cooperation would close. One frontier lab built a text watermark and chose not to ship it: in a 2024 statement quoted by Language Log23, OpenAI said it has “developed a text watermarking method that we continue to consider” but judges it “less robust against globalized tampering” — translation, rewording with another model — “making it trivial to circumvention by bad actors”. Open-weight models sit outside any vendor’s marking reach entirely; the Nature paper itself notes that watermarking models deployed in a decentralized manner is difficult. And even a shipped watermark degrades: Google documents that detector confidence scores “can be greatly reduced when an AI-generated text is thoroughly rewritten, or translated to another language”.

Anthropic now says the same thing about its own reach, and confirms the shape of what is coming. Its key answers for Claude only: a watermark “can’t tell whether the text was written by a different AI (even if that other AI uses watermarking, it would have a different key; it might also use a different watermarking method altogether)”. And this is not an Anthropic project. Around 190 signatories signed the EU Code of Practice on Transparency of AI-Generated Content in July 2026, and Anthropic states plainly that other major model developers “will be implementing their own watermarks”. Many keys, many methods, one silo each.

One force runs the other way, and it deserves stating rather than ignoring: the Code requires signatories to put a watermark-detection interoperability solution in place by 2 February 2027 — a public access point that routes queries, an embedded signpost, a shared consortium, or an equivalent. That is a deadline, not a working system, and it is a routing layer rather than a shared key. It is the strongest reason to expect the picture to improve, and it still does not produce a detector that reads everything, because open-weight models sit outside the whole regime.

So the realistic end-state through at least 2027 is vendor-siloed provenance: each frontier lab able to check its own output, nobody able to check everything, and a clean result from any single detector meaning little.

The removal tool that argues against itself

If provenance cannot be verified, can it at least be scrubbed? There is an open-source project built for exactly that — watermarks-remover24, which began life as a Claude-specific “remove-claude-marks” tool — and the most instructive thing about it is its own README.

It opens with the concession: “Until vendors ship public detectors and keys, no tool can honestly certify ‘this fails the official check.’” And again: “It cannot certify that vendor detectors will fail.” Its deterministic layer strips invisible Unicode characters — which, against a Claude text mark, is now confirmed to be aiming at nothing. “Nothing is added to the text and there are no hidden characters,” Anthropic’s technical explainer states. That leaves only the statistical layer, explicitly “best-effort” — rewriting the text with another model, because a statistical watermark is spread across token choices and “removal means rewording, not restructuring”. Anthropic concedes the same shape from its side: “Light editing probably won’t remove the watermark completely; a complete rewrite where every word is replaced will.” The README then prices the attempt: the rewrite “flattens tone, voice, and precision”; on SEO, marketing and client work “that degradation is real and often visible”; the output “cannot exceed the rewrite model’s ceiling”. It ends on the question that closes the loop: if you are rewriting everything with a cheaper model anyway, generating directly with that model is “simpler, cheaper, and produces the same — or better — end result”. C2PA file metadata, meanwhile, it strips trivially — fragile by design of the file formats, consistent with Anthropic’s own metadata-loss list in the dated section.

There is a compliance edge to this too, and it is worth knowing before anyone proposes scrubbing as a service. Under the same Code of Practice, signatories must make best efforts to preserve markings, must prohibit tampering in their terms and conditions, and must “neither place nor advertise circumvention tools”. The tooling exists and will keep existing; what does not exist is a version of it a serious vendor can put in a proposal.

That is this post’s thesis in miniature. Trying to manage the provenance axis costs you the quality axis — the one that decides whether an editor keeps your piece.

What to ask any provider (including us)

So what does a good answer to the client’s question look like? Not a detection promise — a process. These are the six questions we would put to any content provider, including us; they describe our own outreach link building and listicle engagements, where you approve target publication, URL, anchor and draft before anything goes live.

  1. Are research, drafting and fact-checking separate stages, run by separate agents? Good answer: yes, and they can show you the separation rather than describe it. In ours, each of research, outlining, drafting, claim extraction, fact verification, correction and the final consistency pass is its own agent module with its own system prompt. The drafting agent is handed no tools at all — an empty tool list, not a restricted one — so it cannot search, scrape or fetch, and can only work from the fact file the research stage verified. Fact-checking is deliberately not one agent either: a navigator drives the lookups and a separate judge on a different vendor’s model renders the verdict, so the thing that checks a claim is not the thing that went looking for support for it.

  2. Where are the human gates, and can the machine ship without one? Good answer: the system enforces the gates — a document describing them is not enough. In ours, three are in code: a person selects the idea, approves the outline, and approves the content — and a placement cannot advance until a human sets that final approval. The pipeline cannot publish a placement article on its own.

  3. Are facts verified on opened source pages, or on search snippets? Good answer: pages. In our pipeline, the research agent opens and reads each source page before a fact enters the dossier, and the fact-check stage requires the exact quote or excerpt from the source page as evidence for each claim.

  4. Which vendor’s model writes the body — and can they name it? This is the question that actually determines marking exposure, and it is the one vendors most often answer with a category instead of a name. Because the watermark attaches to words the model chose, a provider whose drafting model is the one doing the marking has a real answer to give about most of the text; a provider who uses that model only to tidy someone else’s paragraph does not. Our own answer, at the level of detail we can stand behind: different stages of our pipeline run on different vendors’ models, and the model that drafts the article body is not the one being watermarked here. Be suspicious of “we use AI responsibly” — it is not an answer to this question.

  5. What exactly do they promise about detection? A provider guaranteeing content “passes AI detection” is promising a check that today nobody can run — per the dated section, Anthropic’s detection API is announced and unshipped — and that even once it ships only Anthropic will be able to run, only against its own marks, and possibly only for verified expert users. What such a provider is really testing against is the post-hoc classifiers below, documented to misfire on humans. The right promise is editorial: it reads well, it’s accurate, and a human approved it.

  6. What happens when a placement is rejected or comes down? The documented risk is editorial rejection — so that is the risk a contract should cover, with replacement or refund terms in writing. We have written up why we offer a satisfaction guarantee as the long-form answer to exactly this question.

If a provider can answer all six specifically, the watermarking news changes nothing for you. If they cannot, watermarking was never your biggest problem.

Two honest caveats

The trigger is broad; the signal is not

When we first published this post, we drew the wrong practical conclusion from the right observation. The trigger genuinely is “processed by” — Anthropic names proofreading and summarization among its examples, and “a human wrote it” is not the clean reassurance people take it for. But we went on to imply that an AI grammar pass therefore leaves a detectable mark on your document, and Anthropic’s technical explainer, published the day after, says the opposite in plain terms:

When Claude proofreads text written by a person, what it gives back has generally only been lightly edited; because nearly all the words are the person’s, there’s very little (if anything) for the watermark to attach to.

And, on a grammar-and-punctuation-only pass: “the watermark can only live in the handful of corrections, which might be too few to register.” The same logic thins the mark wherever the words are forced rather than chosen — Anthropic notes that factual passages and code carry less watermarking than open prose, because there is no free choice for the signal to ride on.

So the honest version is a gradient, not a switch. Text Claude substantially wrote is where the signal lives. Text Claude touched is not, and the regulation draws its own floor under that: nothing below roughly 150 words is required to be marked at all. The failure mode we described — the shop that swears it is “AI-free” while running AI grammar checks — is still wrong about its own process, and that is worth knowing about a vendor. It is not, on this evidence, likely to be a detection problem.

Nobody can verify provenance either way

The verification gap cuts both directions. We cannot prove our content is unmarked any more than a critic can prove a given page is marked — Anthropic’s own framing, per the dated section, is that a mark is a signal rather than proof, and that absence proves nothing. So this post over-claims in neither direction, and you should distrust anyone who does: a vendor guaranteeing “no watermarks” is selling the same unverifiable certainty as a critic claiming to have detected them.

This is a statement about today, and it has an expiry date attached to it. When Anthropic’s detection API ships, one party — Anthropic — will be able to answer one question, about its own marks, with a probability rather than a verdict. That is a narrowing of the gap, not a closing of it, and it is worth revisiting this post when it happens.

Cat and mouse: the founder’s read

Here I’ll switch from “we” to “I”, because this part is judgment, not citation. In twenty years of doing this, I’ve watched the same cycle run over and over: Google announces a hard line, the industry predicts the end of some practice, and then enforcement arrives late, hits the sloppiest operators, and stops short of finishing the job.

The freshest example is the one this post already touched. In its March 2024 announcement, cited above, Google named “very low-value, third-party content produced primarily for ranking purposes and without close oversight of a website owner” — parasite SEO, in the industry’s vocabulary — “publishing this policy two months in advance of enforcement on May 5”. Google’s Search Liaison, quoted by Ranktracker25, said enforcement kicked off on 6 May, the day after the policy took effect. In November 2024, per the updated policy language quoted by Website Builder Expert26, Google closed the first-party-oversight loophole, with Glenn Gabe reporting the affected sections of Forbes Advisor, CNN Underscored and WSJ Buyside effectively deindexed.

That’s the cat catching a mouse — some mice. What I still see in the wild, years after the first “we will enforce” statement, is the same practice, alive and doing fine. That’s not a criticism of Google; it’s the observable pace of enforcement against a web-scale problem — and this is a policy where detection is easy: the content is public and the pattern is obvious. Now scale that expectation to provenance, where, per the dated section, nobody can even run the check.

Two decades of predicted crackdowns were supposed to end this industry several times over, and it’s still here — because what survives every cycle is the same thing: work that holds up under the scrutiny it was always going to get. The watermark is not the problem. It never was.

You now know what a good answer to our client’s question sounds like. Take the six questions above to whoever builds your content, and listen for specifics. If the answers are vague, the vagueness is the finding.

Share

Want results like these?

Book a free strategy session and we'll show you exactly how editorial link building can grow your organic traffic.

No contracts. No commitment. Just a conversation.

Sources

  1. 1.AnthropicHow Claude marks AI-generated content · undated · the page dates its commitments to 2 August 2026 in its body text
  2. 2.AnthropicHow Claude's text watermark works · read August 2026 · the second, technical explainer; published after this post first ran and carrying no date of its own
  3. 3.NatureScalable watermarking for identifying large language model outputs · published 2024
  4. 4.EU AI ActArticle 50: Transparency Obligations for Providers and Deployers of Certain AI Systems · undated · undated article text of Regulation (EU) 2024/1689
  5. 5.European CommissionAI Act — Shaping Europe's digital future · undated · no page date; dates stated in body text
  6. 6.Bird & BirdTaking the EU AI Act to Practice: The Final Transparency Code of Practice · published 2026 · a measure-by-measure reading of the final Code, quoting its commitments and annex tables
  7. 7.Google DeepMindSynthID · undated
  8. 8.GoogleHow we're increasing transparency for gen AI content with the C2PA · published 2024
  9. 9.Google Search CentralSpam policies for Google web search · last updated May 15, 2026
  10. 10.Google Search CentralGoogle Search's guidance on using generative AI content on your website · undated
  11. 11.Google Search Central BlogGoogle Search's guidance about AI-generated content · published February 8, 2023 · the date is the page's byline
  12. 12.Google, The KeywordNew ways we're tackling spammy, low-quality content on Search · published March 2024 · carries an in-body update dated April 26, 2024
  13. 13.GraphiteMore Articles Are Now Created by AI Than Humans · undated
  14. 14.WIREDHow WIRED Will Use Generative AI Tools · last updated May 22, 2023 · per an in-body note
  15. 15.Clarkesworld MagazineSubmission Guidelines · read August 2026 · a live page with no policy date stamp
  16. 16.FuturismSports Illustrated Published Articles by Fake, AI-Generated Writers · undated
  17. 17.CNETCNET Is Testing an AI Engine. Here's What We've Learned, Mistakes and All · undated
  18. 18.Muck Rack, via The Manila TimesMuck Rack's 2026 State of Journalism Report Finds 82% of Journalists Use AI · published March 19, 2026 · the date is in the URL path
  19. 19.arXiv (Liang et al.)GPT detectors are biased against non-native English writers · published April 2023 · the arXiv identifier dates the submission
  20. 20.Ars TechnicaOpenAI discontinues its AI writing detector due to "low rate of accuracy" · published July 2023 · the date is in the URL path
  21. 21.Vanderbilt UniversityGuidance on AI detection and why we're disabling Turnitin's AI detector · published August 16, 2023 · the date is in the URL path
  22. 22.Google AI for DevelopersSynthID (Responsible GenAI Toolkit) · undated
  23. 23.Language LogWatermarking AI output, again · published 2024 · quotes OpenAI's May 2024 page and its August 4, 2024 update
  24. 24.GitHubguillaumemeyer/watermarks-remover · undated · no date on page; latest release v0.5.0
  25. 25.RanktrackerGoogle's Site Reputation Abuse Manual Actions in Effect · undated · quotes Google's Search Liaison
  26. 26.Website Builder ExpertGoogle Updates Site Reputation Abuse Policy: What's Changed? · undated · dated November 2024 in body text; quotes Google's updated policy language

Related posts