The Thesis

Anthropic has named its first marked Claude models and opened a detection API to an approved list. It has published an accuracy figure for neither. The commitment can be checked against Anthropic's own documentation, but the accuracy cannot be checked by anyone outside that list. Institutions drafting provenance clauses before December are writing rules around a capability they cannot test.

The Signal

Three gaps between what has been promised and what can be checked.

Signal 01Marking covers two models out of a lineup.

What happened. Article 50(2) of the AI Act is a binding marking duty. Anthropic says it will meet that duty through the Commission's voluntary Code of Practice, and that models launched on or after August 2, 2026 are marked at launch worldwide. Fable 5.1 and Mythos 5.1, released September 1, are the only two named. Everything older sits in the retrofit queue, including Opus 5 and the June releases of Fable 5 and Mythos 5, with no completion date.

Why it matters. Euronews described Anthropic as watermarking all of Claude's output from August 2, but coverage is version-scoped. Marking attaches to a model ID, so a policy written at vendor level misses the question that decides the case: whether the identifier that produced the text carried a mark that day.

Signal 02The detector shipped to an approved list, without a number.

What happened. On August 14 Anthropic disclosed the method, a version of SynthID-Text published by Google DeepMind in Nature in 2024. That places the method family in the peer-reviewed record. Anthropic's own build remains undocumented, with no false-positive rate, no minimum passage length, and no account of how it differs. The detection API entered private preview on September 1, gated behind a request form.

The closest measured evidence comes from Yangxinyu Xie, Xuyang Chen, Zhimei Ren, and Weijie J. Su at Penn, in Harvard Data Science Review. Their paper opens on a threshold set by hand: 4% of human-written essays fell below it, alongside 25.5% of essays using grammar correction the assignment permitted. The authors include that to show why the approach fails. Calibrating against work from students known to have followed the guidelines brings false positives to 4.69% and 4.32%, close to the 5% ceiling the procedure holds.

Why it matters. A method for using watermark evidence fairly already exists in the literature, and nothing in the Act or the Code requires a deployer to use it. Calibration needs the provider to return a continuous p-value, and Anthropic has not said what its API returns.

Signal 03The gap has an expiry date.

What happened. Article 50 transparency obligations applied from August 2, 2026. Regulation (EU) 2026/1744, the Digital Omnibus on AI, gives providers whose systems were on the market before that date until December 2, 2026 to comply with Article 50(2). It is settled law, adopted July 8 and in force July 27. A second date sits in the Code of Practice, where signatories commit to detection interoperability by February 2, 2027. Signing the Code is voluntary, but the Commission monitors signatories against the commitments they sign, and Anthropic has tied its Article 50(2) compliance to the Code with no alternative route on the table.

Why it matters. Institutions will settle what a mark means before most can test one, and policies written to a deadline are rarely reopened.

The Verification Test

Claim under test. Anthropic meets the timetable: the lineup marked by December 2, third-party detection usable by February 2, 2027.

Test. Watch the Help Center article for the remaining model IDs, and watch what the preview detector returns. Searchlight adds a bar the Act does not impose: an error rate at a stated token count by February.

Fail smell. The detector settles into a verdict with no error rate and no passage-length floor, enough for an institution to act on and too thin for an individual to challenge.

The Playbook

Four steps for institutions acting on a provenance signal.

Step 01

Ask which model IDs carry a mark today, and ask again at every version change. A provider that supports watermarking may still serve unmarked output from an older identifier.

Step 02

Set a floor on what a check may run against. Anthropic says the mark thins out on factual passages, on code, and on proofreading. A 300-word answer may hold too few marked choices to support a finding.

Step 03

Decide what a positive result may be used for, and by whom. Bar watermark evidence as a sole basis until an error rate exists, and name who reviews a contested result. No vendor or regulator has published a dispute path.

Step 04

Demand subgroup error rates. Xie et al. found essays by non-native English speakers carried stronger raw signals after identical permitted edits, because the model edited them more heavily. Their weighted method reduces the gap but still lands above 5% in several configurations.

The Metric

Two separate tests at alpha 0.05. Uncalibrated watermark test: 4% of human-written essays and 25.5% of essays with permitted grammar correction fall below the threshold. Calibrated guideline-compliance test: false positives of 4.69% for standard conformal and 4.32% for hierarchical conformal methods. Study uses Gumbel-Max on Phi-4-mini-instruct; these are not Claude accuracy estimates.
Xie, Chen, Ren, and Su, Harvard Data Science Review 8.1, February 2026. Read the study · Download chart (SVG)

Uncalibrated. Share of essays whose watermark p-value falls below α = 0.05, testing whether the text is entirely human-written.

Calibrated. Share of guideline-following essays whose conformal p-value falls below α = 0.05, testing whether the submission followed the permitted AI guidelines.

The threshold is a ceiling on how often the test should fire on text that satisfies the null. The two halves ask different questions: the top asks whether a human wrote the text, the bottom asks whether the student followed the rules, and only the second is the one an instructor has. The authors include the top pair to show why a hand-set threshold fails. All four figures come from Gumbel-Max watermarking on Phi-4-mini-instruct, and Claude uses a different scheme on a different model.

The Lens — Horizon Search Institute

Responsible AI. A detected mark is evidence that a marked Claude model touched the text. An absent mark says nothing about origin, as Anthropic's documentation states.

Human Performance. The bias that discredited the first generation of AI-text classifiers reappears in the raw watermark signal.

Governance & Diplomacy. Anthropic says it cannot yet scope watermarking by region, so the EU rule applies everywhere. Brussels sets a global default through a US company's compliance decision.

Governance & Diplomacy. Marking is due December 2 and interoperability February 2. For most of that window only approved organizations can test the detector.

Links Worth Your Time

How Claude marks AI-generated content Anthropic's Help Center article, now naming Fable 5.1 and Mythos 5.1. Read the Limitations section first.

Claude Watermark Detector Access Request Form The private preview gate itself. Educational organizations, researchers, and media are all eligible categories, which means most of our readership can apply today.

Watermark in the Classroom: A Conformal Framework for Adaptive AI Usage Detection The Penn group's paper. Section 2 sets out the calibration procedure; Section 4.4 is where the scheme and model dependencies live.

Regulation (EU) 2026/1744 The Digital Omnibus on AI, which is where the December 2 date now sits in operative text.

Sources

  1. Anthropic. How Claude marks AI-generated content. Help Center, updated September 2026. Source ↗
  2. Anthropic. How Claude's text watermark works. August 14, 2026, updated September 1, 2026. Source ↗
  3. Dathathri, S., et al. Scalable watermarking for identifying large language model outputs. Nature 634, 818–823. October 2024. Source ↗
  4. European Commission. Transparency obligations under Article 50 of the AI Act (FAQ). July 24, 2026. Source ↗
  5. European Commission. Guidelines on the implementation of the transparency obligations for certain AI systems under Article 50 of the AI Act (final). July 20, 2026. Source ↗
  6. European Parliament and Council. Regulation (EU) 2024/1689 (Artificial Intelligence Act), Article 50. Source ↗
  7. European Parliament and Council. Regulation (EU) 2026/1744 of 8 July 2026 (Digital Omnibus on AI). Official Journal, July 24, 2026; in force July 27, 2026. Source ↗
  8. Euronews. EU compliance, delivered globally: Anthropic to watermark Claude's output worldwide. August 11, 2026. Source ↗
  9. Fortune. Anthropic plans to add an invisible mark to AI text as the industry scrambles to police AI slop. August 11, 2026. Source ↗
  10. TechCrunch. Anthropic says it will watermark text generated by its AI models. August 11, 2026. Source ↗
  11. TechCrunch. Anthropic shares more details about how Claude's new watermarks will work. August 15, 2026. Source ↗
  12. The Register. Anthropic pledges to embed watermarks to help discern AI slop in sop to EU. August 11, 2026. Source ↗
  13. Dastin, J. Anthropic rolls out Opus 5 AI model in efficiency upgrade. Reuters, July 24, 2026. Source ↗
  14. VentureBeat. Anthropic's Claude Fable 5.1 and Mythos 5.1 arrive with a 75% cost reduction for Fable cache reads. September 1, 2026. Source ↗
  15. Xie, Y., Chen, X., Ren, Z., and Su, W. J. Watermark in the Classroom: A Conformal Framework for Adaptive AI Usage Detection. Harvard Data Science Review 8.1, February 18, 2026. Source ↗
  16. Willecke, L. Transparency obligations for AI-generated content: the Code of Practice adequacy decision and the final EU Commission Guidelines on Article 50. Reed Smith, July 28, 2026. Source ↗

Method and Limitations

Signal-selection window: August 11 – September 5, 2026, reaching back to the August 2 application date of the EU AI Act's Article 50 obligations.

What the sources establish: that Anthropic has committed to model-level text watermarking; that the commitment covers models launched on or after August 2, 2026; that Fable 5.1 and Mythos 5.1, released September 1, are the only models Anthropic names as supported, with the rest of the lineup still being retrofitted on no published schedule; that the method is a version of SynthID-Text; that a detection API is in private preview for a defined set of applicant categories, including regulators, law enforcement, media, fact-checkers, independent researchers, educational organizations, and EU civil society groups; that Anthropic has published no accuracy figures, no minimum passage length, and no dispute procedure for its implementation; that Article 50(2) compliance for systems on the market before August 2 falls due December 2, 2026 under Regulation (EU) 2026/1744; and that Code signatories commit to an interoperability solution for detection by February 2, 2027. Article 50 is binding on all providers. The Code is voluntary to sign, and the Commission has said that enforcement for signatories will focus on monitoring their adherence to it.

What Searchlight infers: that the practical consequence of these facts is a quarter in which institutional expectations will form ahead of any publicly measurable capability. Nothing here establishes that Anthropic will miss the December deadline, that its implementation is inaccurate, or that the preview will fail to produce usable figures. The public record shows an implementation in transition and an accuracy gap that remains open.

On the Xie et al. figures: the 4% and 25.5% pair is the paper's uncalibrated baseline, presented by the authors as a demonstration of why a hand-set threshold fails to hold the intended false-positive rate. The 4.69% and 4.32% figures come from their standard and hierarchical conformal methods respectively, and are reported against a different null hypothesis than the first pair. All four are evidence about a method family rather than about Claude: they come from Gumbel-Max watermarking on Phi-4-mini-instruct, and the paper's Section 4.4 shows detection performance shifting when either the watermarking scheme or the underlying model changes. Claude uses a different scheme on a different model. That is why these numbers are labeled here rather than applied.

Author and institutional conflicts: none.

Issue Credits
Author
Oscar Tirabassi
Managing Editor
Ashwin Telang
Editor-in-Chief
David Lovejoy
Published by Horizon Search Institute, a registered trade name of HSI Research Foundation · EIN 42-1954110 · A Delaware nonprofit corporation · horizonsearch.org