Most of the AI provenance announcements I have read this year promised a property and left out the measurement. The 5 October OpenAI post, “Our approach to EU text provenance rules,” is a different kind of document. It says what the watermark is, where it will run, and, in a chart and three sentences, how often it works. Those are the first detection numbers I have seen from OpenAI for ChatGPT and Codex, so I read them slowly.
The commitments first. API customers anywhere can opt in to text watermarking for “select models,” and it stays off by default. Over the coming weeks OpenAI will add “an invisible watermark” to eligible ChatGPT and Codex text in the EU, across all plans, and not elsewhere: “We are not making text watermarking a global default at launch.” The method, called textGrain, “adds an invisible statistical signal to the model’s word choices.” A detector exists, but access is limited to “approved researchers and expert organizations,” case by case. The post says why: “Given the risk of missed watermarks and false positives, we are not making it publicly available at launch.”
That sentence is the center of the announcement, and the rest of this piece takes it at its word.
The numbers
OpenAI reports its figures at a target false-positive rate of 1%, on prompts from the ELI5 set. Passages of 200 tokens were detected about 80% of the time and passages of 400 tokens about 95% of the time, for content such as psychology (the post’s wording is “about 80% of 200-token passages, compared with about 95% of 400-token passages, for content such as psychology”). “Detection rates were substantially lower for content such as mathematics,” a rate that appears only in a chart, not in the text. Then the edit test, run on 400-token passages: “replacing 10% of words with synonyms reduced detection from about 92% to 66%. Replacing 25% of words reduced it to 17%.”
Three details in that chart are easy to miss.
The first is the 200-token row. At the target false-positive rate, one watermarked 200-token psychology passage in five goes undetected, and mathematics is worse, so one in five is the better of the two content types OpenAI shows. The 400-token row looks better, but the edit-test baseline for the same length is 92%, not 95%, and the post gives no reason for the three-point gap. I would guess different content mixes, and I would not put weight on that guess.
The second is the edit test. A tenth of the words swapped for synonyms takes detection from about 92% to 66%, and a quarter takes it to 17%. A quarter of the words is a heavy edit, but it is the edit a careful student or a content mill would make by accident while “improving” a draft. The post does not publish a number for paraphrase or for translation. It says only that text “may be too short, edited, or translated for detection to work reliably,” and promises to keep studying how watermarks “withstand editing and translation.” That promise is a fair description of where the data is: I could not find a measured translation rate in the post.
The third is the false-positive side. The 1% is a target the detector is tuned to, not a measured rate on a corpus of human writing. On arithmetic alone, 1% of a million human-written passages is 10,000 passages flagged as watermarked, and the post gives no measured figure for real human text to check that against. The technical report is candid about the same point. It derives a false-positive rate of exactly the chosen level, under assumptions of independence, and then says that a fixed deployed key and finite-precision arithmetic “require empirical calibration checks.” The report I read, dated 5 October, is a mathematical document, and it contains no detection-rate experiments that I could find. The blog post says it “will be updated with additional details in the coming weeks.”
OpenAI also reports that watermarking leaves quality largely unchanged on its Astra model at the maximum setting (“we do not see meaningful performance differences”), for example 94.44% against 93.94% on GPQA Diamond. The largest gap in its table is 3.1 points, on Terminal-Bench Science 0.1 (56.90% against 60.00%, with the watermarked run higher); on Terminal-Bench 4.0 the watermarked run also scored higher (56.06% against 53.90%). The post gives no confidence intervals or run counts, so gaps of that size may well be within ordinary run-to-run noise on agentic benchmarks; I take the claim as “no obvious damage,” and no stronger. (I made a similar point about a different vendor’s reliability numbers in Twenty Tries, Two Numbers.)
What the law asks for
The reason any of this is shipping now is Article 50(2) of the EU AI Act. Providers of systems that generate synthetic text must “ensure that the outputs of the AI system are marked in a machine-readable format and detectable as artificially generated or manipulated,” and must “ensure their technical solutions are effective, interoperable, robust and reliable as far as this is technically feasible.” The qualifier matters. It lets a provider argue that a 66% post-edit detection rate is as good as the current state of the art permits.
The dates are less tidy than the article. Article 50 applied from 2 August 2026. The Digital Omnibus on AI, Regulation (EU) 2026/1744, was published in the Official Journal on 24 July and adds, in its recital 38, “a transitional period of four months for providers who have already placed their systems on the market before the 2 August 2026.” The operative text is a new paragraph 4 in Article 111 of the AI Act: providers of systems generating synthetic text “that have been placed on the market before 2 August 2026 shall take the necessary steps in order to comply with Article 50(2) by 2 December 2026.” I have not verified whether ChatGPT and Codex count as systems placed on the market before the cutoff. If they do, an October rollout sits inside the grace window. If they do not, it is late by two months.
The Commission’s Code of Practice on transparency of AI-generated content, finalized on 10 June 2026, is the voluntary instrument OpenAI cites for how the detector is shared. I did not read the Code itself. A law-firm summary of it (Lewis Silkin, 24 July) reports three provisions that bear on the numbers above: watermarking is required for free-form text longer than 200 tokens, even though “reliability may be lower for shorter text”; providers must offer “a detection solution (free of charge as a general rule)”; and by 2 February 2027 they must implement an interoperability mechanism. Treat all three as second-hand until checked against the Code.
If the summary is right, two tensions follow, and I do not see either resolved in what I read. The 200-token threshold is the length at which OpenAI’s own chart shows about 80% detection for psychology text, and lower for mathematics. And a detector that is free “as a general rule” sits awkwardly beside one limited to approved researchers, although OpenAI cites the Code for exactly that restriction. Perhaps “as a general rule” leaves room. It is a question for the Commission, not for me.
The independent evidence is thin
I looked for measurements of deployed text watermarks from anyone other than the vendor, from 2026, and found none for ChatGPT or Claude. What exists is research on academic implementations. A July preprint by Tamim and Khan (arXiv 2607.16010) tested three schemes, including the MarkLLM implementation of Google’s SynthID-Text, against forensic-evidence standards. Across 846 valid paraphrase runs on 15 prompts per method, every initially detected KGW and Unigram text lost its watermark, and SynthID lost 98.3%. The SynthID configuration also flagged 5.4% of paraphrased human-written controls as AI-generated. This is a small preprint, a research implementation rather than a production system, a single attack type, and a sample of 15 prompts; it says nothing direct about textGrain. What it does show is the shape of the problem: robustness to meaning-preserving rewrites is the hard part, and vendors have not so far published it for their deployed systems. OpenAI says textGrain “matched or exceeded” other approaches including SynthID for text, then adds that “strong performance under ideal conditions does not guarantee reliable detection in everyday use.” That is the author of the comparison telling the reader not to over-trust it.
Anthropic announced its own text watermark in August, and its post expects other developers that signed the Code of Practice to follow. Anthropic says the watermark applies to future Claude models, with work under way to add it to models launched before 2 August 2026. Google’s Nature paper on SynthID-Text, published in October 2024, said such watermarking “has not been adopted in production systems” owing to quality, detectability and efficiency requirements, then reported a live trial on nearly 20 million Gemini responses in the same abstract. The EU rules are what now push the approach into general-purpose products.
A prototype from 2022
The oldest piece of this story is the one I use for calibration. In November 2022, Scott Aaronson, then on leave at OpenAI, wrote on his blog that he was building a tool for statistically watermarking a text model’s output, using “a cryptographic pseudorandom function, whose key is known only to OpenAI,” that an OpenAI engineer, Hendrik Kirchner, had built a working prototype, and that “a few hundred tokens seem to be enough to get a reasonable signal.” Almost four years later the shipped system reports 80% at 200 tokens and 95% at 400 for psychology text. The order of magnitude he named matches, though he gave no detection rate. What the prototype could not tell anyone in 2022, and what the blog post of 2026 still cannot, is how that detection holds up once the text has passed through an editor, a translator, or a student’s thesaurus.
How I would read the announcement
The post’s own caveat is the right frame: “The absence of a detected watermark does not prove human authorship.” I would add its mirror image. A detected watermark at a 1% target error rate is evidence about one passage, not a verdict, and a rule that treats it as a verdict will wrongly accuse some writers at exactly the scale the number implies. OpenAI has published a miss rate, which most vendors have not. The next step would be the same table for paraphrase, translation, and a corpus of real human text, and an access policy for the detector that fits the Code it cites.
References
- OpenAI (2026). Our approach to EU text provenance rules. OpenAI, 5 October 2026. The detection rates, quality comparisons and the textGrain comparison are OpenAI’s own claims.
- Li, X., Wen, G., Chen, X., Long, Q., Jain, A., Joly, F., Lam, M., Song, Q., Su, W. (2026). textGrain: Entropy-Calibrated Watermarking for Language Model Text. Technical report accompanying OpenAI’s 5 October post, dated 5 October 2026.
- Regulation (EU) 2024/1689 (Artificial Intelligence Act), Article 50. Official Journal, 12 July 2024. Text read at artificialintelligenceact.eu/article/50.
- Regulation (EU) 2026/1744 (Digital Omnibus on AI), recital 38 and amended Article 111(4). EUR-Lex, published 24 July 2026. The 2 December 2026 date is in Article 111(4) of the AI Act as added by this Regulation.
- European Commission (2026). Code of Practice on Transparency of AI-generated Content. Final code published 10 June 2026.
- Lewis Silkin (2026). The EU’s new AI labelling rules: what every organisation needs to know. 24 July 2026. Secondary summary of the Code; the 200-token, free-detection and 2 February 2027 provisions are taken from it.
- Tamim, Khan (2026). AI Watermark Evidence Fails Forensic Readiness: An Empirical Evaluation. arXiv:2607.16010, 17 July 2026. Preprint.
- Anthropic (2026). Claude text watermark. Anthropic, August 2026.
- Dathathri, S. et al. (2024). Scalable watermarking for identifying large language model outputs. Nature 634:818–823, 23 October 2024. Background.
- Aaronson, S. (2022). My AI Safety Lecture for UT Effective Altruism. Shtetl-Optimized, 28 November 2022. Background.