OpenAI is about to start fingerprinting your homework. On Monday, the company announced that text generated by ChatGPT and its coding tool Codex inside the European Union will soon carry an invisible watermark, a compliance move tied directly to the EU AI Act. The rollout lands over the coming weeks.
It’s called textGrain. You won’t see it. Nobody will. It lives in the words themselves: a subtle statistical pattern nudged into the model’s word choices as it writes. A matching detector can then look at a passage and estimate the odds it came from OpenAI. Under ideal conditions, the detector catches about 80% of watermarked text at 200 tokens and about 95% at 400 tokens, OpenAI said, with a 1% false-positive rate.
That’s the headline promise. The caveats, as always, are doing a lot of the work.
How the ChatGPT textGrain watermark works
Forget invisible ink metaphors. textGrain doesn’t insert hidden characters, metadata, or zero-width spaces you could strip out in a text editor. There’s nothing to delete. Instead, as ChatGPT composes a response, the model is steered with a secret key that sorts its next-word predictions into preferred and non-preferred buckets. Stack hundreds of these tiny nudges across a passage and a detector holding the key can spot the pattern, even after the text has been copied and pasted somewhere else.
OpenAI published a technical report on the method alongside the announcement, co-written with researchers from the University of Pennsylvania and Yale. That’s a meaningful signal about the company’s confidence here: watermarking research has been a graveyard of clever schemes that collapse the moment someone paraphrases the output, so OpenAI is putting peer-scrutiny-level detail behind this one.
The company says the watermark survives the most common escape route: copy-paste. Because the signal is in the word choices themselves, moving text into an email, a report, or a website doesn’t shake it loose. OpenAI also says it saw no meaningful change in model performance with watermarking switched on, and that the watermark can’t identify the user behind the text, read their prompts, or view their chat history.
Shorter passages are harder to pin down, and math content is tougher than prose. There’s less room to hide a statistical signal when the correct answer constrains every word.
What it can’t do
OpenAI is unusually blunt about the limits. In its own testing, replacing just 10% of words with synonyms dropped detection from about 92% to 66%. Another evaluation cited by reporters found that swapping a quarter of the words for synonyms cratered detection to around 17%. Translation, paraphrasing, and short snippets all degrade it.
“We’re expanding our approach to content provenance to include text in response to EU regulatory requirements, while recognizing the significant limitations of current text watermarking technology,” OpenAI said in a post on X.
That honesty is doing real work in the announcement. A detected watermark can show that OpenAI’s system generated or processed part of a text. It can’t tell you who used the system, how much of the text they wrote, who owns it, or whether any of it is accurate. And a missing watermark doesn’t prove a human wrote something. Teachers, editors, and hiring managers hoping for a lie detector are going to be disappointed.
Detector access reflects that caution. Initially, only approved researchers and specialist organizations get to use it, granted case by case. OpenAI says it plans to release the technology as open source eventually, but not until the reliability questions are better answered.
Why the EU gets it first, and only the EU
The EU AI Act’s transparency provisions took effect Aug. 2, 2026. Article 50 requires providers of generative AI systems to mark synthetic text in a way machines can identify. Systems placed on the market before that date get a limited grace period, with the Commission’s guidance pointing to compliance from Dec. 2, 2026, but the direction is clear: in Europe, AI-generated text must be machine-detectable.
So EU users on every ChatGPT and Codex plan will get the watermark automatically. They won’t be able to switch it off. Everyone else — including all US users — gets nothing by default. API developers anywhere in the world can opt in for select models starting Monday, Oct. 5, but the setting ships off by default, and OpenAI says it isn’t making text watermarking a global default at launch.
There’s no similar requirement in the US or India right now, which is why the map looks the way it does. Whether US regulators follow the EU’s lead is the open question, and OpenAI’s careful “not at launch” phrasing leaves the door open.
The bigger picture
OpenAI is not the first mover here. Anthropic introduced invisible watermarks for supported Claude models under the same EU rules less than two months ago, and Google DeepMind has been developing SynthID-Text, its own text watermarking system. The big labs all signed onto the EU’s transparency code of practice over the summer — roughly 190 signatories, including Anthropic, Google, Meta, Microsoft, and OpenAI — so the industry’s compliance posture is “shape the regime, don’t fight it.”
The honest take: textGrain is a compliance tool, not a truth machine. It gives regulators, researchers, and platforms a statistical signal where today there is none, and OpenAI deserves some credit for publishing its failure modes instead of burying them. But anyone expecting it to end the era of undisclosed AI-written content will be let down quickly. The watermark works best on long, untouched prose — and the internet’s slop pipeline runs on short, edited, translated content.
Still, the trajectory matters. Two years ago, “did an AI write this?” was unanswerable. Now the largest model providers are building machine-readable answers into their output by default, at least in Europe. The question isn’t whether the watermark is perfect. It’s whether “good enough for researchers” becomes the floor the rest of the world eventually adopts.
Frequently asked questions
What is OpenAI’s textGrain watermark?
textGrain is an invisible watermark that subtly shapes the statistical pattern of ChatGPT’s and Codex’s word choices so a detector can estimate whether text is AI-generated. It adds no hidden characters or metadata, so it survives copying and pasting.
Will the ChatGPT watermark affect users in the US?
No. The EU rollout is the only default deployment; US users won’t see watermarked text. API developers anywhere in the world can turn on watermarking for select models starting Oct. 5, but it’s off by default.
How accurate is OpenAI’s textGrain detector?
At a 1% false-positive rate, OpenAI says its detector catches about 80% of watermarked 200-token passages and about 95% of 400-token ones. Editing weakens it: swapping 10% of words for synonyms dropped detection from about 92% to 66%.
Does the watermark identify who wrote the text?
No. OpenAI says the watermark cannot identify a user, read prompts, or view chat history, and detector access is initially limited to approved researchers and expert organizations.
Sources: TechCrunch, Engadget, News18, The Decoder
