Saying a document is 'watermarked' is easy. Showing the maths behind that claim is harder, and that's the bit most explainers skip.
This is the deeper companion to our beginner guide: a plain-English look at the statistics an ai watermark detector actually runs, built on the green-list method from Kirchenbauer et al.
We'll walk through green and red token lists, why the test counts distinct bigrams, and a worked z-score example using WatermarkRemoverPro's own live demo figures.
- 01Before each token, the model secretly splits its vocabulary into a green list and a red list, chosen by a key.
- 02The model is nudged, not forced, to prefer green tokens as it writes.
- 03Genuinely marked text ends up with far more green tokens than chance predicts.
- 04A detector holding the same key counts green hits and turns that into a z-score.
- 05The test runs over distinct bigrams, not single tokens, because repeated pairs would skew the count.
- 06WatermarkRemoverPro only tests keys it actually holds, and says so; it never implies broader coverage than that.
What a green list and a red list actually are
Before a language model writes its next word, it has already calculated a probability for every word in its vocabulary. A green-list watermark steps in right at that moment.
Using a secret key and a hash of the words that came just before, the method splits the whole vocabulary into two piles for that single step: a green list and a red list, each roughly half the vocabulary.
That split happens fresh before every single token, so the green list for word 50 of a sentence looks nothing like the green list for word 51. It's not one fixed list applied throughout, since it reshuffles constantly.
How the nudge works, without forcing any word
The model doesn't have to pick a green word. It's simply nudged, given a small boost, toward the green list before it makes its choice.
If a red word is a much stronger fit for the sentence, the model can still choose it. The nudge shifts the odds; it doesn't override the model's judgement.
That's exactly why marked text still reads naturally. There's no telltale pattern of odd word choices to spot by eye, because the signal lives in the statistics, not in the prose.
Why marked text contains more green tokens than chance
Under the null hypothesis, meaning no watermark or the wrong key, roughly half of a document's tokens would land on the green list purely by chance, because that's how the vocabulary was split.
Watermarked text pushes that proportion up. Not to 100%, because the nudge is gentle and real writing still needs plenty of red-list words to make sense, but noticeably above the chance rate.
That excess, more green tokens than an unmarked document would produce, is the entire signal a detector is looking for.
The z-test: turning a token count into a verdict
A z-test compares what was actually observed against what chance alone would predict, and expresses the gap in standard deviations. A z-score of 0 means 'exactly what chance predicts'. A high z-score means the gap is very unlikely to be a coincidence.
In WatermarkRemoverPro's own positive-control test, marked text scored z greater than 8, a p-value below 1 in a million, under the correct key. The same text, checked under a different key, scored at chance.
On the live /verify demo page, a specimen of marked text scores z = 20.45 under the correct key, and z = 0.1, essentially nothing, for the identical text under a different key. That gap is the whole point: the key is what makes the signal visible at all.
Why the test counts distinct bigrams, not single tokens
Kirchenbauer et al.'s original method scores over bigrams, meaning pairs of consecutive tokens, rather than single words, and specifically over distinct bigrams, not every repeated occurrence.
A document that repeats one common phrase many times would otherwise skew a single-token count, making the test easier to fool or easier to accidentally trigger. Counting distinct pairs keeps repeated phrasing from dominating the score.
It also better matches how the green-list split actually works, since the list for each token depends on the token before it, so a pair, not a single word, is the natural unit to test.
Worked example: a z-score walkthrough
The table below walks through what a real check looks like in practice, using WatermarkRemoverPro's own measured figures alongside one illustrative case.
Notice the gap between the correct-key row and the wrong-key row on the exact same underlying text: that's the whole method proven live, not just claimed in a paper.
The final row is a reminder that very short passages simply don't carry enough tokens to score reliably, whichever key is used, because the test needs enough text to work with.
Why the key has to stay secret, and what that means for honesty
If a watermark's key were public, anyone could counterfeit the signal in unmarked text, or strip it from marked text by targeting exactly the tokens the key favours. Secrecy isn't an accident; it's what keeps the method meaningful.
That's why no model vendor publishes its own detection key, and it's also why WatermarkRemoverPro only ever tests the keys it actually holds: its own public reference key, plus any key a vendor or institution has supplied to it directly.
That scope is stated plainly rather than implied to be wider. A result under one key says nothing about text marked under a key WatermarkRemoverPro doesn't have, which is exactly why an absent mark is never treated as proof of human authorship, only as 'no mark found under the keys tested'.
| Scenario | Bigrams scored | Green-list hits | z-score | Verdict |
|---|---|---|---|---|
| Marked text, correct key (positive-control test) | several hundred | well above the ~50% chance rate | z > 8 (p < 1e-6) | Strong statistical signal |
| Same text, wrong key (/verify demo) | same text | close to the ~50% chance rate | z = 0.1 | No signal, effectively chance |
| Marked text, correct key (/verify demo specimen) | full specimen | well above chance | z = 20.45 | Very strong statistical signal |
| Short passage, illustrative example | too few to score reliably | not applicable | insufficient data | No verdict, needs more text |
“z=20.45 isn't a percentage and it isn't a vibe; it's how many standard deviations the green-token count sits from what chance alone would produce.”
Common pitfalls
- Treating a z-score like a percentage confidence figure, rather than what it actually is: a distance from chance.
- Assuming any high z-score proves a document is AI-written by a specific model, rather than marked under a specific key.
- Forgetting that a low z-score under one key says nothing about a different key.
- Expecting a very short passage to produce a reliable score at all.
A detected mark is not proof of authorship, and an absent mark is not proof of human authorship. WatermarkRemoverPro's on-device rewrite can reduce detectable evidence but cannot guarantee defeating a vendor's undisclosed watermark, on any tier.
On WatermarkRemoverPro