WatermarkRemoverPro

Green-List Watermarking: The Stats Behind AI Marks

WatermarkRemoverPro Content Team6 min readDeep dive
A performance analytics dashboard on a screen, representing the statistics behind green-list AI watermark detection

Photo via Unsplash

Saying a document is 'watermarked' is easy. Showing the maths behind that claim is harder, and that's the bit most explainers skip.

This is the deeper companion to our beginner guide: a plain-English look at the statistics an ai watermark detector actually runs, built on the green-list method from Kirchenbauer et al.

We'll walk through green and red token lists, why the test counts distinct bigrams, and a worked z-score example using WatermarkRemoverPro's own live demo figures.

TL;DR
  • 01Before each token, the model secretly splits its vocabulary into a green list and a red list, chosen by a key.
  • 02The model is nudged, not forced, to prefer green tokens as it writes.
  • 03Genuinely marked text ends up with far more green tokens than chance predicts.
  • 04A detector holding the same key counts green hits and turns that into a z-score.
  • 05The test runs over distinct bigrams, not single tokens, because repeated pairs would skew the count.
  • 06WatermarkRemoverPro only tests keys it actually holds, and says so; it never implies broader coverage than that.

What a green list and a red list actually are

Before a language model writes its next word, it has already calculated a probability for every word in its vocabulary. A green-list watermark steps in right at that moment.

Using a secret key and a hash of the words that came just before, the method splits the whole vocabulary into two piles for that single step: a green list and a red list, each roughly half the vocabulary.

That split happens fresh before every single token, so the green list for word 50 of a sentence looks nothing like the green list for word 51. It's not one fixed list applied throughout, since it reshuffles constantly.

How the nudge works, without forcing any word

The model doesn't have to pick a green word. It's simply nudged, given a small boost, toward the green list before it makes its choice.

If a red word is a much stronger fit for the sentence, the model can still choose it. The nudge shifts the odds; it doesn't override the model's judgement.

That's exactly why marked text still reads naturally. There's no telltale pattern of odd word choices to spot by eye, because the signal lives in the statistics, not in the prose.

Why marked text contains more green tokens than chance

Under the null hypothesis, meaning no watermark or the wrong key, roughly half of a document's tokens would land on the green list purely by chance, because that's how the vocabulary was split.

Watermarked text pushes that proportion up. Not to 100%, because the nudge is gentle and real writing still needs plenty of red-list words to make sense, but noticeably above the chance rate.

That excess, more green tokens than an unmarked document would produce, is the entire signal a detector is looking for.

The z-test: turning a token count into a verdict

A z-test compares what was actually observed against what chance alone would predict, and expresses the gap in standard deviations. A z-score of 0 means 'exactly what chance predicts'. A high z-score means the gap is very unlikely to be a coincidence.

In WatermarkRemoverPro's own positive-control test, marked text scored z greater than 8, a p-value below 1 in a million, under the correct key. The same text, checked under a different key, scored at chance.

On the live /verify demo page, a specimen of marked text scores z = 20.45 under the correct key, and z = 0.1, essentially nothing, for the identical text under a different key. That gap is the whole point: the key is what makes the signal visible at all.

Why the test counts distinct bigrams, not single tokens

Kirchenbauer et al.'s original method scores over bigrams, meaning pairs of consecutive tokens, rather than single words, and specifically over distinct bigrams, not every repeated occurrence.

A document that repeats one common phrase many times would otherwise skew a single-token count, making the test easier to fool or easier to accidentally trigger. Counting distinct pairs keeps repeated phrasing from dominating the score.

It also better matches how the green-list split actually works, since the list for each token depends on the token before it, so a pair, not a single word, is the natural unit to test.

Worked example: a z-score walkthrough

The table below walks through what a real check looks like in practice, using WatermarkRemoverPro's own measured figures alongside one illustrative case.

Notice the gap between the correct-key row and the wrong-key row on the exact same underlying text: that's the whole method proven live, not just claimed in a paper.

The final row is a reminder that very short passages simply don't carry enough tokens to score reliably, whichever key is used, because the test needs enough text to work with.

Why the key has to stay secret, and what that means for honesty

If a watermark's key were public, anyone could counterfeit the signal in unmarked text, or strip it from marked text by targeting exactly the tokens the key favours. Secrecy isn't an accident; it's what keeps the method meaningful.

That's why no model vendor publishes its own detection key, and it's also why WatermarkRemoverPro only ever tests the keys it actually holds: its own public reference key, plus any key a vendor or institution has supplied to it directly.

That scope is stated plainly rather than implied to be wider. A result under one key says nothing about text marked under a key WatermarkRemoverPro doesn't have, which is exactly why an absent mark is never treated as proof of human authorship, only as 'no mark found under the keys tested'.

ScenarioBigrams scoredGreen-list hitsz-scoreVerdict
Marked text, correct key (positive-control test)several hundredwell above the ~50% chance ratez > 8 (p < 1e-6)Strong statistical signal
Same text, wrong key (/verify demo)same textclose to the ~50% chance ratez = 0.1No signal, effectively chance
Marked text, correct key (/verify demo specimen)full specimenwell above chancez = 20.45Very strong statistical signal
Short passage, illustrative exampletoo few to score reliablynot applicableinsufficient dataNo verdict, needs more text
A worked z-test walkthrough: the first three rows are WatermarkRemoverPro's own measured and live-demo figures; the fourth is an illustrative example of an under-length document.

“z=20.45 isn't a percentage and it isn't a vibe; it's how many standard deviations the green-token count sits from what chance alone would produce.”

A WatermarkRemoverPro detection engineer, on the /verify page's z=20.45 result

Common pitfalls

  • Treating a z-score like a percentage confidence figure, rather than what it actually is: a distance from chance.
  • Assuming any high z-score proves a document is AI-written by a specific model, rather than marked under a specific key.
  • Forgetting that a low z-score under one key says nothing about a different key.
  • Expecting a very short passage to produce a reliable score at all.

A detected mark is not proof of authorship, and an absent mark is not proof of human authorship. WatermarkRemoverPro's on-device rewrite can reduce detectable evidence but cannot guarantee defeating a vendor's undisclosed watermark, on any tier.

Further reading
Answers, in full

Questions this post answers

What is a z-score in ai watermark detection, in plain terms?
It's a measure of how far the observed green-token count sits from what pure chance would produce, counted in standard deviations. A z-score near 0 looks like chance. A z-score like 20.45 is a very large, very unlikely-to-be-coincidence gap.
Why does green-list watermarking use bigrams instead of single words?
Because scoring distinct pairs of tokens, rather than single tokens, stops a repeated word or phrase from skewing the count. It keeps the statistical test fair and matches how the green-list split is actually generated, using the token that came before.
Why do ai watermark detectors need a secret key, and what happens if it leaks?
The key controls which tokens are green at each step. If it leaked, the signal could be counterfeited into unmarked text or specifically stripped from marked text, so it's kept private, which is also why any given detector can only test the keys it's actually been given.
Does a high z-score ever mean the same as '100% AI-written'?
No. It means the text is statistically very unlikely to be unmarked under that specific key, a narrower claim than 'this was written by AI'. It says nothing about text generated without a watermark at all, or marked under a different key.