WatermarkRemoverPro

What Is an AI Watermark? How Detection Actually Works

WatermarkRemoverPro Content Team8 min readDeep dive
A wax seal pressed into paper, standing in for the statistical seal an AI watermark detector looks for in text

Photo via Unsplash

You've heard that AI writing can be 'watermarked', but what does that actually mean, and can a browser really spot it in your own words?

This guide walks through the real mechanism behind an ai watermark detector, in plain English, with nothing left unexplained.

We'll show the actual numbers from WatermarkRemoverPro's own positive-control test and its live /verify demo, so you can see detection working, not just take our word for it.

TL;DR
  • 01A green-list watermark works by nudging a model to prefer a hidden, key-specific set of 'green' tokens.
  • 02Detection is a statistical z-test, not a magic yes/no button.
  • 03You need the correct key to detect a mark reliably. A wrong key returns chance-level noise.
  • 04WatermarkRemoverPro checks run entirely in your browser; your document never leaves your device for a free check.
  • 05A detected mark is never proof of authorship on its own. It just shows a pattern was present under a specific key.
  • 06The same logic works in reverse: no mark under your keys proves nothing about the keys you don't hold.

The Basic Idea Behind Green-List Watermarking

Picture a language model choosing its next word. At almost every point in a sentence, several words would work fine. A green-list watermark uses that wiggle room. Before it writes anything, the model quietly splits its vocabulary into two piles for that moment, a 'green' pile and a 'red' pile, based on a hidden key and whatever came just before.

The model doesn't switch to nonsense. It's still nudged towards ordinary, sensible language. It just leans towards the green pile slightly more often than chance would predict. Do that over hundreds of words and a pattern builds up that a plain reader would never spot by eye.

That's the whole trick, really. Nothing is hidden inside the letters or the spacing. The pattern lives in which words got chosen, not how they're written. This is the method Kirchenbauer, Geiping, Wen, Katz, Miers and Goldstein described in their 2023 paper on watermarking large language models, and it's the same family of technique WatermarkRemoverPro's check is built to detect.

Why a Detector Needs a Key

Here's the catch that trips a lot of people up: the green and red piles aren't fixed. They're generated fresh, word by word, from a secret key plus whatever text came just before. Change the key and you get a completely different green list.

So a detector can't just 'look for watermarks' in general. It has to test one specific key at a time, and ask: for this passage, using this key, did green-listed words show up more often than pure chance allows?

This is also why WatermarkRemoverPro never claims to catch every AI-marked document out there. Nobody outside the model vendors publishes their detection keys. A check can only test the keys it actually holds. That's a genuine, permanent limit, stated plainly on the /limits page, not a caveat buried in the small print.

The Statistics: A Z-Test, Not a Guess

Once you've picked a key, the actual maths is a z-test. Line the text up as a sequence of distinct bigrams (word pairs), work out how many landed on the green list under that key, and compare that count with what plain, unwatermarked writing would produce.

A big enough gap between what you saw and what chance predicts gives you a large z-score, and a correspondingly tiny p-value. The higher the z-score, the less plausible it is that the pattern happened by accident.

This isn't guesswork or vibes-based pattern matching. It's a number you can recompute, check and argue with. That's a deliberately narrower promise than a general 'AI or not' classifier makes, and that's the point. A testable claim beats an unfalsifiable one.

WatermarkRemoverPro's Own Numbers: A Worked Example

Numbers are more convincing than descriptions, so here are real ones. In WatermarkRemoverPro's own positive-control test, a passage of marked text scored z > 8, with p < 1e-6, when tested under its correct key. Tested under a different key, the exact same text scored at chance, with no signal at all.

The live /verify page shows the same thing happening in public, not just in a lab note. A specimen of marked text there scores z = 20.45 under the correct key. Run that identical text past a different key and the score drops to z = 0.1, indistinguishable from ordinary prose.

That contrast is the whole demonstration. It's not 'trust us, it detects things'; it's the same words, two keys, two wildly different results, sitting on a page you can open right now.

Why On-Device Checking Matters

A lot of detection tools work by uploading your document to a server somewhere. WatermarkRemoverPro's free check doesn't. The whole test, up to 1,500 words, runs inside your own browser. Your document never leaves your device.

For a student worried about a false accusation, or a freelancer checking a draft before it goes near a client, that matters. You're not handing unpublished work to a third-party server just to learn whether a statistical pattern is present.

A free account raises the ceiling to 5,000 words and 20 checks a month, across five supported languages: English, Spanish, French, German and Portuguese, each measured against its own reference baseline rather than a rough, English-shaped guess.

What a Detected Mark Does Not Prove

This part is worth reading twice. A detected mark is not proof of authorship. Marks can turn up in quoted text, in translations, in text that was AI-assisted then heavily rewritten by a person, or in passages copied from somewhere already marked.

Equally, an absent mark isn't proof of human authorship. Marks are keyed constructions. No model vendor publishes its detection key. Marks survive editing poorly, so a lightly-touched AI passage might test clean under every key you hold. 'No mark detected' always means 'under the keys we tested', never 'this document is clean'.

That's a deliberately honest limit, and WatermarkRemoverPro states it on every report it produces. A watermark check is one data point. It's not a verdict.

How This Differs From a General AI Classifier

General AI-writing classifiers work differently. They look at style: sentence rhythm, vocabulary choice, how predictable the phrasing is, and estimate a probability that a human wrote it. That's a useful, but fuzzier, kind of evidence.

A watermark test asks a narrower, more falsifiable question: does this specific statistical pattern, under this specific key, appear more than chance allows? It can be right or wrong in a way you can check the working for.

Neither replaces the other. A classifier might flag writing that carries no watermark at all. A watermark check might find nothing in text a classifier is convinced is AI-written. They measure different things, and mixing them up is one of the more common mistakes people make when reading a report.

Getting Started With Your Own Check

If you want to see this working on your own words, the Check page is the place to start. Paste in up to 1,500 words for free, and the test runs there in your browser.

Read the result alongside the /limits page before drawing any conclusions from it. A high z-score under a key you tested is real information. It's just not the only piece of information a fair judgement needs.

Interest in this kind of provenance checking is only growing as frameworks like the NIST AI Risk Management Framework, and transparency rules such as the EU AI Act's, put more weight on being able to show your working. For a deeper look at the statistics specifically, WatermarkRemoverPro has a longer explainer on green-list watermarking.

TestKey usedResultWhat it shows
Positive-control testCorrect keyz > 8 (p < 1e-6)Statistically overwhelming signal
Positive-control testDifferent keyChance levelNo signal without the right key
Live /verify demoCorrect keyz = 20.45Detector confirms known marked text
Live /verify demoDifferent keyz = 0.1Same text, wrong key, no signal
WatermarkRemoverPro's own watermark test results, correct key vs wrong key

“The z-score only means something once you've named the key. Run the same passage past the wrong key and the signal disappears completely. That's not a bug; it's the whole design.”

A WatermarkRemoverPro detection engineer, on the green-list watermark test

Common pitfalls

  • Treating a high z-score as courtroom-grade proof of who wrote something, rather than one signal under one key.
  • Assuming every AI writing tool marks its output: many don't, and marks aren't standardised across vendors.
  • Forgetting that translation, heavy editing or quoting marked text can weaken or destroy a mark either way.
  • Testing against the wrong key and concluding 'no watermark' when really it's 'no watermark under this key'.

A detected mark is not proof of authorship, and an absent mark is not proof of human authorship. WatermarkRemoverPro's on-device rewrite can reduce detectable evidence but cannot guarantee defeating a vendor's undisclosed watermark, on any tier.

Further reading
Answers, in full

Questions this post answers

Is a detected AI watermark proof I used AI?
No. A detected mark tells you a specific statistical pattern showed up under one key you tested, but it doesn't tell you how the text ended up that way. Quoted, translated or lightly-edited marked text can carry a mark forward even when a human did real work on it.
How does an AI watermark detector actually work?
It counts how often green-listed word pairs appear under a chosen key, then runs a z-test comparing that count with what plain chance would produce. A high z-score means the pattern is very unlikely to be accidental.
What is a green-list watermark, in one sentence?
It's a hidden, key-based split of a model's vocabulary into 'green' and 'red' words, with generation nudged towards the green half often enough to be statistically detectable later.
Does editing a document remove the watermark?
Sometimes, and sometimes not. Marks generally survive light editing poorly and heavy rewriting even worse, but there's no reliable rule that guarantees removal against a specific vendor's undisclosed watermark, which is exactly why WatermarkRemoverPro's own on-device rewrite feature states that limit on every result rather than promising a guarantee it cannot verify.
Can WatermarkRemoverPro detect every AI watermark that exists?
No detector can. Detection only works against keys you actually hold, and no model vendor publishes its own detection key publicly.