Ask whether an AI detector is accurate and the honest answer is: accurate in which language?
Detection tools built and tuned mostly on English text do not automatically carry that reliability into Spanish, French, German or Portuguese.
This piece looks at why AI detector accuracy by language genuinely differs, what WatermarkRemoverPro measured to handle it properly, and why guessing across a language boundary is worse than saying "unsupported".
- 01The keyed maths behind a watermark test does not care what language the text is in.
- 02But measuring typical writing style needs a reference sample in the SAME language.
- 03WatermarkRemoverPro built its five reference baselines from real, contemporary Wikipedia prose, around 597,000 words in total.
- 04Independent research shows detectors misjudging text across language and register lines, most sharply against non-native English writers.
- 05WatermarkRemoverPro reports "unsupported" rather than testing text against the wrong language's baseline.
Why language changes the accuracy question
Most AI detection tools started life trained and tested on English. That is not a criticism so much as a fact of where the early research and the largest datasets happened to sit.
The trouble comes when the same tool, tuned on English patterns, gets pointed at a French essay or a German report and produces a confident-looking score anyway. Confidence and accuracy are not the same thing, and the gap between them widens fast once you cross a language boundary.
The watermark maths itself is language-independent
Here is the useful part. A keyed green-list watermark test, the method described in Kirchenbauer, Geiping, Wen, Katz, Miers and Goldstein's 2023 paper, "A Watermark for Large Language Models", works by checking whether a specific, keyed set of tokens turns up more often than chance would predict.
That statistical logic does not care what language the tokens belong to. A z-test over bigrams works the same way whether the text is in English or Portuguese, provided you have the right key and the right tokenisation for that language.
But style-based measurement needs a same-language reference
The keyed test is only half the picture. To judge whether a piece of writing looks statistically ordinary or unusual for its language, you need a baseline of real, typical writing in that exact language to compare against.
Borrow a French baseline to judge an English document, or vice versa, and you are comparing apples to a completely different orchard. Sentence rhythms, common bigrams, and typical vocabulary spread all differ language by language, so a substitute-language comparison produces numbers that look plausible and mean nothing.
How the five language baselines were built
WatermarkRemoverPro measured a separate reference baseline for each of its five supported languages (English, Spanish, French, German and Portuguese), using real, contemporary Wikipedia prose rather than synthetic or translated text.
Every source document was run through the engine's own language identifier before being kept, so a mislabelled or mixed-language page could not slip into the wrong baseline. In total that comes to roughly 597,000 words across the five languages, from 216,157 for English down to 76,142 for Portuguese, shown in the table below.
Corpus size and corpus quality are not the same lever, and it is worth being precise about why both matter. A baseline built from a smaller set of genuine, contemporary, verified-by-language prose is more useful than a larger one padded out with duplicate, scraped or machine-translated text, because what actually drives an accurate comparison is how well the sample represents ordinary sentence rhythm and vocabulary spread in that language today, not the raw word count sitting in a spreadsheet. That is precisely why every one of the roughly 597,000 words behind WatermarkRemoverPro's five baselines passed through the same per-document language check before being counted, rather than simply pooling whatever text was available and hoping sheer volume would smooth out the noise on its own.
A concrete case makes the point clearer. Take a 500-word German cover letter, written formally, with the longer compound nouns and clause structures typical of that register. Checked correctly against the German baseline, those features compare against genuine German business writing, which contains plenty of the same patterns, so the letter reads as statistically unremarkable. Checked instead against the English baseline, something WatermarkRemoverPro deliberately refuses to do, that same compound-noun density and clause structure would look completely alien next to typical English sentence patterns, and could produce a misleadingly unusual-looking score for writing that is, in its own language, entirely ordinary.
What goes wrong when a detector crosses language lines
The sharpest illustration of this problem is not really about language at all. It is about register instead. Liang, Yuksekgonul, Mao, Wu and Zou's 2023 study, "GPT detectors are biased against non-native English writers," found that the detectors they tested consistently misclassified non-native English writing as AI-generated, while accurately identifying writing from native speakers.
The likely cause is that non-native writing often carries simpler sentence structure and a narrower vocabulary range, traits that some detectors had learned to associate with AI output, purely because AI text also tends to look statistically "smooth" in similar ways. It is a reminder that a detector trained mostly on one register of one language can carry hidden assumptions into every score it produces, even within a single language.
Why WatermarkRemoverPro refuses to guess with a substitute-language baseline
Given all that, WatermarkRemoverPro takes a deliberately narrow position: if a document is not written in one of the five supported languages, or the language cannot be confidently identified, the check reports "unsupported" rather than quietly substituting a different language's baseline and producing a number anyway.
A wrong-but-confident-looking result is worse than no result. Saying "unsupported" costs nothing except a slightly less satisfying screen. Saying "here's a score" built on the wrong reference data could cost someone their credibility.
What this means if your language isn't covered
If you write in a language outside the current five, WatermarkRemoverPro will not force a result out of the wrong baseline. That is a limit worth knowing before you rely on the tool, not after.
Within the five supported languages (English, Spanish, French, German and Portuguese), each one gets its own measured reference, checked on the Check page or via the language-specific landing pages, so a result in French is being judged against genuine French writing, not a translated proxy for it.
| Language | Reference corpus size (words) |
|---|---|
| English | 216,157 |
| Spanish | 92,667 |
| French | 92,643 |
| German | 112,172 |
| Portuguese | 76,142 |
“People assume a bigger model automatically means better multilingual accuracy. What actually moves the needle is whether you measured a proper same-language reference sample, or borrowed one from somewhere else and hoped for the best.”
Common pitfalls
- Assuming a detector's advertised accuracy figure applies equally across every language it accepts.
- Running non-native English writing through a general classifier and treating a flagged result as final, without accounting for known register bias.
- Trusting a score for a language the tool does not clearly state it has a dedicated reference baseline for.
- Confusing "the maths works the same everywhere" with "the accuracy is the same everywhere": the test logic travels, but the reference data does not, unless it was measured separately.
A detected mark is not proof of authorship, and an absent mark is not proof of human authorship. WatermarkRemoverPro's on-device rewrite can reduce detectable evidence but cannot guarantee defeating a vendor's undisclosed watermark, on any tier.
On WatermarkRemoverPro