WatermarkRemoverPro

Check English writing for an AI provenance mark

WatermarkRemoverPro supports English with its own measured reference baseline. English has the largest reference corpus of the five, and is the language most institutional detectors were built and evaluated on.

Why a per-language baseline matters

The provenance-mark test itself is language-independent, since it counts word pairs against a keyed partition, and that arithmetic does not care what language the words are in.

The style measurement is a different matter. It compares your document to a reference corpus, so the reference has to be in the same language or the comparison is meaningless. Every deviation would simply be measuring the language difference.

WatermarkRemoverPro measured its English baseline from contemporary English prose, and every document in that corpus was verified to be English by the engine's own language identifier before it was included.

If the language cannot be determined

When the engine cannot confidently identify a document’s language it stops and asks, rather than picking the closest match. Analysing against the wrong baseline produces a real-looking number that means nothing at all, and a real-looking meaningless number is worse than no number.

You can also set the language explicitly before running the check.

Wherever this page describes a result: a detected mark is not proof of authorship, and an absent mark is not proof of human authorship. WatermarkRemoverPro's on-device rewrite can reduce detectable evidence but cannot guarantee defeating a vendor's undisclosed watermark, on any tier.

Answers, in full

Questions people actually ask

Is the check less accurate in English than in English?
The provenance-mark test behaves the same in every language, since it does not use a language model. The style measurement is as good as its corpus, and the English corpus is currently the largest. Each result reports the size and retrieval date of the corpus it was measured against, so you can judge it yourself.
What about languages that are not supported?
They are reported as unsupported rather than analysed against a substitute baseline. Adding a language means measuring a real corpus for it, not adding a name to a list.