Ask five different tools for an AI detection false positive rate and you'll get five different numbers, measured five different ways.
This piece lays the published figures side by side: Turnitin, GPTZero, Originality.ai, and the independent research on non-native English writers, and explains why a single 'X% accurate' headline rarely tells the whole story.
The short version: methodology matters more than the number itself.
- 01'False positive rate' means different things depending on what level (document vs sentence) and what test set a company used.
- 02Turnitin's own blog reports under 1% at document level and around 4% at sentence level, on an 800,000-document test.
- 03GPTZero states a debiased approximately 1% false positive rate for ESL learners and 99% accuracy, but these are the vendor's own claims.
- 04Independent research (Liang et al.) found detectors studied were biased against non-native English writers, flagging their genuine writing as AI far more often than native writers' work.
- 05No published figure here is directly comparable to another: different tools, different tests, different definitions.
What 'False Positive' Actually Means Here
A false positive, in this context, is when a detector flags genuine human writing as AI-generated. It sounds simple. In practice, it depends heavily on what you're measuring, a whole document, or a single sentence, and what counted as 'genuinely human' in the test set to begin with.
Those two choices alone can move a headline number by several percentage points, before you've even got to how the detector itself works.
Why the Published Numbers Vary So Much
Different companies test against different document sets, written by different people, in different contexts. A test built from clean native-English academic essays will produce a different error rate than one built from a mix of languages, writing levels and genres.
The level of measurement matters just as much. A document-level false positive rate (did the whole essay get wrongly flagged) is almost always lower than a sentence-level one, because a single wrongly-flagged sentence in an otherwise correctly-cleared document doesn't move the document-level number at all.
None of this means the numbers are dishonest. It means they're answering narrower questions than the marketing copy around them sometimes suggests.
Turnitin's Own Figures
Turnitin has published two figures worth knowing. At document level, they report a false positive rate under 1%, specifically for documents where over 20% of the content was flagged as AI writing, based on an 800,000-document test set.
At sentence level, their own figure is closer to 4%. They've also reported that errors aren't evenly spread: 54% of false-positive sentences sit right next to genuine AI-written text, at the point where styles change mid-document.
Independent education press, including K-12 Dive, has covered Turnitin acknowledging these higher sentence-level rates publicly, which is worth noting as a point in favour of taking their own figures at face value.
GPTZero and Originality.ai's Vendor Claims
GPTZero's own site states 99% accuracy, with a claimed 96.5% accuracy on mixed human/AI documents, and describes itself as de-biased for ESL learners with a stated false positive rate around 1% for that group specifically.
Originality.ai's own site claims 97.8% accuracy for its multilingual model, citing peer-reviewed third-party studies, though the specific false positive rate isn't given on that page. Their FAQ acknowledges false positives happen and says the company shares figures in its own separate accuracy study.
These are vendor-published claims. They're not necessarily wrong, but they're a different category of evidence from independently reviewed research, and it's worth reading them with that in mind.
The Non-Native English Writer Bias Problem
The most important independent finding here comes from Liang, Yuksekgonul, Mao, Wu and Zou, published on arXiv. They found that the detectors they studied consistently misclassified genuine writing by non-native English speakers as AI-generated, while accurately identifying native English writers' work.
That's a bias problem hiding inside an otherwise reasonable-sounding accuracy figure. A tool can post a low overall false positive rate while still getting it wrong disproportionately for one group of writers, and an overall percentage won't show you that on its own.
This is exactly the kind of thing worth checking for directly if you're a non-native English writer worried about a flag, rather than trusting a single headline number to cover your situation.
The Comparison Table, and Its Limits
The table below pulls these figures together in one place. Read the caption carefully: these are a mix of vendor-published claims and independent research, measured in different ways, on different test sets. They are not directly comparable, one line against another, as if they were competing scores in the same race.
What they're useful for is spotting the pattern across all of them: every serious source here, including the vendors themselves, acknowledges that false positives happen. None claims perfection. That consistency, across otherwise very different numbers, is arguably the most trustworthy thing in the table.
How a Watermark Check Differs From All of This
Everything above concerns style-based classifiers: tools estimating whether writing sounds AI-generated. A statistical watermark check, like WatermarkRemoverPro's, is a different kind of measurement entirely: it tests for a specific, keyed pattern, not a style.
In its own positive-control test, WatermarkRemoverPro's check scored z > 8 (p < 1e-6) for correctly-keyed marked text, and chance-level for the same text under a different key, numbers that come from a testable statistical procedure rather than a trained style classifier's probability estimate.
That doesn't make it immune to false positives in some looser sense; it makes a different kind of claim altogether, so it doesn't belong on the same comparison line as a style classifier's accuracy figure.
Reading Any Single Number Responsibly
If you take one thing from all this, let it be a habit: whenever you see a detector's accuracy or false positive figure quoted, ask what it was measured against, at what level, and whether it's the vendor's own claim or someone else's independent study.
A number without that context is a headline, not evidence. With it, you can actually judge whether it applies to your situation.
| Source | Claimed / measured figure | Level / context |
|---|---|---|
| Turnitin (own blog) | Under 1% | Document-level, documents with over 20% AI writing, 800,000-document test |
| Turnitin (own blog) | Approximately 4% | Sentence-level, more common at human/AI transitions |
| GPTZero (own site) | Approximately 1% | Vendor-stated, for ESL writers specifically, after de-biasing work |
| GPTZero (own site) | 99% accuracy / 96.5% on mixed documents | Vendor-stated overall accuracy claims |
| Originality.ai (own site) | 97.8% accuracy claimed | Vendor-stated, multilingual model; specific FPR not published on that page |
| Liang et al., arXiv 2304.02819 | Detectors studied were biased against non-native English writers | Independent research finding, not a single percentage |
“Every one of these numbers is true, and none of them are interchangeable. A document-level rate and a sentence-level rate answer different questions, even from the same company.”
Common pitfalls
- Quoting one vendor's headline accuracy number as if it applies to every kind of document and every kind of writer.
- Ignoring the difference between document-level and sentence-level false positive rates when they come from the very same source.
- Treating a vendor's own published figure as equivalent to independent, peer-reviewed research, when the two carry different weight.
- Forgetting that non-native English writers have been shown, in independent research, to be flagged more often, a bias a single 'accuracy' number hides.
A detected mark is not proof of authorship, and an absent mark is not proof of human authorship. WatermarkRemoverPro's on-device rewrite can reduce detectable evidence but cannot guarantee defeating a vendor's undisclosed watermark, on any tier.
On WatermarkRemoverPro