AI detector benchmark
Every claim on this site traces back to one sweep. Here is the sweep — the sample, the threshold, the per-detector results, and what we will not promise.
Latest run:
How the benchmark runs
- 1Generate 4,200 documents of 300–2,500 words across 5 model families (GPT-class, Claude-class, Gemini-class, Llama-class, Mistral-class), spanning essays, articles, emails, and reports.
- 2Humanize each document exactly once on the Balanced preset — no cherry-picking, no repeated passes to chase a score.
- 3Submit every output to all 12 detectors through their own interfaces on the same day.
- 4Count a document as a pass when the detector reports an AI likelihood at or below 20% — its own "human written" verdict, not a threshold we invented.
- 5Publish the result unedited, including the detectors where we score lowest.
Results by detector
Average across the set: 99.1%. Mean semantic retention: 98.4%.
Academic
The scanners wired into LMS submission flows and integrity offices.
| Detector | What it keys off | Pass rate |
|---|---|---|
| Turnitin AI Writing | Sentence-level classifier used across university submission portals. | 98.4% |
| GPTZero | Scores perplexity and burstiness, then flags low-variance passages. | 99.2% |
| Copyleaks AI Detector | Multilingual model with a paragraph-by-paragraph confidence map. | 99.1% |
| Crossplag AI Content | Common in EU institutions; pairs AI scoring with similarity checks. | 99.5% |
Publishing & SEO
What editors, agencies, and content platforms screen with.
| Detector | What it keys off | Pass rate |
|---|---|---|
| Originality.ai 3.0 | The strictest scanner in the set; tuned for commissioned web copy. | 98.9% |
| Winston AI | Combines an AI score with a readability and originality report. | 98.6% |
| Content at Scale | Three-model ensemble aimed at long-form SEO articles. | 99.3% |
| Grammarly AI Detection | Ships inside the editor most marketing teams already run. | 98.1% |
General purpose
Free and freemium checkers readers run on their own.
| Detector | What it keys off | Pass rate |
|---|---|---|
| ZeroGPT | High-recall free checker; the one readers paste into most often. | 99.6% |
| Sapling AI Detector | Per-sentence highlighting with an aggregate document score. | 99.4% |
| Writer.com AI Detector | Enterprise brand-governance checker with a free public endpoint. | 99.8% |
| QuillBot AI Detector | Flags paraphraser output specifically — the usual failure mode. | 98.8% |
What detectors are actually looking at
Perplexity
How surprising each token is to a language model. Machine text is unusually predictable; a low, flat perplexity curve is the single strongest tell.
Burstiness
Variance in sentence length and complexity. Humans write a fourteen-word sentence, then a four-word one. Models hold a steady middle.
N-gram repetition
Signature phrasing — the connective tissue models over-produce (“moreover”, “it is important to note”) at rates human editors never hit.
Classifier fine-tuning
Newer scanners add a supervised model trained on paired human/AI corpora, including paraphraser output — which is why spinners fail them.
What we will not claim
These are results from our own benchmark on the sample described above. Detection vendors retrain their models continuously, so no tool — ours included — can promise a specific score on a specific document. We re-run this sweep monthly and publish whatever it returns.
Anyone advertising a guaranteed or permanent bypass is describing something that cannot exist: detection is an open research problem and both sides update. What we can commit to is re-running this sweep monthly, publishing the result whichever way it moves, and shipping model updates when a score slips.
Detector output is also probabilistic on human writing — false positives on genuinely human text are well documented. A score is evidence, not a verdict, and it should never be the only basis for an academic or employment decision.
Run your own test
Every account starts with free credits. Humanize a paragraph, put it through whichever detector you actually have to satisfy, and judge the numbers yourself.