Best AI Detector: 8 Tools Tested on a 60-Document Corpus (2026)

Every AI detector vendor claims 98% accuracy or better. We bought access to eight of them — GPTZero, Originality.ai, Copyleaks, QuillBot, Turnitin, Sapling, Winston AI, and ZeroGPT — and ran the same 60-document corpus through each: human writing, raw model output, and edited AI text. The raw-AI numbers are excellent. The edited-AI numbers are not. And the false positives on genuine human writing are the finding that should worry everyone.

Filed 2026-07-27 Read 10 min Method How we work
In short
  • Every detector we tested catches raw, unedited AI text at 80–100%. Vendor accuracy claims are broadly honest about that scenario and misleading about every other one.
  • One pass through a mid-tier humanizer plus a light manual edit dropped catch rates to between 15% and 45% on every tool. No text detector survives determined editing.
  • False positives on human writing are the real story. Every detector except GPTZero and QuillBot flagged at least two of our twenty human documents, and most of those flags landed on writing by non-native English speakers.
  • Category winners: Originality.ai for paid accuracy, QuillBot for free checks, Sapling for API access, and Turnitin for education only if the score is treated as a conversation starter, never as evidence.
  • Text detectors stop at text. They do nothing for AI music or AI images, which run on entirely different classifier stacks.

Every roundup of the best AI detector tools opens with a vendor accuracy claim — 99%, 99.8%, "virtually zero false positives". We wanted to know what those numbers survive contact with, so we built a 60-document corpus and bought retail access to the eight detectors people actually face in 2026: GPTZero, Originality.ai, Copyleaks, QuillBot's free checker, Turnitin, Sapling, Winston AI, and ZeroGPT.

The corpus is deliberately adversarial in the way real life is. Twenty documents are genuinely human — essays, blog posts, and reports, five of them written by non-native English speakers. Twenty are raw, unedited output from GPT-5, Claude, and Gemini. Twenty are that same AI output after one pass through a mid-tier humanizer and a light manual edit — the minimum effort anyone trying to evade detection actually applies.

The verdict, in advance: every detector is very good at the easy job and poor at the hard one. Raw AI text gets caught almost every time. Edited AI text mostly walks through. And the false positives on human writing — concentrated on the non-native subset — are the finding we think matters most, because a false positive is not a rounding error when it lands on a student's academic record.

How we tested

Each of the 60 documents went through every detector at its default settings, on a paid tier where one existed. We normalised every tool's output to a simple call: flagged as AI, or passed as human, at the tool's own threshold. A detector scores a catch when it flags an AI document and a false positive when it flags a human one.

Three design choices worth stating. First, the edited-AI set is mid-effort on purpose — one humanizer pass plus fifteen minutes of manual editing per document, not the strongest tools from our humanizer benchmark, which pass at far higher rates. Second, the human set includes non-native English writing because the published research says that is where detectors fail worst, and we wanted to see it directly. Third, we tested Turnitin through institutional access, since it sells no individual tier; the rest were retail accounts.

Sample sizes are what they are: 20 documents per category is enough to rank tools and expose failure modes, not enough to quote a false-positive rate to a decimal place. We report counts, not percentages dressed up as precision.

At-a-glance: eight detectors compared

Detector Price (as of mid-2026) Free tier Claimed accuracy Raw AI caught Edited AI caught Human falsely flagged Modality
Originality.ai From ~$14.95/mo or one-time credits No 99%+ 20/20 9/20 2/20 Text
Copyleaks From ~$9.99/mo, credit-based Limited trial 99%+ 20/20 8/20 2/20 Text + source code
GPTZero Free tier; paid from ~$15/mo ~10k words/mo 99% 19/20 6/20 1/20 Text
Turnitin Institutional licence only No 98%, <1% FPR claimed 19/20 5/20 2/20 Text (LMS-integrated)
Sapling Paid from ~$25/mo Limited checks 97%+ 19/20 7/20 2/20 Text + API
Winston AI From ~$12/mo Small trial 99.98% claimed 19/20 6/20 2/20 Text + scanned docs (OCR)
QuillBot Free Effectively unlimited Not published 17/20 7/20 3/20 Text
ZeroGPT Free; paid from ~$9.99/mo Yes, capped 98% claimed 16/20 4/20 4/20 Text

Read the table columns left to right and the story tells itself: claimed accuracy is uniform and observed accuracy is not, the raw-AI column is a landslide, the edited-AI column is a collapse, and the false-positive column is where the tools genuinely differ.

The headline finding: false positives are the story

Six of the eight detectors flagged at least one genuinely human document. ZeroGPT flagged four of twenty. And the flags were not randomly distributed — across all eight tools, the large majority of false positives landed on the five documents written by non-native English speakers.

This is consistent with the research record. A widely cited Stanford study found that GPT detectors misclassified a majority of essays by non-native English writers as AI-generated, while making almost no such errors on native-speaker essays from the same prompt set. The mechanism is unglamorous: detectors score perplexity and burstiness, non-native writers tend toward simpler vocabulary and more uniform sentence structure, and the statistical shadow of careful second-language writing looks like the statistical shadow of a language model.

The consequences are not hypothetical. Students have been hauled into misconduct hearings on nothing but a detector score; some institutions, Vanderbilt most prominently, disabled Turnitin's AI indicator rather than keep adjudicating accusations they could not verify. Turnitin itself has walked its false-positive claims into ever more caveated language. When a tool's errors fall predictably on one demographic, "98% accurate" is not a defence — it is a description of who pays for the other 2%.

Our practical rule: any workflow where a detector score triggers punishment, rather than a conversation, is a workflow that will eventually punish an innocent person.

The eight detectors, reviewed

1. Originality.ai — the strongest paid all-rounder

Originality.ai caught all twenty raw-AI documents and led the field on edited AI at 9/20 — the only tool to catch more than 40% of the edited set. Two false positives on our human corpus, both on non-native writing, against a marketing claim of a sub-2% false-positive rate. It is built for publishers and agencies: credit-based scanning, a Chrome extension, site-wide scans, team seats. As of mid-2026 pricing starts around $14.95 monthly, with one-time credit packs for lighter use.

Trust it for: publisher and agency screening at volume. Distrust it for: anything punitive. Its aggression against edited AI is exactly what raises its false-positive floor.

2. Copyleaks — the enterprise pick with an honest weakness

Copyleaks matched Originality.ai on raw AI at 20/20 and came second on edited text at 8/20, with the broadest integration surface in the group — LMS plugins, an API, source-code detection, and support for dozens of languages. It also flagged two human documents, and independent benchmarks report a materially higher rate — one 2026 benchmark measured its false-positive rate at 11%. Our full Copyleaks review covers where its multilingual strength and its flag-happiness both come from.

Trust it for: organisations that need detection wired into existing systems. Distrust it for: adjudicating individual documents without human review.

3. GPTZero — the false-positive conservative

GPTZero missed one raw-AI document and only caught 6/20 of the edited set, but it flagged exactly one human document — the best false-positive performance of any paid tool we ran. That trade is deliberate: GPTZero's public positioning since the education backlash has been to tune conservative and publish its error philosophy. The free tier (around 10,000 words monthly) is genuinely useful, and the writing-analysis dashboard is the most transparent in the group, showing sentence-level highlighting instead of a bare percentage.

Trust it for: situations where a false accusation costs more than a missed detection — which is most situations involving people. Distrust it for: catching edited AI. It is the most evadeable of the top tier.

4. Turnitin — the education incumbent, on probation

Turnitin is the detector most students will actually face, purely by install base — it sits inside the LMS at most universities. On our corpus it caught 19/20 raw-AI documents, managed only 5/20 on edited text, and flagged two human documents despite a claimed sub-1% false-positive rate. Its real product is not accuracy but integration and audit trail. The controversy record — institutions disabling it, its own quietly revised claims — is significant enough that we cover it separately in our Turnitin AI detector review.

Trust it for: a first-pass signal inside an academic-integrity process with human judgment downstream. Distrust it for: being the process. A Turnitin percentage is not evidence.

5. Sapling — the developer's detector

Sapling caught 19/20 raw and 7/20 edited — quietly third-best against edited text — with two human false positives. Its differentiation is the API: clean documentation, per-request pricing, and sentence-level scores that slot into moderation pipelines without UI scraping. The web interface is spartan, which we mean as a compliment. Paid access starts around $25 monthly as of mid-2026.

Trust it for: building detection into your own product or editorial pipeline. Distrust it for: nothing unusual — it shares every category-wide limit below.

6. Winston AI — solid, oversold

Winston AI advertises 99.98% accuracy, which is the least credible claim on this page, but the tool underneath is respectable: 19/20 raw, 6/20 edited, two human false positives, plus an OCR mode that scans photographed or scanned documents — genuinely useful for handwritten-then-typed academic workflows. Pricing from around $12 monthly.

Trust it for: scanned-document workflows nobody else handles well. Distrust it for: its own marketing numbers.

7. QuillBot — the best free detector, with a conflict of interest

QuillBot's free detector caught 17/20 raw-AI documents, no sign-up and no meaningful cap — the best free option we tested. Two caveats. It caught only 7/20 of the edited set, and it flagged three human documents, two of them written by non-native English speakers — the worst false-positive count of any tool here except ZeroGPT. And QuillBot's core business is a paraphraser — the company sells both the evasion tool and the detector, an arrangement we unpack in our QuillBot AI detector review.

Trust it for: free sanity checks on unedited text. Distrust it for: anything that has been paraphrased — including by QuillBot.

ZeroGPT is one of the most-searched names in the category and posted the weakest numbers we measured: 16/20 on raw AI, 4/20 on edited, and four false positives on twenty human documents. A tool that misses a fifth of raw model output while flagging a fifth of human writing is generating noise in both directions.

Trust it for: very little beyond a free second opinion. Distrust it for: any decision that matters.

Winners by category

Best paid detector: Originality.ai. Highest combined catch rate, tolerable false-positive count, priced for professional use.

Best free detector: QuillBot. Unlimited, accurate on raw AI, minimal false positives. Just know its blind spot is anything edited.

Best for educators: Turnitin — conditionally. It wins on integration, not accuracy. The condition is a written policy that no student is sanctioned on a score alone. If your institution cannot commit to that, GPTZero's conservative tuning is the more defensible choice.

Best API: Sapling. The cleanest developer surface and quietly strong edited-text numbers. Copyleaks is the alternative if you need multilingual coverage and LMS hooks.

What every text detector cannot do

These limits are structural. No update fixes them, and any vendor implying otherwise is selling.

The output is a probability, not a proof. There is no hidden mark being read (our ChatGPT watermark breakdown covers why text watermarking remains mostly unshipped). A detector infers from statistics, and statistics have tails on both sides.

Editing defeats detection. Our mid-effort edited set dropped every tool below 50%. The strong humanizers from our humanize AI text benchmark pass at 84–92% against production detectors. Detection catches the lazy and loses to the diligent, and that asymmetry is permanent — the full mechanics are in our bypass AI detection explainer.

The errors are biased. Non-native English writers, formulaic genres, and heavily templated professional prose all score AI-like. The people most exposed to false accusation are the people least equipped to contest it.

The target moves. Every new model generation shifts the token statistics detectors are trained on. A benchmark run in July 2026 — including this one — is a measurement, not a guarantee, which is why we re-run our corpora quarterly.

Text detectors stop at text

A question we get constantly: will any of these tools tell you whether a song or an image is AI-generated? No. Text detectors score written-language statistics, and none of that transfers to audio or pixels.

AI music detection is a separate stack — spectral classifiers trained on Suno, Udio, and Stable Audio output, which we benchmarked in our AI music detector guide. AI image detection is different again: visual-artifact models plus embedded watermarks like SynthID, covered in our AI image detector and AI watermark detector guides. And the removal side of those modalities has its own tooling — for readers dealing with audio or image artifacts rather than text, our Undetectr review covers the one remover that passed our audio benchmark. None of that applies to writing; the modalities share an arms race, not a toolset.

What this means for you

If you are screening content — as a publisher, an agency, or an educator — the detectors on this page are usable as long as you use them for what they are: probabilistic signals that are excellent against raw model output, weak against edited text, and dangerous when treated as verdicts. Pick Originality.ai for volume screening, GPTZero when false accusations cost more than misses, and put a human between every flag and every consequence.

If you are being screened, understand both edges of the data. Raw AI text will be caught — the 95–100% catch rates are real. And genuinely human writing sometimes gets flagged anyway, especially if English is not your first language; if that happens to you, the research record and the institutional retreats documented above are your appeal material.

The honest summary of the best AI detectors in 2026 is that they are good tools for a narrow job, marketed as arbiters of a broad one. The narrow job — flagging unedited model output at scale — they do well. The broad one — telling you with certainty who wrote something — no statistical classifier can do, and the sooner scores stop being treated as verdicts, the fewer innocent writers will pay for the difference.

Frequently asked

Questions readers ask.

On our 60-document corpus, Originality.ai posted the strongest overall numbers: it caught every raw-AI document, held up best against edited AI text, and kept false positives lower than most rivals. Copyleaks and GPTZero were close behind, each with a different trade-off — Copyleaks caught slightly less edited text, GPTZero produced the fewest false positives of any paid tool. There is no single best answer for every user, which is why we split the verdict into category winners: best paid, best free, best for educators, and best API.

It depends entirely on what the text has been through. On raw, unedited model output, Originality.ai, Copyleaks, GPTZero, Turnitin, Sapling, and Winston AI all caught 19 or 20 of our 20 raw-AI documents — effectively tied. On AI text that had been through one humanizer pass and a light manual edit, Originality.ai led at 9 of 20 caught, and everything else fell further. Vendor claims of 99% accuracy describe the first scenario only. No detector we tested is accurate against deliberately edited text.

QuillBot's free detector is the best free option we tested: no sign-up, no meaningful word cap, and a raw-AI catch rate of 17 out of 20. Its weaknesses are edited AI text, where it caught 7 of 20 documents, and false positives — three of our twenty human documents came back flagged. ZeroGPT is also free but produced the most false positives of any tool we ran. Treat free detectors as a first-pass sanity check, not a verdict.

More often than the marketing admits. Across our 20-document human set, six of the eight detectors flagged at least one genuinely human document, and ZeroGPT flagged four. Most flags landed on the five documents written by non-native English speakers — consistent with published research from Stanford showing detectors disproportionately misclassify non-native writing as AI. A single flag on a 20-document set is a small sample, but the direction matches the documented cases of students falsely accused by institutional detectors. The false-positive risk is real and it is not evenly distributed.

Mostly no. We put each raw-AI document through one pass of a mid-tier humanizer followed by a light manual edit, and catch rates collapsed on every tool: from 95–100% on raw output to 15–45% on the edited versions. Originality.ai held up best at 9 of 20; QuillBot caught 3 of 20. This matches what we found from the other direction in our humanizer benchmark, where the two strongest humanizers pushed pass rates above 84% against five production detectors. Detection is reliable against laziness and unreliable against effort.

Not as evidence on their own. Detector output is a probability, not a proof, and the false-positive burden falls hardest on non-native English speakers and on students with formulaic writing styles. Several institutions, including Vanderbilt, publicly disabled Turnitin's AI indicator over exactly this concern. The defensible use is as a conversation starter: a flag prompts a discussion about drafts, process, and revision history, none of which a score can substitute for. Punishing a student on a detector score alone is indistinguishable, from the student's side, from being punished by a coin flip weighted against them.

No. Text detectors score the statistical properties of written language — perplexity, burstiness, token distribution — and none of that applies to audio or pixels. AI music detection runs on spectral classifiers trained on generator output, and AI image detection runs on visual-artifact models and embedded watermarks like SynthID. Each modality is a separate arms race with separate tools; our AI music detector and AI image detector guides cover those stacks with the same corpus-style testing.

The verdict, in one sentence: Undetectr.

Every detector on this page is probabilistic, beatable by editing, and hardest on non-native writers — use the scores as signals, never as verdicts. If you are on the other side of the gate, our humanize AI text benchmark names the tools that passed. Undetectr, the tool we recommend across the rest of the site, handles audio and image artifacts, not text.