Updated: August 4, 2026

Best AI Detectors: Accuracy, Bulk Cost and Data Retention Compared

Search best ai detectors and three of the four articles you land on are published by companies that sell AI detectors. Scribbr ranks its own premium detector first and its

Search best ai detectors and three of the four articles you land on are published by companies that sell AI detectors. Scribbr ranks its own premium detector first and its free detector third. GPTZero and YouScan each put themselves at number one. Only Cybernews has nothing in the race.

Now look at what those tests found.

ToolScribbr’s own testCybernews (independent)Vendor’s own claim
GPTZero52%, 10th of 1270%About 99%
Scribbr84%, 1st in its own test50%, 11th of 1284%
Sapling68%10%Not published
QuillBot78%80%Not published
Originality.AI76%70%76 to 94%

Sapling moves 58 points between Scribbr’s own test and Cybernews’. GPTZero’s own number is close to double the 52% it scored in the most methodical test here, the one Scribbr ran and won. These results are not reproducible, so a single accuracy percentage cannot justify a purchase.

We sell no AI detector, and no vendor paid for a place in this ranking.

The evidence, ranked by how much weight it can carry

  • Independent academic with disclosed funding: the Chicago Booth working paper by Brian Jabarian and Alex Imas (also NBER Working Paper 34223), which tested 1,992 pre-2020 human texts against 1,992 AI texts across six genres and four frontier LLMs. Funded by Booth’s Center for Applied AI, the Becker Friedman Institute and Google Cloud research credits. No vendor money, no declared conflicts. Alongside it, Russell, Karpinska and Iyyer (Maryland, UMass Amherst and Microsoft, NSF and Open Philanthropy funded) on humanized text.
  • Independent but not academic: Cybernews and PCWorld. Real tests, small samples, no peer review.
  • Vendor self-reports: marketing, until an outside party reproduces them.

What we ranked on

  • False-positive rate, and who measured it. A percentage with no named tester behind it decides nothing.
  • Behaviour on human-edited AI. Your writers edit drafts. That is the workflow detection has to survive.
  • Bulk and API capability, with a real cost per 1,000 words. Checking one document is not your problem. Checking four hundred is.
  • Data retention. You are uploading a client’s unpublished copy to a third party.

We assign no star ratings. Scoring tools on numbers we just called unreproducible would be incoherent.

ToolBest forIndependent evidenceFalse-positive postureAPI and bulkCost per 1,000 words
Originality.AIAgencies scanning whole client sitesOne academic study plus two testsBelow 1% (Chicago Booth)Enterprise tier only$0.65 to $0.91
GPTZeroTeams that need 250-file batchesOne academic study plus three testsBelow 1% (Chicago Booth)Yes, 250-file batchNot published
CopyleaksTeams of 10 to 25 needing bundled seatsNone publishedUndocumented outside the companyYes, pricing not publishedAbout $0.30
Winston AISmall teams wanting unlimited seatsNone publishedNo published figureYes, from the entry tierAbout $0.05 to $0.10
QuetextModelling a predictable volume billNone publishedNo published figureYes, at professional tier$0.045 to $0.16
QuillBotOne editor spot-checking a submissionTwo non-academic testsNo published figureNo published pathFlat, unmetered
SaplingLarge operations needing SSO and self-hostingTwo non-academic tests, 58 points apartNo published figureMetered APINot published
TurnitinCannot be bought outside an institutionNoneUnder 1%, the vendor’s own claimInstitutional onlyNot sold to teams

The order below reflects evidence quality and fit for a team buying checks at volume. It is not a composite score, because no honest composite exists.

1. Originality.AI: Best Workflow for Agencies, With a Privacy Trade

originality ai homepage

Every other tool here assumes you are checking one document. Originality.AI assumes you are checking a client’s entire site, which is why agencies keep landing on it. Full-site scans, custom tagging, team management, a Chrome extension, and 365-day scan history on Enterprise.

  • Pricing: Pro runs on credits at 1 credit per 100 words. Enterprise lists at $179 per month, or $136.58 per month annualised.
  • Cost per 1,000 words: $0.65 on Pro, $0.91 on Enterprise.
  • API and bulk: full-site scanning, team seats and tagging. API access is gated to Enterprise and is not available on Pro.
  • Data handling: absent an opt-out, submitted content may be retained and used in anonymised form for testing and training. School and student customers are excluded from training data by default. SOC 2 covers the consumer tier.

The evidence splits cleanly. Chicago Booth measured its false-positive rate below 1%, which is genuinely good: it clears human writing without much drama. The same paper measured false negatives of 10 to 40% depending on which model generated the text, so it misses AI writing far more often than GPTZero did on that corpus.

Cybernews’ independent test scored it 70%. Scribbr’s own test, the one that placed Scribbr’s premium detector first, scored it 76%. Originality.AI’s marketing range of 76 to 94% sits above every independent measurement of it.

Then the privacy trade. Originality.AI’s default is training on anonymised submissions unless you actively opt out. If you upload a client’s unpublished copy, that is a question for your legal team before the first scan, not a footnote in the terms.

This is also the tool that flagged Kimberly Gasuras, a news reporter of 24 years, on the WritersAccess platform. The FAQ below covers what that cost her.

Against GPTZero: Originality.AI wins on workflow by a distance and loses on false negatives and API access. If your bottleneck is scanning a 400-page client site, take the workflow. If your bottleneck is trusting a single flag, take GPTZero, provided you can live with text retention by default.

2. GPTZero: Strong Independent Results, Weak Marketing Discipline

gptzero homepage

When an independent academic paper rated a rival detector above it, GPTZero published a reanalysis of that paper’s data on its own blog. The post claims 99.5% for GPTZero against 99.1% for the tool that beat it, and calls that “40% fewer errors”. The researchers, it says, failed to use “the correct field in our API”.

The post links the paper. It never names the authors, Brian Jabarian and Alex Imas.

A vendor re-litigating independent research in its own favour is the cleanest argument for reading the studies yourself.

The tool is better than the marketing around it.

  • Pricing: published plan tiers only. No concrete API rate appears on its pricing pages.
  • Cost per 1,000 words: not published, so ask for a written per-word quote before comparing it against Originality.AI or Quetext.
  • API and bulk: batch upload of up to 250 files on the Professional tier, plus an API with sample code in 17 languages.
  • Data handling: submitted text is stored by default and may be retained to improve its classifiers. SOC 2 Type II with continuous monitoring via Drata, and stated FERPA and GDPR alignment.

Credit where it is earned. The Chicago Booth paper measured GPTZero’s false-positive rate below 1%, 96% accuracy on short passages, and false negatives of 0 to 2%, the strongest result on that corpus of any tool ranked here. Against that: 52% in Scribbr’s own test (10th of 12), 70% in Cybernews’ and 62% in PCWorld’s. On lightly human-edited AI, Cybernews measured 9%.

The 99% figure deserves its own paragraph. RAID is a real benchmark, from Dugan et al. at ACL 2024, covering more than 10 million generations across 11 LLMs and 12 adversarial attacks. Its public leaderboard and README list no commercial vendors, GPTZero included.

The 99% comes from GPTZero’s own blog, on a subset it filtered to “modern LLMs like GPT-4”. Treat it as marketing.

The verdict: a good detector and an unreliable narrator. Shortlist it only if you can live with text retention by default.

3. Copyleaks: Most Seats Per Dollar, No Independent Accuracy Data

Copyleaks homepage

Twenty-five team seats bundled into one plan at $74.99 per month annualised. Nothing else here comes close to that ratio, and for a mid-sized editorial team that is the difference between a line item and a budget conversation.

  • Pricing: Pro at $74.99 per month annualised, covering 250,000 words per month and 25 team seats.
  • Cost per 1,000 words: about $0.30 at the Pro tier, computed from the plan’s own word allowance.
  • API and bulk: website scanning and cross-language detection. API pricing is not published, and the enterprise and education tiers route through sales.
  • Data handling: uploaded text is stored as Customer Data and can be used to improve Copyleaks’ systems unless you opt out. Holds PCI-DSS, SOC 2 and SOC 3 certifications, and states GDPR compliance.

Copyleaks does not appear in the scored results of any independent test covered in this article, and it is not among the tools evaluated in the Chicago Booth paper. Its false-positive posture is undocumented by anyone outside the company.

That gap is yours to close. Build a blind set of 20 pieces whose provenance you already know, ten human and ten AI, and push it through a trial before you commit to an annual contract.

Where it is weak

  • No published API pricing, which makes cost-per-volume planning harder than with Originality.AI or Quetext.
  • Default data use for system improvement, the same posture as Originality.AI and weaker than Winston AI’s stated consent requirement.
  • Certification breadth is a security claim, not an accuracy claim. Do not let a SOC 3 badge stand in for detection evidence.

Best for: teams of 10 to 25 who want plagiarism and AI checking under one seat.

Skip if: your buying criterion is documented false-positive performance.

4. Winston AI: Cheapest Way to Give a Whole Team Access

winston ai homepage

Per-seat pricing punishes the workflow you actually want: every editor checking every submission. Winston AI’s Elite tier takes the seat count out of the equation.

  • Pricing: Essential at $10 per month for 100,000 words, or $120 per year. Elite at $26 per month for 500,000 words with unlimited team members.
  • Cost per 1,000 words: about $0.10 on Essential and about $0.05 on Elite, computed from the published word allowances. Elite sits at the bottom of the disclosed range here, alongside Quetext at professional volume.
  • API and bulk: API access from the Essential tier upward, covering text, image and plagiarism detection. Scans are capped at 200,000 characters.
  • Data handling: states it neither trains on individually identifiable submitted content without explicit consent nor sells or licenses it. Content is retained while the report exists in your dashboard, deleted within 30 days of report deletion, with encrypted backups persisting 90 days.

No independent study covered here evaluates Winston AI, and it appears in neither the Cybernews nor Scribbr scored results. Unit economics are the reason to shortlist it. Nobody outside Winston has measured the accuracy, so the trial period has to.

Where it is weak

  • A default retention window measured in months, with no zero-retention option published by anyone here. That matters if your client contracts carry a training-data exclusion clause.
  • No published false-positive figure from any party, vendor included.
  • The 200,000 character scan cap forces splitting on long-form and full-site work.

Against Copyleaks: Winston is cheaper per word and better if you want everyone checking without counting seats. Copyleaks is stronger on certifications and bundled plagiarism depth. A five-person team on a tight budget takes Winston. A 20-person team facing a procurement checklist takes Copyleaks.

5. Quetext: Clearest Published Price at Volume

Quetext publishes something almost nobody else here does: a rate that keeps scaling. $4.49 per 100,000 words per month for detector-only usage, all the way up to 14.9 million words a month. Build your spend model on that.

  • Pricing: $7.99 per month for 50,000 words on the detector-only plan. Professional volume pricing runs at $4.49 per 100,000 words per month.
  • Cost per 1,000 words: about $0.16 at the entry tier, falling to about $0.045 at professional volume.
  • API and bulk: bulk upload of up to 100 files, plus API access, both included at the professional tier.
  • Data handling: not documented in enough detail to compare against Winston AI’s stated consent requirement. Ask for the data processing addendum before you upload client work.

No independent test covered here includes Quetext in its scored results, and it does not appear in the Chicago Booth paper. Trial it against a blind set alongside whatever you already use, then read the disagreements rather than the headline scores. Two tools splitting on the same piece tells you more than either number alone.

Where it is weak

  • Built as a plagiarism checker first, with AI detection as a secondary feature.
  • Sells an AI humanizer alongside the detector, the same dual incentive QuillBot carries.
  • Zero third-party accuracy data.

The verdict: the right shortlist entry when your constraint is a predictable per-word bill, and the wrong one when your constraint is defensible evidence.

6. QuillBot: Fine for Spot Checks, Sold Alongside the Bypass

QuillBot homepage

Both independent tests put QuillBot in the upper half, and nothing else here managed that: 78% in Scribbr’s own test, 80% in Cybernews’. Then the result on lightly human-edited AI in that same Cybernews test: 0%, the steepest collapse recorded.

  • Pricing: Premium at roughly $8.33 per month annualised, bundling unlimited AI content detection with paraphrasing, grammar checking and summarising.
  • Cost per 1,000 words: effectively flat, since detection is unmetered inside the subscription.
  • API and bulk: no standalone API pricing or batch limits published. A team plan exists with usage dashboards, centralised billing and data controls, priced on request.
  • Data handling: team plans advertise data-control options. If you are uploading client copy on an individual plan, those are the terms to read.

The same subscription that detects AI writing also includes unlimited AI humanising. One company sells the test and the way around the test. For a content team, that changes what a clean result means: the writer who submitted the piece may hold the same subscription you are checking it with.

Where it is weak

  • Not built for agency-scale checking. Detection rides along inside a writing assistant.
  • No published bulk or API path, so it cannot sit in a pipeline.
  • The 0% result on edited AI makes it unusable for the workflow most teams are actually trying to police.

Best for: an individual editor spot-checking a single submission on a small budget.

Skip if: your writers are permitted to edit AI drafts, or you need a record you could show a client.

7. Sapling: Built for Large Operations, With the Category’s Widest Accuracy Swing

Sapling scored 68% in Scribbr’s own test and 10% in Cybernews’. Fifty-eight points on the same product. That is why these percentages cannot buy anything.

  • Pricing: Enterprise from $15 per seat per month with a 10-seat minimum. Individual Pro at $25 per month, or $12 per month billed annually, with detection as one feature of the plan.
  • Cost per 1,000 words: not published, and no word or credit limits are published either. Get a metered API quote in writing before you compare it on cost.
  • API and bulk: a separate usage-based metered API plan. Enterprise adds team analytics, domain administration, bulk provisioning, SSO and SCIM, plus self-hosting.
  • Data handling: self-hosting is the control that matters. If client copy must never leave your own infrastructure, Sapling is one of the few tools that answers that.

One genuine credit: Sapling names current models in its claims, including GPT-5, Claude 4.5 and Gemini 2.5. A marketing claim, not a measurement, but more current than the corpus behind the most-cited independent test here.

Where it is weak

  • The widest independent-score swing in the category, with no explanation offered by anyone.
  • The worst procurement transparency of any tool ranked here.
  • Detection is a feature of a writing assistant, not a scoped product with published limits.

Against GPTZero: Sapling wins on admin controls and self-hosting, GPTZero wins on documented independent false-positive performance. If your blocker is IT and data residency, take Sapling. If it is defending a flag to a freelancer, take GPTZero.

8. Turnitin: The One You Cannot Buy, and the One Its Own Customers Are Switching Off

Curtin University disabled Turnitin’s AI-writing detection across all campuses and study periods effective 1 January 2026. The University of Queensland did the same in mid-2025, calling the feature flawed and unreliable. Both kept standard text-matching plagiarism checks running, and that distinction carries the point: paying institutional customers are switching off AI-origin detection specifically, not plagiarism detection.

So why include it? Because Turnitin is sold only through institutional licensing, and a content team, agency or publisher cannot buy it at all. Its appearance in general “best AI detector” rankings, including those published by GPTZero and YouScan, two companies that sell detectors, is a wasted slot for the reader those rankings claim to serve.

  • Pricing: institutional licensing only. Not available to agencies, publishers or in-house content teams.
  • Vendor claim: under 1% false positives, per Turnitin’s own marketing, which is a self-report rather than an independent measurement.
  • Independent evidence: none of the four competing tests reviewed for this article include Turnitin in their scored results.

Dr Mark A. Bassett of Charles Sturt University, cited in the academic reasoning around Curtin’s decision, has publicly called AI-detection technology “deeply flawed”.

The verdict: the detector with the deepest institutional footprint is one you cannot license, and one some of its own customers are switching off, which is worth remembering the next time a vendor offers institutional pedigree as reassurance.

Frequently Asked Questions

How accurate are AI detectors?

There is no reproducible answer, which is itself the answer. The same tools land 30 to 50 points apart in the table above, depending on who ran the test.

The best current independent evidence is the Chicago Booth working paper by Jabarian and Imas: 1,992 pre-2020 human texts against 1,992 AI texts, six genres, four frontier LLMs, non-vendor funding disclosed. GPTZero handled false positives well there; Originality.AI’s false negatives ran 10 to 40%. The same paper found open-source detectors misclassifying up to 78% of genuinely human text, with RoBERTa, the baseline, called unsuitable for high-stakes use.

Can AI detectors be wrong about a human writer?

Yes, and it has already cost people work. Gizmodo reported the case of Kimberly Gasuras, a news reporter of 24 years, suspended from WritersAccess after an Originality.AI flag. In the same reporting, an Ohio copywriter lost an estimated 90% of his income when a three-year client cut ties over one 95% score. A writer publishing as Michael Berben was fired despite supplying full Google Docs revision history.

The structural finding is worse. Stanford’s Liang, Yuksekgonul, Mao, Wu and Zou (Patterns, 2023) ran seven commercial detectors over TOEFL essays written under supervised exam conditions where AI use was impossible, and found false-positive rates above 61%, against near-perfect accuracy on US eighth-grade native-speaker essays in the same test.

The mechanism is perplexity: detectors flag low lexical variability, and non-native writing is systematically less varied. A property of the method, not a bug awaiting a patch. If you commission from non-native English writers, that is contractual and reputational exposure that no model update will remove.

Two voices worth carrying. Debora Weber-Wulff of HTW Berlin: “AI detection companies are in the business of selling snake oil.” And Bars Juhasz, co-founder of Undetectable AI, whose business depends on detectors existing, who calls near-99% accuracy claims impossible, with livelihoods at stake.

Do AI detectors work on Claude and Gemini output?

Yes, and in the one independent test that covered them, Claude output was caught more reliably than ChatGPT. Cybernews measured Originality.AI and GPTZero at 100% on Claude, with Perplexity output close behind. The Chicago Booth paper generated its AI corpus from four frontier models rather than one.

The caveat is the test date, not the model. Scribbr’s test, the most-cited here, still covers only GPT-3.5 and GPT-4 despite a July 2026 revision. Ask a vendor which models sit in its current evaluation set, and when that set was last rebuilt.

Can a writer beat an AI detector by editing the draft?

Usually, yes. Cybernews found detection fell to 0% for QuillBot and 9% for GPTZero on lightly human-edited AI. Russell, Karpinska and Iyyer measured Binoculars at a 6.7% true positive rate and Fast-DetectGPT at 23.3% on humanized text.

If your policy permits AI drafts that writers then edit, most of this category will miss it, and no tool ranked here publishes an independent result that survives it.

Does Google penalise AI-generated content?

No. Google’s Search Central guidance from February 2023, still current, says content is assessed on quality regardless of how it was produced. What it enforces against is scaled content abuse: automation used primarily to manipulate rankings.

The Helpful Content system folded into core ranking in March 2024 and evaluates at site level. Buy detection to verify what a writer delivered and disclosed, not as insurance against a penalty that does not exist.

What should you do when a detector flags a freelancer’s work?

Never act on one score. Set your acceptable false-positive ceiling in writing before you start; the Chicago Booth authors use 0.5% as a strict cap. Ask for the outline approval trail, revision history, and original reporting a model could not produce.

Run a second tool built on a different method, and treat disagreement as the signal to slow down. Then put three clauses in the contract: AI disclosure, a training-data exclusion covering client materials, and a liability cap tied to the fee.

Blogs

Blogs & Articles

August 4, 2026

Instantly vs Apollo: Which Cold Outreach Tool Actually Fits an Agency?

After checking both vendors’ pricing pages, help centre documentation and partner programmes against the six things agencies actually get caught by, Instantly and Apollo come…

Read More

October 16, 2025

Instantly AI Review: Is This Cold Email Tool Worth It?

Instantly AI is a cold email outreach platform designed to help businesses scale their outbound marketing campaigns. From unlimited inboxes to deliverability tools and a…

Read More

Get Started