Are AI Detectors Accurate? Why Two Tools Disagree

By DeAIze Team · ·1788 words

Are AI Detectors Accurate? Why Two Tools Disagree

Are AI detectors accurate? Not as a verdict. They estimate probability from patterns in your text, they disagree with each other, and they flag plenty of human writing. A score is a signal about the words on the page, not proof about who wrote them. The sentence-level view is what you can actually use.

Key takeaways

  • A detector score is a probability estimate, not proof of authorship. No tool can see who typed the words.
  • Two detectors disagree because they use different signals, training data and thresholds. Same draft, different answer.
  • Human writing gets flagged. In our own measurement on 2026-09-19, 17 of 46 hand-written DeAIze guides crossed a 30% flag threshold.
  • Sentence-level scores beat one document number. They show you which lines carry the signal.
  • You can check your own draft in minutes and fix the specific sentences that read machine-made.
  • Never rewrite someone else’s ideas to dodge a score. The argument, data and conclusions have to be yours.

What changed that made these scores harder to trust

Detection stopped being a novelty and became infrastructure. Turnitin’s AI writing indicator now appears in marking workflows at institutions that use it. GPTZero, Copyleaks, Originality.ai and Sapling all sell seats to schools, publishers, agencies and hiring teams. A number that used to be a curiosity is now attached to grades, contracts and job applications.

At the same time, the things detectors look for have spread into ordinary human writing. Predictive text finishes your sentence. Translation tools smooth your grammar. Editors run drafts through rewriting passes. A non-native English writer who learned formal academic phrasing writes in exactly the register that detectors associate with generated text.

The result is that a score now has real consequences while remaining genuinely uncertain. That mismatch is the whole problem. Nothing in the tooling changed to make the number a fact. The stakes just went up.

Why two detectors disagree on the same draft

Every detector is a model trained on a different mix of text, scoring a different set of signals, and cutting the result at a different threshold. Four things drive the disagreement.

Different training data. A detector trained heavily on one model’s output learns that model’s habits. Feed it text from a different model, or from a human who happens to share a habit, and the confidence drops or jumps for reasons that have nothing to do with authorship.

Different signals. Some tools lean on perplexity and burstiness, the statistical predictability of your word choices and the variation in sentence length. Some lean on classifier features learned end to end. Some combine both and add their own heuristics. Perplexity-based tools are sensitive to vocabulary; classifier-based tools are sensitive to structure.

Different thresholds. One vendor ships a conservative cut that flags only high-confidence passages. Another ships an aggressive cut that flags anything above a low bar. The same draft produces two different labels with no change to the text.

Different granularity. A document-level score averages everything, so a single machine-sounding paragraph can drag a human document over the line. A sentence-level score shows you where the signal actually sits.

If you want to see which of your sentences carry the signal, run the draft through the AI detector first. It scores the document and highlights each sentence, so you can argue with a specific line instead of a number you cannot inspect.

What this means for you, honestly

If you wrote the draft yourself and got a high score, you are not imagining the problem. Our own measurement on 2026-09-19 makes the point. We scored 46 hand-written DeAIze guides with our detector at a 30% flag threshold: mean AI score 27%, median 25%, range 0-55%. Seventeen of the 46 were flagged. Those are documents written by people, about their own work, with no model in the loop.

So a flag on human writing is normal, not exotic. That cuts both ways, and it is worth being straight about the second direction: unedited model output in the same run scored a mean of 35% with a range of 25-49%. The two distributions overlap. A single score cannot separate them cleanly, and any tool that claims otherwise is overselling what a probability estimate can do.

The practical consequence is that a score is evidence about the text, not about you. If you are accused, the useful move is to show your process: drafts, notes, version history, sources. If you are self-checking before submitting, the useful move is to find the specific lines that read machine-made and fix those. Both paths run through sentence-level inspection.

One more honest limit. Our measurement also found 3 AI-tell phrases per 100 words in the human sample versus 1 per 100 words in the unedited model output. Humans, in other words, used more of the stock phrases people associate with AI than the model did. Surface tells are a weak guide. Structure and rhythm matter more.

If your draft is AI-assisted rather than fully yours, the rules are different and worth reading properly. AI-assisted writing for students covers where the line sits at most institutions.

How to check your own text sentence by sentence

You do not need a subscription to do this. You need to stop reading your draft as a whole and start reading it line by line.

  1. Score the document, then ignore the document score. Open the draft in a detector that reports per-sentence output. The headline number is the least actionable thing on the screen.
  2. Sort the sentences by AI probability. The top handful are your worklist. Everything below is noise for now.
  3. Read each flagged sentence aloud. Machine-made sentences tend to be grammatically perfect, evenly weighted, and slightly empty. If you can delete the sentence without losing an idea, that is the tell.
  4. Check sentence-length variance by eye. Look at five consecutive sentences. If they are all roughly the same length and shape, that uniformity is what the detector is reacting to. Real writing is lumpy.
  5. Look for the stock moves. Openers that announce what the paragraph will do. Closers that restate it. Transition phrases doing no work. Three-item lists where two items would do.
  6. Re-score after edits. Change one thing at a time where you can, so you learn which edits move the reading and which do not.

For a deeper explanation of what the models are actually measuring, how AI detection works walks through the signals without the marketing.

How to fix the flagged sentences

Fixing is more mechanical than people expect, and it is mostly about restoring specificity.

Problem patternWhat it looks likeWhat to do instead
Uniform rhythmFive sentences of similar length and structureSplit one, merge two, start one with a conjunction
Empty framing“This section will explore the key factors”Delete it and start with the factor
Abstract nouns“implementation”, “optimisation”, “considerations”Name the thing: the file, the deadline, the person
Even hedgingEvery claim softened the same wayCommit to the claim, or cut it
Restated endingsFinal sentence repeats the firstEnd on the newest fact
Stock phrasesThe vocabulary everyone now recognisesSay it the way you would say it out loud

Here is an illustration, not a real score. Before: “It is important to consider the various factors that contribute to effective content strategy, as these elements play a crucial role in achieving optimal outcomes.” After: “Content strategy fails for boring reasons. You picked the wrong channel, or you stopped publishing in week three.” Same idea, different density. The second one has a subject, a verb and a claim.

If you are working through a longer draft, cleaning up machine-made patterns goes through this at paragraph level, and how to rework ChatGPT text covers keeping your meaning intact while you cut the padding.

One warning about the shortcut. Tools that rewrite everything, sentence by sentence, will change your meaning along with your rhythm. If a rewrite pass cannot tell you whether the meaning survived, you are trading one problem for a worse one. DeAIze’s humanizer runs a closed loop for this reason: it scores the text, rewrites only the flagged sentences, keeps a rewrite only when semantic similarity says the meaning held, then re-scores and shows you before and after. On the example paragraph on the site, a 92% AI score becomes 4%.

The limits are real. We report an average AI-score reduction of 60%+ after humanizing, with a best case of 71%, and those are our own predicted scores, not a verdict from any third-party detector. Cross-checks run against real third-party detector APIs where configured, including GPTZero, Turnitin, Copyleaks, Originality.ai and Sapling, rather than trusting an internal guess. No tool, ours included, can guarantee what another detector will say.

When a detector is right and you still have a problem

Sometimes the score is accurate and the draft is the issue. If you generated a first pass with a model and submitted it mostly unchanged, the detector is doing its job. The fix is not to disguise the text. It is to do the work the draft skipped: check the claims, add the examples only you have, cut the sections you cannot defend in a conversation.

That is also the honest boundary of this whole category. A humanizer reduces false positives on your own writing and cleans up your own AI-assisted drafts. It is not a way to pass off someone else’s work, and no serious tool should be sold that way. If you are a non-native English writer getting flagged for formal phrasing you were taught, why an AI detector flags non-native English writers is the relevant read.

Ready to see which sentences read machine-made?

Stop arguing with a single number. Get the line-by-line view and work from there.

  • Run the free AI detector — 200 words on signup, no credit card, sentence-level highlighting so you can see exactly which lines carry the signal.
  • Open the draft editor — rewrite only the flagged sentences, keep a change only when the meaning survives, then re-score and compare before and after.
  • Check pricing — pay-as-you-go credits that never expire, no subscription. Starter $20.99 for 1,000 words, Standard $44.99 for 3,000, Pro $99.99 for 10,000, Max $199.99 for 30,000.
  • Create an account and test it on the paragraph you are actually worried about. Uploads take .txt, .docx and .pdf, and rewriting a Word document preserves fonts, styles and tables.

Detection and rewriting cover English and 9 other languages, the interface ships in 10. Your text is processed for the current task only, never used for training and never resold. And keep the framing straight: a detector score is a probability estimate, not proof of authorship. Use it to find weak sentences in your own work.


Part of our guide to draft editor.

Want to clean up a machine-made draft?

200 free words once you confirm your email — check, rewrite, read it back.

Get started free

Related reading

Frequently Asked Questions

Are AI detectors accurate enough to prove someone used ChatGPT?

No. A detector produces a probability estimate from patterns in the text, and it cannot observe how the text was made. Our own measurement on 2026-09-19 flagged 17 of 46 hand-written guides at a 30% threshold, which is a lot of human writing caught. Treat any score as one signal among several, and pair it with drafts, notes and version history before drawing conclusions.

Why do GPTZero and Turnitin give different scores for the same essay?

They use different training data, different underlying signals and different flag thresholds, and they report at different granularity. One may average a whole document while another highlights passages. Nothing about your essay changed between the two runs. The disagreement is a property of the tools, not evidence that one of them caught you.

Can human writing get flagged as AI?

Yes, regularly. Formal academic phrasing, translated text, heavily edited prose and writing that follows a template all share features detectors associate with generated text. In our 2026-09-19 run, the human sample averaged 27% and the unedited model output averaged 35%, with overlapping ranges. The two groups are not cleanly separable by a single score.

How do I check my own draft without paying for a detector?

Use a detector that reports per-sentence scores and read the flagged lines aloud. Look for uniform sentence lengths, empty framing sentences, abstract nouns and restated endings. Fix the worst handful, then re-score. You can do a useful pass on a short excerpt for free, which is usually enough to see whether the problem is local or spread through the whole draft.

Will humanizing text guarantee it passes a specific detector?

No, and you should be sceptical of any tool that promises it. DeAIze reports an average AI-score reduction of 60%+ after humanizing, best case 71%, but those are our own predicted scores rather than a third-party verdict, and cross-checks against detectors like GPTZero, Turnitin, Copyleaks, Originality.ai and Sapling run where configured. Different detectors will still disagree.