How to Check If Text Is AI Written: A Triage Method

By DeAIze Team · ·1503 words

How to Check If Text Is AI Written: A Triage Method

To check if text is AI written, run three signals in order: a detector score you read at sentence level, source questions you ask the writer before you mention AI, and metadata checks on the file itself. Treat each as evidence, not proof. No single signal — including ours — establishes authorship, so the decision comes from how the signals combine.

What does an AI detector score actually measure?

A detector reads patterns. Sentence rhythm, word predictability, how evenly the ideas are spaced, whether the vocabulary clusters around the most likely next word. It outputs a probability estimate that the text resembles machine-generated writing, not a finding of fact about who typed it.

That distinction matters when you are the person who has to act. A score of 92% does not mean “92% certain this person cheated.” It means the text carries a lot of the statistical signal the model associates with generated prose. Human writing can carry that signal too — heavily edited drafts, formulaic genres, second-language phrasing, and text written under time pressure all push the number up. If you want the mechanics, our breakdown of how AI detection works walks through the underlying signals.

This is why sentence-level output beats a single document number. One flagged paragraph inside an otherwise ordinary essay looks very different from every sentence scoring high. The first is often a stylistic artefact. The second is worth a conversation.

Run our own numbers as a sanity check on how noisy this is. On 2026-09-13 we scored two corpora with the DeAIze detector (local ensemble, lite tier) at a 30% flag threshold. A set of 23 hand-written DeAIze guides scored a mean of 21% AI, median 22%, range 0-46% — and 5 of those 23 were still flagged at that threshold. Twelve unedited model outputs scored a mean of 33%, median 34%, range 4-46%. Human writing and raw model output overlapped heavily. That is our own measurement, not a third-party verdict, and it is the clearest argument I can give you against acting on one number.

Which source questions should you ask before mentioning AI?

Ask about the work, not the tool. The moment you say “did you use AI,” you have changed the conversation into an interrogation, and you will get a defensive answer either way. Instead, open with questions whose answers you can verify.

  1. Walk me through how this came together. Ask for the order of work: reading, notes, drafting, revising. A writer with a real process describes mess — false starts, a section they cut, a source they abandoned.
  2. Where did the argument come from? Ask what made them choose this claim over the obvious alternative. This is the question AI-assisted work struggles with most, because the choice was never made.
  3. Can you point me to the version before this one? Ask for the draft history, not the final file.
  4. What would you change if you had another week? Writers who own their work have a list ready. Writers who assembled it often do not.
  5. Which part was hardest? Specific difficulty is a strong signal of authorship. Vague difficulty is not proof of anything — some people genuinely do not remember.

Listen for substance, not confidence. A nervous writer who can explain their reasoning is in a completely different category from a fluent writer who cannot say why paragraph three exists. Anxiety is not evidence.

If you need a defensible before-and-after for your own notes, DeAIze scores each sentence and rewrites only the flagged lines, then re-scores and shows both numbers. The free tier gives you 200 words on signup, no card — enough to test the workflow on one paragraph before you decide anything. See pricing for the paid tiers.

What metadata and version history can tell you

Metadata is the quietest signal and sometimes the most useful, because it is hard to fake casually and easy to check.

CheckWhat it can showWhat it cannot show
File properties (author, created, modified)Whether the file was created in one sitting, or by a different account than the rest of the submissionWhether the words themselves were generated
Version history in Google Docs or WordA real drafting trail — long edits, cut passages, commentsAnything, if the writer drafted elsewhere and pasted in
Track changes left onWhether revision actually happenedWhether an earlier draft was machine-written
Paste behaviourLarge single insertions with no edit trailWhy the insertion happened
Formatting residueOdd heading styles, leftover markdown, straight quotes in a curly-quote documentAuthorship

A clean drafting history is mild evidence of authorship. A missing one proves nothing — plenty of honest writers compose in a notes app and paste once. Read the table as a list of things you may ask about, not a list of things that convict.

How do you combine the signals into a decision?

Sort each submission into one of three buckets. The point is to stop yourself treating a high score as a verdict.

  • Low concern. Low or mixed detector scores, a plausible process story, and a drafting trail that exists. File it and move on.
  • Worth a conversation. High detector score with most sentences flagged, weak answers on reasoning, and no drafting history. You still do not accuse. You ask for the next draft with notes, or you ask them to talk you through the argument in person.
  • Escalate on process, not on detection. Multiple conflicting signals plus a policy that already defines what counts as misrepresentation. Handle it through your institution’s existing procedure, with the detector output as one item in a file, never as the finding.

The bucket you choose should depend on what you would say if the writer asked you to justify it in writing. If the honest answer is “the tool said so,” you are not there yet.

This is the part editors and teachers find hardest, and it is worth naming plainly: the cost of a false accusation is higher than the cost of one missed case. A student who is wrongly accused in front of a class remembers it for years. A single undetected submission is a bad day. Weight your process accordingly. Our guide to false positives and what to do when you are accused is written for the other side of this desk, and it is worth reading before you send an email.

What do you do when the signals conflict?

They will conflict. That is the normal case, not the exception.

High detector score, convincing source answers. Trust the process evidence over the score. Ask for the draft history if you want a paper trail, then close it. Formulaic writing and heavy editing both produce high readings on legitimate work.

Low detector score, evasive answers. A low score is not a clearance. If the writer cannot explain their own argument, that is a writing problem you can address directly — ask for a revision with reasoning attached, which is a normal editorial request and does not require mentioning AI at all.

Everything flags, and the writer is a non-native English speaker. This is the most common false-positive pattern in the wild. Predictable phrasing and simpler sentence structures push scores up. Read our piece on why detectors flag non-native English writers before you decide anything, and weight the source conversation much more heavily.

Two detectors disagree. They often do, because they weight different signals and use different thresholds. Disagreement is information: it tells you the text sits in the ambiguous zone. Compare the sentence-level output rather than the headline numbers and see whether the same lines are flagged by both. If they are not, your evidence is weak.

The writer admits using AI for parts. Now you are in policy territory, not detection territory. Check what your institution or publication actually says about assisted drafting, then apply that rule to the specific use. Many policies permit assistance and prohibit misrepresentation of authorship. That is a different question from whether the text reads machine-like — and it is the question you should have been asking all along.

One more thing the numbers support: AI tells are not a reliable shortcut. In the same 2026-09-13 measurement, our human-written sample ran about 4 AI-tell phrases per 100 words, and the unedited model output ran about 1 per 100 words. The supposedly tell-tale constructions showed up more often in the human writing. Any checklist that tells you to count “delve” and “moreover” is measuring the wrong thing.

Write your triage steps down before the next submission arrives. A process you can show the writer is the only kind that survives a challenge — and it is the only kind that lets you act with a clear conscience when the signals genuinely line up.


Measurement note: figures in this article come from our own detector run on 2026-09-13 — n = 23 hand-written guides versus n = 12 unedited model outputs, median AI score 22% and 34% respectively. Method: /methodology/.

Part of our guide to ai detector.

Want to clean up a machine-made draft?

200 free words once you confirm your email — check, rewrite, read it back.

Get started free

Related reading

Frequently Asked Questions

How can I check if text is AI written without accusing anyone wrongly?

Use three signals together: a detector score read at sentence level, source questions about the writer's process, and metadata or version-history checks. No single one is proof. If all three point the same way, ask the writer to walk you through their reasoning before you raise AI at all. If they conflict, trust the process evidence over the score.

Is a high AI detector score proof that a student cheated?

No. Detector output is a probability estimate about how much the text resembles machine-generated writing, not a finding about authorship. Human drafts, formulaic genres and second-language phrasing all produce high readings. In our own 2026-09-13 measurement of 23 hand-written guides, 5 were still flagged at a 30% threshold.

Which AI detector should teachers use for submissions?

One that shows sentence-level output rather than a single document number, because you need to see which lines carry the signal. GPTZero, Turnitin, Copyleaks and Originality.ai all report differently, and they often disagree on the same essay. Whichever you use, treat the result as one input into a triage decision, never as the decision itself.

What metadata should I check on a submitted document?

File properties for author and creation date, version history in Google Docs or Word for a real editing trail, whether track changes were left on, and formatting residue such as leftover markdown or mismatched quote styles. A clean drafting history is mild evidence of authorship. A missing one proves nothing, since many honest writers compose elsewhere and paste once.

What should I do when two AI detectors give different results?

Treat the disagreement as information: it means the text sits in the ambiguous zone. Compare sentence-level output rather than headline scores and check whether the same lines are flagged by both. If they are not, your detection evidence is weak, and you should lean on the source conversation and metadata checks instead.