How Do AI Detectors Work? What the Score Means and Why It Is Often Wrong
AI detectors do not detect AI. They measure how predictable a piece of writing is and guess. This is how the main methods work, what a percentage score actually tells you, why non-native writers get flagged, and how to use a detector without being misled by it.
Paste a paragraph into an AI detector and it hands back a number: 87% AI, 12% human, “likely generated”. The number looks like a measurement. It is not. It is a guess about how predictable the writing is, produced by a model that has never seen the person who wrote it and has no way of knowing what happened at the keyboard.
That distinction explains almost everything that goes wrong with these tools: why a non-native English speaker’s honest essay gets flagged, why a lightly edited ChatGPT draft sails through, and why two detectors can disagree about the same paragraph by 60 points. This article explains the three ways detectors work, what the score really represents, and how to use one sensibly. It is the background piece for our individual reviews of QuillBot’s detector and Copyleaks, and for our look at what UK universities actually run.
What an AI detector is measuring
A language model writes by predicting the next word, over and over, choosing words that are likely given everything before them. The output is therefore, on average, more predictable than human writing: more even sentence lengths, more common word choices, fewer odd turns of phrase.
An AI detector runs that logic backwards. It asks: if a language model had produced this text, how surprised would it be by each word? Low surprise suggests machine writing. High surprise suggests a person. Everything else is engineering around that one idea.
Two terms come up constantly, mostly because GPTZero popularised them:
- Perplexity is a measure of how surprising the text is to a model. Low perplexity means each word was easy to predict.
- Burstiness is how much that surprise varies across the piece. Human writing tends to swing between plain sentences and unusual ones. Machine writing tends to stay level.
Neither is a property of who wrote the text. They are properties of the text itself, and plenty of human writing has low perplexity and low burstiness. Legal boilerplate does. Technical documentation does. So does a careful essay written by someone using their second language, which is where the trouble starts.
The three methods detectors use
- 1. Statistical scoringMeasure perplexity and burstiness directly with a reference model. Fast and cheap. Easily fooled by editing, and biased against plain, formal writing.
- 2. Trained classifiersTrain a model on large sets of labelled human and AI text and let it learn the difference. This is what Turnitin, Copyleaks, Originality.ai and QuillBot run. Better accuracy, but only on writing that resembles the training data.
- 3. WatermarkingThe AI vendor embeds a hidden statistical pattern in the words it chooses, which a matching detector can later find. Reliable in principle, but only if the model that wrote the text applied a watermark and the detector holds the key. Editing degrades it.
Statistical scoring
The simplest detectors compute perplexity using an open language model and apply a threshold. They need no training data and run instantly, which is why so many free checkers exist. They are also the easiest to defeat. Swap a few words for less common synonyms, vary the sentence lengths, and the score moves. They also have the strongest bias against writers whose vocabulary is narrower or more formal, for reasons covered below.
Trained classifiers
The serious commercial detectors are classifiers. The vendor assembles a large corpus of text labelled human or AI, often millions of samples across many models and subjects, and trains a neural network to tell them apart. The classifier learns many subtle signals at once rather than relying on one or two statistics.
This works well when the text you test looks like the text the classifier was trained on: unedited output from a major model, in English, on a general topic, at a reasonable length. It works less well as you move away from that. New model releases, unusual subject matter, heavy editing, short passages and non-English text all reduce accuracy, and the vendor’s published figure will have been measured under the favourable conditions.
The most honest illustration of the limits came from OpenAI itself. It launched an AI text classifier in January 2023 and withdrew it six months later, saying it had a low rate of accuracy. In its own evaluation it correctly identified only 26% of AI-written text while labelling 9% of human text as AI.
Watermarking
Watermarking is different in kind. Instead of trying to spot machine writing after the fact, the model marks it at the point of generation by nudging its word choices in a pattern invisible to readers but detectable statistically. Google has published and open-sourced a text watermarking scheme, SynthID-Text, and other vendors have built their own.
The catch is coverage. A watermark only helps if the text came from a model that applied one, and the detector belongs to the same vendor or has been given the key. Text from any other model, or text that has been paraphrased, translated or substantially rewritten, carries a weakened watermark or none at all. Watermarking may eventually matter a great deal for identifying content from a specific vendor’s products. It does not make general-purpose detection reliable today.
What the percentage actually means
A score of “80% AI” is not a claim that 80% of the words were written by a machine. Depending on the tool it means one of two things, and vendors are not always clear which:
- A probability. The classifier is 80% confident the passage as a whole was machine-generated.
- A proportion. The tool has scored the text sentence by sentence and 80% of sentences crossed its threshold.
Either way, the number is calibrated on the vendor’s own test set, not on your writing, your students or your freelancers. A detector that is 95% accurate on its benchmark can be considerably worse on the particular kind of text you feed it, and you have no way of knowing by how much.
Length matters too. Most vendors quietly advise a minimum, typically a few hundred words. Below that the statistics have too little to work with, and a single unusual sentence can swing the result. A score on a two-line email or a social post is close to meaningless.
Why the errors fall where they do
Detectors make two kinds of mistake. A false positive flags human writing as AI. A false negative lets AI writing through. Both matter, but they do not fall evenly.
Measured results from published evaluations, not vendor claims.
Stanford evaluation of seven detectors on non-native English essays (Liang et al., 2023); large-scale evaluation across 805 samples, both cited in the Higher Education Policy Institute analysis of July 2026.
False positives cluster on plain writing. The Stanford study is the one to remember. Seven detectors were run on essays written for the TOEFL English test by non-native speakers, and more than 61% were classified as AI-generated, against near-perfect results on essays by native speakers. The detectors were not finding AI. They were finding a smaller vocabulary and more conventional sentence structure, which is what writing in a second language produces and also what a language model produces. One signal, two causes, and nothing in the text to separate them. The same effect catches people who write in a deliberately plain style, people with some forms of neurodivergence, and anyone producing formulaic text such as reports and procedures.
False negatives cluster on edited text. Every detector performs best on raw, unedited model output, and that is the least common thing anyone submits. Researchers at the University of Maryland showed in 2023 that running AI text through a paraphrasing tool was enough to defeat the detectors of the day, and the humanizer tools sold for exactly this purpose have only improved since. Even ordinary human editing of a draft moves the statistics towards “human”. A detector is therefore most reliable on the cases that need it least.
The base rate problem
Suppose a detector has a genuine 1% false positive rate, which would be excellent, and a university runs it across 20,000 essays a year. Two hundred honest students are flagged. If the real rate of AI misuse is low, those two hundred can easily outnumber the students correctly caught, and the flagged group will be weighted towards international students. A small error rate multiplied by a large volume of honest work is still a lot of wrong accusations. This is why the sensible position, including Turnitin’s own guidance, is that a detection score should prompt a conversation and never settle one.
How to use a detector without being misled
Detectors are not useless. They are useful for one thing: sorting a large pile of text into “look closer” and “probably fine”. Used that way, with the limits understood, they save time. Used as a verdict, they do harm.
- Treat the score as a signal, not a finding. Nothing above about 20% should surprise you on human text, and nothing below 80% on AI text.
- Test long passages. Several hundred words at least. Ignore scores on short fragments.
- Run more than one tool if the decision matters. Disagreement between detectors is itself information: it usually means the text sits in the grey zone where none of them is reliable.
- Look for corroboration that is not statistical. Version history, drafts, the writer’s notes, a conversation about the argument. Turnitin’s Authorship-style reports and Google Docs history tell you more than any percentage.
- Never make an accusation, a hiring decision or a payment decision on a score alone. If you are setting policy for a team or an institution, write that sentence into it.
- Know why you are checking. Google has been explicit that it rewards helpful content regardless of how it was produced, so for marketing teams the question is whether the writing is good and true, not whether a detector likes it. We cover that in Is AI content bad for SEO?.
Which detector should you use?
The honest answer is that the differences between the well-known tools are smaller than the difference between using any of them sensibly and using any of them badly. That said:
- QuillBot AI Detector is free, fast and sentence-level, which makes it a reasonable first look. Independent testing puts it around 80% overall, with high false positives on some human writing.
- Copyleaks is stronger on unedited AI text and sold into education, with the usual drop on edited and short content.
- Turnitin’s indicator is what most UK universities run. Students rarely see it. Our article on AI detectors in UK universities covers what it means to be flagged and what to do.
- GPTZero, Originality.ai, Winston, Scribbr and the rest each have a niche. We are working through them one at a time with the same test set, and will link the reviews here as they are published.
Frequently asked questions
Can AI detectors be 100% accurate?
No, and no credible vendor claims it in the small print. The signals a detector uses exist in human writing too, so there is always a trade-off between catching AI text and wrongly flagging people.
Why did my own writing get flagged as AI?
Most often because it is clear and consistent: even sentence lengths, common vocabulary, a formal register. Detectors read that as predictability. Writing in a second language, following a template, or writing about a technical subject all make it more likely.
Why do two detectors give different results for the same text?
They are different models trained on different data with different thresholds. Disagreement is normal, especially on text that has been edited or mixes human and AI work. It is a sign the text is in the zone where none of them is reliable.
Can a detector tell which AI wrote something?
Some claim to, and watermark-based detection can identify text from a specific vendor’s models. General-purpose detectors cannot do this reliably, and new model releases regularly change the patterns they rely on.
Does Google use AI detectors to rank content?
Google has said its focus is on the quality and helpfulness of content rather than how it was produced. There is no evidence that a detector score influences ranking. Thin, inaccurate or unoriginal content performs badly whether a person or a model wrote it.
Is it possible to make AI text undetectable?
Yes, and the tools that do it are widely sold. That is one more reason a detection score cannot be treated as proof in either direction.
In short
AI detectors estimate how predictable a piece of writing is and infer authorship from that. The inference is often right on unedited machine output and often wrong on plain, formal or second-language human writing. The score is a prompt to look closer, calibrated on someone else’s data, and it should never be the only reason for a decision about a person.
If your team is deciding how to handle AI-assisted writing, whether that is a content operation, a school or a compliance function, this is the kind of question our AI governance training is built around: what the tools can and cannot tell you, and how to write a policy that holds up when someone challenges it.
Related reading: QuillBot AI Detector: is it accurate? · Copyleaks AI detector review · Which AI detectors UK universities use · Is AI content bad for SEO?
Jon Goodey
Founder & CEO
Jon is the founder of Indexify, helping UK businesses leverage AI and data-driven strategies for marketing success. With expertise in SEO, digital PR, and AI automation, he's passionate about sharing insights that drive real results.
Related Resources
Continue Reading
- More Articles - Latest marketing insights
- Learning Hub - Free educational tracks
- Case Studies - Real client results
Our Services
- Digital PR - Earn quality backlinks
- Technical SEO - Site optimisation
- Marketing Analytics - Data-driven insights
- SEO Training - Private courses
Ready to Put These Insights Into Action?
Explore our services or get in touch to discuss your marketing goals.