AI & Technology

Which AI Detectors Do UK Universities Actually Use, and Should They?

Turnitin is still the default across UK higher education. The evidence on accuracy is considerably worse than most students, and many staff, realise. This is what the research shows and what it means in practice.

JG
Jon Goodey
Founder & CEO
12 min read

If you are a student in the UK, the detector your work will most likely pass through is Turnitin’s AI writing indicator, embedded in the same submission system that has handled plagiarism checking for two decades. Copyleaks and GPTZero appear at some institutions, usually alongside rather than instead.

That much is straightforward. What is not straightforward, and what has become considerably harder to ignore during 2026, is whether those tools are accurate enough for the weight being placed on them.

What the research actually shows

The headline claims from detector vendors and the findings from independent evaluation are a long way apart.

Independent evaluation findings

These are measured results from published studies, not vendor accuracy claims.

  • Non-native essays wrongly flagged61%+
  • Accuracy on unmodified AI text39.5%
  • Accuracy after light obfuscation17.4%

Stanford evaluation of seven AI detectors on non-native English speaker essays; large-scale evaluation across 805 samples. Both cited in the Higher Education Policy Institute analysis published 20 July 2026.

Take the first figure slowly, because it is the one that matters most. A Stanford study testing seven AI detectors found that more than 61% of essays written by non-native English speakers were misclassified as AI-generated, against near-perfect accuracy on essays by native speakers.

The detectors are not, in that case, detecting AI. They are detecting a narrower vocabulary and more conventional sentence construction, which is what writing in a second language tends to produce and also what a language model tends to produce. Two different causes, one signal, and no way to tell them apart from the output alone.

The second and third figures cover the other failure direction. Across 805 samples, average accuracy on unmodified AI text was 39.5%, falling to 17.4% once simple obfuscation was applied. A tool that misses more than eight in ten deliberately disguised submissions while wrongly flagging six in ten honest ones from international students is not performing the job it is being used for.

Who this lands on

The distribution of harm is not random, and it compounds an existing inequity.

International students make up 24% of UK higher education enrolment and 51% of postgraduate students. They are also the group the detectors most reliably misclassify, and the group least equipped to contest an accusation: often unfamiliar with the appeals process, worried about visa consequences, and facing an allegation supported by a number that looks objective.

Of four Office of the Independent Adjudicator cases published in July 2025, three involved international or second-language students. Three were upheld or partly upheld against the universities.

Research cited in the HEPI analysis found that these students frequently use AI tools for exactly the reasons you would expect: to overcome language barriers and to navigate unfamiliar academic conventions. Using a grammar tool to make a well-understood argument comprehensible is a very different act from having a model produce the argument. Detection scores cannot distinguish between them.

What institutions are doing about it

Some have stopped. The University of Waterloo disabled its detection tooling in September 2025. Curtin University followed in January 2026. UCLA and UC San Diego moved earlier, in 2024.

In the UK, as of July 2026, no major institution has formally discontinued AI detection. What has changed is the tone of the conversation inside them. Judy Williams of Queen’s University Belfast put the objection concisely: “AI detection tools are not the solution”, and the follow-on question, “what are we actually trying to assess?”, is the one doing the real work.

The HEPI analysis in July 2026 recommended that universities suspend the use of detection scores as primary evidence in misconduct proceedings pending independent validation, move towards process-based assessment with staged submissions and oral defences, embed AI literacy in the curriculum rather than policing it at the margins, and seek joint QAA and OIA guidance making clear that a detection score alone cannot ground a disciplinary finding.

That last recommendation is the pivotal one, and it is worth being precise about it, because it is already how most well-run misconduct processes are supposed to operate.

What a detection score is and is not
  1. It is a signalA statistical estimate that text resembles machine-generated writing. Legitimately a prompt to look more closely at a submission.
  2. It is not evidenceIt cannot establish what happened. Turnitin's own guidance is that the indicator supports rather than replaces academic judgement.
  3. It is not reproducibleScores can shift between runs and between model versions. A finding of fact should not rest on a number that may change next term.
  4. It is not neutralIts errors fall disproportionately on one identifiable group of students. That is a fairness problem, not a tuning problem.

If you have been wrongly flagged

Practical, and worth knowing before it happens rather than after.

Ask what the allegation actually is. A score is not an allegation. You are entitled to know what you are said to have done, and to see the evidence being relied on.

Produce your process, not your innocence. Proving a negative is not possible. Showing your working is. Version history in Word or Google Docs, notes, outlines, library loan records, browser history, drafts sent to a friend. This is why keeping the trail matters even when you have done nothing wrong.

Ask for an oral discussion of your own work. Nothing settles authorship faster than a student explaining their argument in their own words. Most people who wrote something can talk about it fluently. Most people who did not, cannot.

Bring your students’ union. They handle these cases routinely, they know the institution’s procedures, and you should not go into a meeting alone.

Cite the evidence. The false positive research is public, peer-reviewed and directly relevant to the reliability of the evidence against you. Referring to it is entirely legitimate.

Appeal, and go to the OIA if needed. The Office of the Independent Adjudicator has upheld complaints in this exact category. Exhaust the internal process first, because the OIA requires a completion of procedures letter.

If you are setting policy

Three things, in order of impact.

Never let a score stand alone. Require corroborating evidence and a conversation with the student before any finding. If your process permits an outcome on the basis of a percentage, it will eventually produce an unjust one.

Track your flags by demographic. If your flag rate for international students is materially higher than for home students, the tool is telling you about language proficiency, not about misconduct. You cannot know this without measuring it, and very few institutions currently do.

Assess what detection cannot reach. Staged submissions, annotated drafts, oral components, in-class work, and assignments that require the student’s own context and experience. This is more work to design and it is the only durable answer, because detection accuracy is not going to improve fast enough to catch up with generation.

The wider point for anyone outside universities

The same failure mode is arriving in recruitment, where AI screening tools are used to filter candidates, and in workplaces where managers use detectors on staff writing.

The lesson generalises cleanly: a probabilistic signal with an uneven error distribution is a reasonable prompt to look more closely and an unreasonable basis for a decision about a person. Under UK law that distinction now has teeth, since the Data (Use and Access) Act 2025 brought in safeguards for solely automated decisions, including a right to human review and a right to contest the outcome.

If you are building AI into decisions about people, that principle belongs in your policy before the tool goes live. We have written separately about what belongs in a UK AI policy and about governance structures that mid-market organisations will actually use.

In short

Turnitin remains the UK default. The independent evidence on detector accuracy is poor in both directions, and the errors fall hardest on international and second-language students, who are a quarter of UK undergraduates and half of postgraduates. Institutions elsewhere have begun switching detection off. UK institutions have not, but the sector conversation moved considerably during 2026, and the direction of travel is towards assessment design rather than detection.

For anyone on the receiving end of a flag: a score is a prompt to look, not a finding of fact, and the research supporting that view is public and citable.

Sources: Higher Education Policy Institute analysis published 20 July 2026; Stanford evaluation of seven AI detectors; large-scale detector evaluation across 805 samples; Office of the Independent Adjudicator case summaries published July 2025.

JG

Jon Goodey

Founder & CEO

Jon is the founder of Indexify, helping UK businesses leverage AI and data-driven strategies for marketing success. With expertise in SEO, digital PR, and AI automation, he's passionate about sharing insights that drive real results.

Related Resources

Continue Reading

Our Services

Ready to Put These Insights Into Action?

Explore our services or get in touch to discuss your marketing goals.