Yes — AI detectors are wrong often enough that a score alone should never decide a misconduct case. A Stanford study found detectors wrongly flagged around 61% of essays by non-native English speakers as AI-generated, and in 2026 the Higher Education Policy Institute stated plainly that a detector score should never be sufficient evidence of misconduct on its own.
Detectors score how predictable writing is, not whether a machine wrote it — which is why careful, formal or second-language writing gets flagged even when a human wrote every word.
Why the detectors get it wrong
Detectors don't actually detect AI. There is no watermark in a paragraph, no fingerprint left behind. What these tools measure is how predictable a piece of writing is — how closely each word follows the word a statistical model would have expected next.
Clean, formal, textbook-correct academic prose scores as machine-like on that measure, because it is exactly the kind of writing a model was trained to imitate. So the students flagged most often are frequently the ones writing most carefully: people working in a second, third or fourth language, who lean on safe, correct, well-drilled constructions rather than risking an idiom. A detector reads “safe and correct” as “not human enough.”
A detector score should never be sufficient evidence of misconduct on its own.HEPI, August 2026
Neurodivergent students hit the same problem from a different direction — consistent structure, repeated signposting and even sentence rhythm are all things a marking rubric rewards and a detector penalises.
How often is “often”?
Precisely enough that universities have started backing away from the tools. In 2026 the Office of the Independent Adjudicator upheld complaints from students — including non-native English speakers and an autistic student — who were accused on the strength of a detector score and later cleared. Some universities outside the UK have switched their detectors off entirely as too unreliable to act on. Most UK universities have not, which is the awkward middle we are currently in: the tools are known to be shaky, and they are still switched on.
A high score is not evidence of cheating, and a low score is not a clean bill of health. Both directions of error are documented. Nobody in this process should be treating the number as a verdict — including you.
What this means if you've been flagged
The tool being simply wrong is a real, documented possibility — not a long shot you're clutching at. If you wrote your essay, you are entitled to see the evidence beyond the number, to explain how you wrote it, and to have your side properly heard under your university's own misconduct procedure.
Practically: keep your drafts, your notes, your reading list and your version history. Boring proof of a boring, human process is the strongest thing you can bring to a meeting — far stronger than arguing about percentages. If your students' union runs an academic advice service, use it; that is precisely what it exists for.
For the full step-by-step, read what to do if you're wrongly accused of using AI.
SafeGrade scans your work privately and shows you what's raising flags, so you can fix it in your own words. Your essay is never stored, never used for training, never shared.
Run one free scan →No card. No catch. First scan free.