Do AI Detectors Actually Work? What the Research Says

Do AI Detectors Actually Work? What the Research Says

Loch Ness · VERIFIED OR VIBES? · AI detectors By Camille G., founder of Loch Ness · Published September 29, 2026 · Last updated September 29, 2026 Short answer: sometimes, and not reliably enough to accuse anyone. Independent studies keep finding the same pattern: detectors miss a lot of AI text, especially once it’s been…


Loch Ness · VERIFIED OR VIBES? · AI detectors

By Camille G., founder of Loch Ness · Published September 29, 2026 · Last updated September 29, 2026

Short answer: sometimes, and not reliably enough to accuse anyone. Independent studies keep finding the same pattern: detectors miss a lot of AI text, especially once it’s been edited, and they wrongly flag real human writing, most often from people writing in a second language. A detector score is a guess with a percentage attached, not evidence.

I build scoring tools myself, including the algorithm behind our creativity test, so I read these numbers with a developer’s eye. A score tells you how closely a text matches a pattern. It can’t tell you who sat at the keyboard.

Human Origin Label · For writers You wrote every word. Make that verifiable. Show us your drafts and version history once, live or in a screen recording. Readers and clients get a public page that backs you up, not a detector score. Get verified, it’s free → Not ready for a review? Get the free self-declared badge →

The claim: “An AI detector can tell whether a human or a machine wrote this”

That’s the promise behind every “AI score” you see: paste a text, get a percentage, know the truth. Schools, publishers and clients use these scores to make real decisions about real people. So it matters how often they’re right.

What the research says

The largest independent test. In 2023, a team led by Debora Weber-Wulff tested 14 detection tools, including the commercial systems Turnitin and PlagiarismCheck. Every tool scored below 80% accuracy and only five scored above 70%. The tools leaned towards calling texts human-written, and they did worse once the AI text had been edited, paraphrased or machine-translated. The authors concluded that the available tools are neither accurate nor reliable (International Journal for Educational Integrity, 2023).

To be fair to the detectors, the same study noted one exception: in the test of AI-generated documents, only Turnitin classified all of them correctly (AI Weekly summary of the study). So some tools catch raw, unedited AI text well. The problem is everything around that case.

Even OpenAI gave up on its own detector. OpenAI released an AI text classifier in January 2023. By its own figures, it correctly flagged 26% of AI-written text as likely AI, and wrongly flagged human writing 9% of the time. It was withdrawn on July 20, 2023, because of its low rate of accuracy (OpenAI).

The group that pays the price. Stanford researchers ran 91 essays written by non-native English speakers for the TOEFL exam through seven popular detectors. On average, the detectors wrongly labeled 61% of these human essays as AI-generated, and 89 of the 91 essays were flagged by at least one detector. Essays by US eighth-graders, meanwhile, were classified almost perfectly (Liang et al., Patterns, 2023).

Why they get it wrong

Many detectors rely heavily on how predictable a text is. AI models tend to pick likely words, so very predictable writing looks “machine-like”. But plenty of humans write predictably too: people writing in a second language, students following a template, anyone writing plainly on purpose. In the Stanford study, the essays every detector flagged were exactly the ones with the most predictable wording (Patterns).

The flip side is just as awkward: when researchers asked ChatGPT to rewrite those same essays with fancier vocabulary, the detectors called them human (ScienceDaily). Predictability is a style, not a fingerprint. I go deeper into how false positives happen, and who gets hit, in why AI detectors flag human writing.

Verdict: 🟡 Partly verified. Detectors can pick up raw, unedited AI text some of the time. But they miss edited AI text, they wrongly flag real writers, and they are hardest on people writing in a second language. A detector score can start a conversation. It should never end one.

If a detector says you used AI

Don’t argue with the percentage. Show your process instead: the version history of your document, your drafts, your notes. That’s evidence a score can’t match, and it’s the fastest way to settle things. Step by step: how to prove you wrote it yourself.

A better question than “what does the detector say?”

Ask how the text was made. Did the writer use AI to check spelling, to brainstorm, to draft a paragraph? Those aren’t the same thing, and most readers care about the difference. Where to draw the line honestly: writing with AI, where the line is.

And if you want readers to know your work is yours before anyone runs a detector, a label backed by a real review of your process says more than any score. I compare the options in human-written content labels for authors and bloggers.

Want someone independent to check your writing process and say so publicly? See how the Human Origin review works →

Full disclosure: I created Human Origin, a label that reviews creators’ process and publishes what was checked. I’m telling you so you can weigh what I say about it. Everything above rests on the sources below, not on my label.

Sources

Update log

September 29, 2026 — First publication.