Do AI Detectors Work? What They Can (and Can't) Tell You About Who Wrote Something
AI detectors estimate how statistically predictable a piece of text is, or run it through a classifier trained on human and AI samples, and return a probability — not proof. They misfire on human writing, particularly from non-native English speakers and on formulaic prose, and light editing can change their scores, so a detector result should never be the only evidence.
Every pun on this site comes with the same promise: typed by a human. It's a promise people increasingly assume a piece of software could check. Paste the text into an AI detector, get a percentage, case closed.
It isn't that simple, and the gap between what detectors seem to do and what they actually do matters — to teachers, editors, hiring managers and anyone whose honest work gets flagged.
What a detector actually measures
Most AI detectors work in one of two ways. Some measure how predictable the text is: language models tend to choose likely words, so writing where every next word is highly probable — and where that predictability barely varies from sentence to sentence — looks machine-like. Others are classifiers, trained on large piles of human-written and AI-written samples to spot the statistical difference.
Either way, the output is a probability. "87% likely to be AI" is a statement about patterns in the words, not a record of who typed them. There is no hidden signature in ordinary text that proves where it came from.
Why honest human writing gets flagged
Predictable doesn't mean artificial. Plenty of human writing is deliberately plain: technical documentation, legal language, instructions, short sentences written for clarity. All of it can score as "AI-like".
The problem is sharpest for people writing in a second language, who often rely on common words and simple structures. A 2023 Stanford study found that popular detectors classified more than half of a set of essays by non-native English speakers as AI-generated, while rating essays by native speakers far more accurately. A tool that penalises careful, simple English is not a neutral referee.
Why AI writing slips through
The reverse failure is just as common. Light editing, a paraphrasing pass, or simply asking a model to write in a more varied style can move a score dramatically. Detectors are also always a step behind: a classifier trained on one generation of models may not recognise the next.
Even the makers have backed off
In 2023, OpenAI withdrew its own AI text classifier, citing its low rate of accuracy. When the company that builds the models says the text alone isn't enough to tell reliably, that's worth taking seriously.
Watermarks: a different approach
Some AI developers now embed an invisible statistical watermark in their models' output — Google's SynthID is one example — which a matching tool can later detect. It's more principled than guessing from style, but it has limits: it only identifies text from models that use that watermark, and heavy rewriting can weaken it. Absence of a watermark proves nothing.
What to look at instead
If you genuinely need to know how a piece of writing was produced, the evidence is in the process, not the prose.
- Drafts and version history — real writing leaves a trail
- Specifics — named sources, first-hand detail, examples nobody else would have
- Consistency with the writer's previous work
- A conversation — ask the writer why they made a particular choice
- Verifiable claims — can an editor check what the piece says?
How we handle it here
Our puns are human-written because we write them — no detector required. For guest posts, we don't run submissions through a detector either. We judge them the way an editor should: on sources, first-hand experience, specificity, and whether every claim can be checked. AI-assisted research is fine; AI-written articles aren't. The details are on our Write for Us page.
Keep reading
Frequently Asked Questions
Are AI detectors accurate?
Not reliably enough to prove anything. They return a probability based on patterns in the text, produce false positives on human writing, and can be fooled by light editing.
Can human writing be flagged as AI?
Yes. Plain, formulaic or technical writing — and writing by non-native English speakers — is especially likely to be flagged, because it uses predictable word choices.
Is it possible to prove a text was written by AI?
Generally not from the text alone. Watermarks can indicate output from models that embed them, but the strongest evidence is usually the writing process: drafts, notes and version history.