How a detector claims to work
It looks at two statistical properties. How predictable each word is given the words before it, and how much sentence structure varies across the piece. AI text tends to be more predictable and more even.
So does the writing of a careful student working in their second or third language, who chooses safe vocabulary and keeps sentences to a familiar shape.
That is the entire problem, in one sentence. The signal the detector reads is not "machine", it is "unsurprising" — and plenty of human writing is unsurprising for reasons that have nothing to do with cheating.
What the evidence actually shows
A 2023 study published in *Patterns* tested seven detectors on TOEFL essays written by non-native English speakers. More than half were flagged as AI-generated. Essays by US school students were almost all correctly cleared. When the same non-native essays were rewritten with richer vocabulary, the flags largely disappeared — which tells you exactly what was being measured.
OpenAI withdrew its own detection tool in July 2023, citing a low rate of accuracy. The company with the most access to how these models write could not build a reliable detector for their output.
Vendors publish low false-positive rates from their own testing. Independent evaluations have generally found higher rates, and higher still for second-language writers. Treat vendor figures as a claim, not a finding.
The arithmetic nobody does
Suppose a detector really does have a 1% false positive rate — better than most claims. You have 30 students. Run every piece of homework through it and you can expect a wrong flag against an innocent student roughly one class in three, week after week.
Meanwhile, running AI text through a free paraphrasing tool defeats most detectors in about eight seconds.
So the tool punishes the honest and misses the determined. That is the worst possible error profile, and no improvement in accuracy fixes its direction.
What to do instead
Collect process, not just product. Ask for the draft, the notes, the plan. Word and Google Docs both keep version history for free, and it is already on. A document that appeared in one paste at 11pm looks different from one that grew over four days.
Have the two-minute conversation, with everyone. "Talk me through why you chose this example." Do it with every student for every major piece and it is a normal part of the course, not an accusation. Do it with one student and it is an interrogation. The difference is entirely in whether it is routine.
Get a baseline. One piece of in-class writing early in the term, on paper, tells you what each student sounds like. Not as evidence to convict — as the thing that makes you notice, honestly, when something reads oddly.
Write the policy before you need it. Students should know what is permitted before they sit down, not after you are unhappy.
If you are worried about a specific piece
Do not open with the accusation, and never with a score. A detector percentage is not evidence and cannot be shown to a parent.
Open with curiosity: "This is interesting — tell me how you got to this argument." A student who wrote it will talk happily for two minutes. A student who did not will usually tell you, or make it obvious without you having to say the word.
And hold the line internally. If a colleague or a head wants to act on a detector score alone, the answer is that the tool has a known bias against second-language writers, its own makers withdrew theirs for inaccuracy, and a wrong accusation costs a child far more than a missed one costs the school.
Before you move on