What your institution's policy probably says
Almost every policy lands in one of four postures:
- Prohibited for assessed work. Any AI use is misconduct.
- Permitted with disclosure. Use it, declare what you used it for.
- Permitted at some stages only. Commonly: fine for research, brainstorming and proofreading; not for drafting, analysis or generating content.
- Left to the module leader. Increasingly the most common, and the most confusing.
If yours is the fourth — and there is a good chance it is — then the university's public policy page is not your answer. The answer is in the assignment brief, and if the brief is silent, the answer is in an email you have not sent yet.
Send it. "Is it acceptable to use an AI tool to generate practice questions and to check my grammar for this essay? I want to be sure before I start." Ask before, not after. Keep the reply. A dated email from the person marking your work is worth more than any argument you can construct later.
Also know: policies are being rewritten constantly. A rule you learned in first year may not be the rule now. Check each term.
The detectors
Here is the part nobody tells students clearly.
OpenAI released an AI Text Classifier in January 2023 — a detector for its own model's output. It withdrew it that July, citing low accuracy. Its own published figures: it identified about 26% of AI-written text, and wrongly flagged about 9% of human-written text as AI. The company that built the generator could not reliably detect the generator.
A 2023 Stanford study (Liang and colleagues, published in *Patterns*) ran seven commercial detectors over TOEFL essays written by non-native English speakers, and over essays by US eighth-graders. The eighth-graders were classified correctly almost every time. More than half of the non-native speakers' essays were flagged as AI-generated. The reason is mechanical: the detectors keyed on low lexical variety and simple sentence construction. A careful second-language writer produces exactly that.
Several universities responded by turning the tools off. Vanderbilt disabled Turnitin's AI detector in 2023 and said so publicly. Turnitin has claimed roughly a 1% false positive rate at document level — and even taken at face value, across twenty thousand essays in a term, one percent is two hundred students accused of something they did not do.
Why this is not a solvable engineering problem
A detector cannot examine provenance. There is no watermark in ordinary text. What it can do is measure statistical typicality: how predictable each word is given the ones before it. Machine text is, on average, more predictable.
So is clear, plain, careful human writing. So is writing by someone with a smaller working vocabulary in the language of assessment. So is the prose of a student who was taught to write simply, which is what good teachers teach.
Everything that makes writing easy to read makes it look machine-made. That is not a bug to be patched. It is what the measurement is.
Protect yourself before you need to
Being wrongly accused is genuinely frightening, and it is a defence built in advance or not at all.
- Write in something with revision history. Google Docs and Word both keep timestamped version history. A document that grew over eleven sessions across two weeks looks nothing like one pasted in whole, and the history is inspectable.
- Keep the mess. Your outline, your crossed-out plan, the photo of the notebook page, the draft with the abandoned second argument.
- Keep the sources you actually read, with your own notes on them.
- Be ready to talk it through. In a viva-style conversation, the person who wrote it can explain why they cut the third section. That is the evidence that persuades.
Start the habit now. The record is worthless if you begin assembling it after the email arrives.
The thing that has to be said
Detectors failing is not permission.
The argument for doing your own work is not "you will be caught," because increasingly you will not be. The argument is the one from lesson two: the exam hall exists, the professional situation exists, and both are populated by people who can tell in ten minutes of conversation whether you know your subject. The detector was never the thing that was going to find out.
Before you move on