Where the line falls
AI can help where the criterion is visible in the text and does not require knowing the student. It cannot help where the judgement is about the person — the effort, the progress, the odd but valid route to a right answer. And its failures there are not random noise. They are patterned, and the pattern falls on the same students every time.
What it does usefully
- Feedback on drafts, not grades. "Point to three places where the evidence does not support the claim being made." That is useful to a student on Tuesday in a way your marking on Friday is not.
- A second opinion against your own rubric. Mark the set yourself, have it mark the same set, then read only the disagreements. The disagreements are the information.
- Comment banks. Generate thirty comments for common problems, then choose. Faster than typing, and you still decide which one this student gets.
- Closed-form answers, with the working shown. If it cannot show why an answer is wrong, do not accept its verdict.
Three biases, and they land on the same children
Length. Longer answers score higher, largely independent of quality. This is well documented in essay scoring, for human markers as well as machines. The machine just applies it without tiring.
Fluency. A student writing in their third language, making a correct and original argument in plain sentences, scores below a fluent, empty answer. This is the one to watch, because it hits precisely the students for whom fair marking matters most.
Unusual correctness. The student who solved it a different way. A marker matching against an expected shape marks that down. You recognise it and give full credit — and then you make her show the class.
And there is a fourth thing it simply cannot see: whether this is progress. Is this the first time Kwame has used paragraphs? That is often the most important fact about the piece and it exists only in your memory of him.
A protocol that catches the problem in ten minutes
Mark five scripts yourself. Then have the AI mark the same five. Compare.
You are not checking whether it agrees with you. You are looking for the shape of its disagreement. If it consistently rewarded the longer answers, you will see that in five. If it punished the student with weaker English and a sharper argument, you will see that too. Do this once with any new tool, and again if you change how you prompt it.
The grade you cannot explain
Do not let AI assign a final grade that goes on a record without you standing behind it.
The reason is not that it is always worse than you. On some narrow tasks its agreement with trained human markers is respectable. The reason is that a parent will ask why, and "the system gave it a 6" is not an answer a teacher can give. There is no appeal path, no reasoning you can walk through, and nothing you can teach from. A grade you cannot defend is not a grade — it is a number.
Use it to find things. Use your own judgement to decide things.
Before you upload thirty essays
Student work is personal data, often about children. Putting a class set with names into a consumer chatbot may breach your school's rules or your country's law, and this varies enormously — what is routine in one country is unlawful in another.
The practical minimum: strip names before anything is uploaded, ask your school what tools are approved, and if nobody knows, treat that as a no rather than a yes. This is contested and unsettled territory. Do not let a colleague's confidence stand in for a policy.
Before you move on