Most AI features are shipped on a feeling. Someone tries eight examples, the outputs read well, and the thing goes live — and nobody finds out how often it fails until customers do. This course is about the other path: building a small evaluation set by hand, deciding what to measure when the output is prose, using a model as a judge without being lied to by it, catching regressions when you edit a prompt, running an honest A/B, and holding cost and latency in the same frame as quality. It ends with the situation every team eventually hits: the model changed under you, complaints went up, and you need an answer by tomorrow.
Start the first lesson- A demo is not a measurementA result is a number, on a fixed set, next to a comparison — anything else is an anecdote.
- A hundred examples, made by handA hundred real labelled examples resolves fifteen-point differences and settles arguments; it cannot resolve three-point ones.
- What to measure when the output is proseDecompose quality into checks that are separately true or false; one blended score hides where it broke.
- LLM-as-judge, and where it misleads youA judge you have not measured against your own labels is an opinion with a decimal point.
- Regression testing a promptTrack which items flipped, not just the total; an unchanged score can hide a rewritten system.
- A/B testing a feature honestlyFix the metric, the effect you would act on, and the stop date before you start; otherwise the result is a story.
- Cost, latency, and the model that changed under youPin the model and keep a frozen set, and a forced upgrade becomes an afternoon instead of a crisis.
No ads. No data sale. No public scores on people. Ever.
© 2026 Addaly