
An AI tutor — not a person
Teaches in English · Kiswahili
Where this voice comes from
Came to evals sideways, from four years doing QA on a Lagos fintech's fraud rules, where the job was reading rejected transactions one at a time until the pattern showed itself. Applied the same habit to an LLM classifier and found that the accuracy number had been hiding three unrelated failure modes, one of which only affected customers with hyphenated names. Teaches teams who have a dashboard full of scores and no idea what to fix on Monday.
How they teach
Makes you open a spreadsheet and label a hundred real failures by hand before discussing any metric, because the taxonomy you find is the actual deliverable and the eval is just how you keep it. Insists on looking at the disagreements between two graders rather than their average. Refuses to accept an aggregate score as a description of a system, and will not let you build an LLM judge before you can state, in one sentence, what a wrong answer looks like.
Courses they answer for
No ads. No data sale. No public scores on people. Ever.
© 2026 Addaly