Counting the four ways to be right and wrong
An accuracy figure hides the argument
"The model is 88% accurate" is the least informative sentence in applied machine learning. It combines two entirely different mistakes into one number, and it averages over groups of people who are experiencing the system very differently. Everything useful about fairness starts by pulling that number apart, and the tool for it is old, simple and does not require any software.
For a system that says yes or no, every case falls into one of four boxes. Using a loan model as the example, where "yes" means approved and the truth is whether the person would have repaid:
- True positive — approved, and would have repaid. The system was right.
- False positive — approved, and would have defaulted. The lender loses money.
- False negative — rejected, and would have repaid. The applicant loses the loan they deserved. The lender never finds out.
- True negative — rejected, and would have defaulted. Right again.
Four boxes. Now compute them separately for each group of people, and the argument becomes visible.
A worked example, with the arithmetic shown
Two groups, a thousand applicants each. The model is applied identically to both.
Group A. Approved and repaid: 630. Approved and defaulted: 70. Rejected who would have repaid: 90. Rejected who would have defaulted: 210.
Group B. Approved and repaid: 350. Approved and defaulted: 50. Rejected who would have repaid: 150. Rejected who would have defaulted: 450.
Now four rates, each answering a different question. Do the division yourself as you read; the numbers are chosen to divide cleanly.
Approval rate — of everyone who applied, who got a loan. Group A: 700 of 1,000, so 70%. Group B: 400 of 1,000, so 40%.
True positive rate, also called recall or sensitivity — of the people who would have repaid, who got approved. Group A: 630 of 720, so 87.5%. Group B: 350 of 500, so 70%. This is the rate that matters to a creditworthy applicant. Being in Group B and deserving a loan means a 30% chance of refusal, against 12.5% in Group A.
False positive rate — of the people who would have defaulted, who got approved anyway. Group A: 70 of 280, so 25%. Group B: 50 of 500, so 10%.
Precision, or positive predictive value — of the people the model approved, who actually repaid. Group A: 630 of 700, so 90%. Group B: 350 of 400, so 87.5%.
Stop and look at what just happened. On precision the two groups are almost identical: 90% against 87.5%. A lender could truthfully say "an approval means the same thing regardless of group". On the true positive rate they are not remotely identical: a qualified Group B applicant is nearly three times as likely to be turned away. Both statements describe the same four numbers.
This is not a hypothetical construction. It is the exact structure of the argument about the COMPAS recidivism tool, which the next two lessons take apart.
Which rate belongs to whom
The reason people talk past each other is that each rate answers the question of a different party.
The institution cares about precision and about the false positive rate, because those are its losses. The individual who was rejected cares about the false negative rate, because that is their harm. The regulator often looks at the approval rate, because that is what shows up as an outcome gap in the population. A defendant in a risk-scoring system cares about being wrongly labelled high risk, which is a false positive in that framing.
None of these people is confused. They are computing different fractions of the same table, and the fractions genuinely disagree.
How to do this yourself
You need three columns: the group, the model's decision, and what actually happened. Then count. In a spreadsheet this is a pivot table and takes ten minutes. In Python it is a couple of lines with pandas, free and running on any laptop:
import pandas as pd
df = pd.read_csv("decisions.csv")
tab = df.groupby(["group", "decision", "outcome"]).size()
print(tab)The hard part is never the arithmetic. It is the third column. You know what actually happened only for people the system said yes to — the rejected applicant never gets the chance to repay. That gap has a name, the selective labels problem, and it is the last lesson in this module.
One discipline before you start: decide which rate you will report before you compute all four. Otherwise you will compute four, find the one that looks best, and report that, and you will do it without noticing.
The one thing to keep
Split every decision into the four outcome boxes and compute the rates separately per group: precision can look nearly equal while the true positive rate differs by a factor of three on the same numbers.
Before you move on
Of 500 Group B applicants who would have repaid, 350 were approved. Of 400 Group B applicants approved, 350 repaid. A lender reports "our approvals are equally reliable for both groups" and a campaigner reports "qualified Group B applicants are turned away three times as often". Who is computing what?
Pick the one you would defend. Nobody sees your answer.