Three definitions of fair, stated precisely
"Fair" is not one thing
Once you can compute the rates, the word "fair" splits into several claims that sound alike in English and are different in arithmetic. Three of them dominate the literature and the law. Each is defensible. Each implies something the others do not. You have to choose, and the choice is a value judgement wearing a mathematical costume.
Definition one: equal outcomes
Demographic parity, sometimes called statistical parity or independence. The proportion of each group receiving the favourable decision is the same. If 60% of Group A applicants are approved, 60% of Group B applicants are approved.
What it is good for: cases where the outcome itself is the thing being distributed and you doubt the ground truth. Advertising, outreach, shortlisting for interview, allocation of scarce places. If you believe the historical labels are contaminated — that "did well in the job" was itself measured by biased managers — then insisting on equal outcomes is a way of refusing to inherit that.
What it costs: if the groups genuinely differ on the outcome you are predicting, enforcing equal rates means approving less qualified members of one group and rejecting more qualified members of the other. Individuals bear that. Note also that it can be satisfied cynically: approve a random 40% of Group B and you have parity with terrible decisions.
Definition two: equal error rates
Equalised odds, or separation. Among people who would in fact repay, the approval rate is equal across groups; among people who would in fact default, the approval rate is equal too. In the language of the previous lesson: equal true positive rates and equal false positive rates.
A weaker version, equal opportunity, requires only the first: qualified people have the same chance regardless of group.
What it is good for: this matches most people's intuition about individual fairness in a system where the outcome is real. It says: if you would have repaid, your group should not change your odds of being believed. It is the definition ProPublica implicitly used when analysing COMPAS.
What it costs: it requires you to trust the labels. If "defaulted" or "reoffended" is itself measured unequally — and rearrest is a famously unequal measurement of reoffending — then equalising errors against a contaminated target simply encodes the contamination more evenly.
Definition three: the score means the same thing
Calibration by group, or sufficiency. Among everyone assigned a score of 0.7, the same proportion actually repay, whichever group they are in. A risk score of "high" carries the same real risk for everyone.
What it is good for: this is what a decision-maker needs in order to use a score sensibly at all. If "high risk" means 60% for one group and 30% for another, the number is not a number, and anyone acting on it is applying different standards without knowing it. It is the definition the makers of COMPAS used to defend the tool, and their defence on this axis was correct.
What it costs: calibration is compatible with large disparities in who is wrongly labelled. A well-calibrated score can still produce far more false positives in one group than another, and it usually does when the base rates differ.
Also on the list, briefly
Individual fairness — similar people should be treated similarly — sounds like the cleanest principle of all and hides the whole problem in the word "similar". Defining similarity requires deciding which differences are legitimate, which is the original question.
Counterfactual fairness — the decision should be unchanged if the person's group membership had been different, holding everything else fixed — is philosophically attractive and usually unmeasurable, because group membership is upstream of half the other variables. Change someone's caste and you change their school, their neighbourhood and their surname.
How to choose
A usable procedure, which will not make anyone comfortable and is better than the alternative of choosing by accident:
- Name the harm you most want to avoid. Wrongly denying an opportunity, or wrongly granting one, or a visible gap in aggregate outcomes.
- Ask whether you trust the label. If the ground truth is itself a record of unequal treatment, definitions that condition on it inherit that, and parity-style definitions become more defensible.
- Ask whether a human will read the score. If yes, calibration is close to mandatory, because a miscalibrated score silently misleads every reader.
- Write down the definition you chose, and what you are therefore giving up, and why. Dated, in a file, next to the model.
That last step is the whole professional practice. Nobody can give you a system that is fair in every sense at once, which is the subject of the next lesson.
The one thing to keep
Demographic parity equalises outcomes, equalised odds equalises errors, and calibration makes the score mean one thing — each is defensible, each implies a different sacrifice, and choosing one is a value judgement you should record.
Before you move on
A hospital builds a score that clinicians read directly to prioritise follow-up appointments. Which fairness property is closest to mandatory here, and why?
Pick the one you would defend. Nobody sees your answer.