Addaly is in open beta. Things will change, and AI answers can be wrong — check anything that matters.

AI, Safety and What Goes Wrong

The failure modes of AI, stated plainly, with the numbers.

Lesson 56 of 738 min

The labour inside the machine

Automation built by hand

Every polite, safe, helpful assistant is partly a product of human labour that is not visible in the product, does not appear in the marketing, and is performed under conditions most users would not accept for themselves.

The work has three main forms.

Annotation. Labelling images, transcribing speech, marking whether text is toxic, drawing boxes around objects, ranking search results. Millions of people do this, mostly through platforms that route microtasks to a global workforce.

Preference rating. Comparing two model outputs and choosing the better one. This is the raw material of alignment training. The politeness of a commercial model is, quite literally, the aggregated aesthetic and moral judgement of the people paid to make these comparisons.

Safety annotation and content moderation. Reading and classifying the material a model must learn not to produce: violence, abuse, sexual content involving children, self-harm. Someone reads all of it.

Conditions

This is not obscure. It has been reported, litigated and studied.

In January 2023 TIME reported that workers in Nairobi, employed by an outsourcing firm, had been paid roughly US$1.32 to $2 an hour to label harmful content for an AI safety pipeline, and described sustained psychological effects; the contract was ended early. Kenyan content moderators working for social media platforms have brought litigation over conditions and mental health consequences, and courts there have allowed cases to proceed.

Academic work on crowd platforms has found median effective wages well below the minimum wage in the countries where the requesters are based, unpaid time spent searching for tasks, and pay withheld for work rejected without explanation or appeal. Contracts frequently classify workers as independent contractors with no sick pay, no notice and no route to challenge a rejection or a deactivation — which is the algorithmic management problem of the previous module, applied to the people who build the models.

Non-disclosure agreements are common, which is part of why the work is invisible.

Why this belongs in a safety course

Three reasons, none of them sentimental.

It is a working condition question with a technical consequence. Rating quality depends on rater wellbeing, training and time per item. Rushed, distressed or poorly briefed raters produce noisy preference data, and noisy preference data produces models with incoherent values. Bad labour conditions show up in the product.

It relocates the harm rather than removing it. When a model refuses to produce something disturbing, that refusal was purchased by a person who read the disturbing thing. The harm did not vanish; it moved to somebody with less power, usually in another country.

It is the clearest test of whether "AI ethics" means anything. An organisation publishing a responsible AI charter while buying labelling at two dollars an hour through three layers of subcontracting has stated a value and priced it.

What is changing

Some movement, unevenly.

Some buyers now publish supplier standards covering wages, mental health support, rotation limits for exposure to distressing content, and appeal rights. Some specialist vendors compete on labour standards. The African Content Moderators Union, formed in 2023, was the first of its kind. And synthetic preference data — models generating training signal for other models — reduces demand for some categories, while raising its own questions about models trained on their own kind.

What you can do

If you buy this work, ask specific questions: what is the effective hourly rate after unpaid search time, is there a rejection appeal process, what psychological support exists for people exposed to harmful content, and how many subcontracting layers separate you from the worker. Vendors that treat these as reasonable questions are the ones to use.

If you use these products, the useful contribution is not guilt. It is refusing the framing where this labour is invisible — asking about it when procurement decisions are made, and not repeating the claim that these systems are automatic.

And if you are considering doing this work, know that it is real skilled work, that rates vary enormously between platforms, and that the specialist vendors generally pay better than the open microtask markets.

The one thing to keep

Model safety and politeness are produced by annotation, preference rating and content moderation labour that is often low-paid, psychologically hazardous and contractually precarious — and rushed or distressed raters produce noisy preference data, so the conditions show up in the model.

Before you move on

Why is the quality of preference-rating labour a technical issue and not only an ethical one?

Pick the one you would defend. Nobody sees your answer.

No ads. No data sale. No public scores on people. Ever.

© 2026 Addaly