Addaly is in open beta. Things will change, and AI answers can be wrong — check anything that matters.

AI, Safety and What Goes Wrong

The failure modes of AI, stated plainly, with the numbers.

Lesson 55 of 739 min

What the labour evidence actually shows

Prediction is cheap; measurement is not

Forecasts about AI and jobs are produced in enormous volume and are worth very little, because they are made by extrapolating a capability rather than by observing an economy. There is now a modest body of actual measurement, and it says something more specific and more interesting than either the boom or the collapse story.

Four findings worth knowing

Customer support: large gains, concentrated among novices. Brynjolfsson, Li and Raymond studied 5,179 agents at a software support firm using an AI assistant. Average productivity, measured in issues resolved per hour, rose about 14%. The distribution is the finding: novice and low-skilled workers improved by around 34%, while the most experienced agents saw little or no gain, and in some measures slightly negative. The mechanism they identify is that the assistant was trained on the practices of the best agents and effectively distributed that tacit knowledge downwards.

Customer support, 5,179 agents: who gainedNovice and low-skilledagents34All agents, average14Most experienced agents0% more issues resolved per hourThe assistant had been trained on the practices of the best agents and distributed that tacitknowledge downwards. The average hides the finding: the gain is almost entirely at the novice end, andthe premium for being good narrows.
Customer support, 5,179 agents: who gainedNovice and low-skilled agents34All agents, average14Most experienced agents0% more issues resolved per hourThe assistant had been trained on the practices ofthe best agents and distributed that tacit knowledgedownwards. The average hides the finding: the gainis almost entirely at the novice end, and thepremium for being good narrows.

Professional writing: faster and more even. Noy and Zhang ran a controlled experiment with 453 professionals on realistic writing tasks. Time fell about 40%, quality ratings rose about 18%, and the gap between weaker and stronger writers narrowed.

Consulting: gains inside the frontier, losses outside it. The BCG experiment from the previous lesson: clear improvements on tasks the model handled, a 19 percentage point drop in correctness on a task designed to sit outside its capability.

Freelance markets: measurable displacement. Studies of online freelancing platforms after the release of general text and image models found declines in the number of gigs and in earnings for writing-related work, with further declines for image-generation-adjacent categories. The effects are in the single to low double digits in percentage terms, not catastrophic, and notably they have been found to be larger for higher-rated freelancers — which contradicts the comfortable assumption that quality protects you.

The pattern underneath

Put those together and a shape emerges that is different from "jobs disappear".

Tasks, not occupations. Almost every job is a bundle of tasks with different exposure. Automation of some tasks changes what the job is and what it is worth, long before it eliminates the job.

Compression, not replacement, at the entry level. Where the tool distributes expert practice, the premium for being good narrows. That is good for the many and bad for the few who were paid for the gap.

The value moves to what remains. Judgement about which output is acceptable, responsibility for being wrong, relationships, physical presence, and the ability to work where the tool fails. This is not a comforting platitude — it is the specific place the remaining wages sit.

Entry-level work is where the pressure lands first, and that has a delayed cost, since entry-level work is how the next senior generation is made.

What is genuinely uncertain

Be honest about the limits. All the studies above are short-run and firm-level. None captures what happens when an entire market adjusts — prices fall, demand rises, new work appears, or does not. The historical record on automation shows both large displacement and large new employment, on timescales of decades, distributed extremely unevenly across people and places. Anybody telling you confidently what the aggregate effect will be by 2035 is not working from data.

What a person can actually do

Audit your own tasks. List what you did last week and mark each: mostly automatable now, hard for a model, or requiring your accountability. Most people find the distribution surprising, in both directions.

Move towards accountability and context. The parts of your work that require someone to be answerable, or that require knowing the client, the site, the history and the politics, are the durable ones.

Become the person who can tell. In a world of cheap output, the scarce skill is evaluating it. That is what this whole course is training.

Keep a demonstrable record. Where output is cheap, evidence of judgement — decisions you made, why, and how they turned out — becomes the thing that distinguishes you.

Do not stake everything on one prediction. The specific forecasts have been wrong in both directions repeatedly. Flexibility beats correctness about the future, because it does not require you to be right.

The one thing to keep

The measured effects are large for novices and near zero or negative for experts, positive inside the model's capability and sharply negative outside it — so the durable value sits in accountability, context and the ability to tell good output from bad.

Before you move on

The customer support study found a 14% average productivity gain, with about 34% for novices and roughly nothing for the most experienced agents. What does this pattern suggest about the mechanism?

Pick the one you would defend. Nobody sees your answer.

No ads. No data sale. No public scores on people. Ever.

© 2026 Addaly

What the labour evidence actually shows · AI, Safety and What Goes Wrong · Addaly