What the studies actually found
Four results worth knowing, and one that stings
You will be asked, by a manager or by yourself, whether this actually helps. The honest answer is: measurably yes for some work, measurably no for other work, and the boundary is not where intuition puts it. Four findings are worth carrying in your head.
Writing tasks get faster and better. A 2023 experiment published in Science gave 453 college-educated professionals realistic mid-level writing jobs — press releases, short reports, analysis plans. The group with a chatbot took about 17 minutes where the control group took about 27, and independent graders scored their work higher. The most interesting part was the distribution: the gap between the weakest and the strongest writers narrowed. The people who gained most were the ones who had been struggling.
The same pattern shows up in support work. A study of over 5,000 customer-support agents at a software firm found roughly 14% more issues resolved per hour with an AI assistant — and around 34% for the newest, lowest-performing agents, with almost no measurable gain for the most experienced. The assistant was, in effect, distributing what the best agents already knew.
The gains are large inside a boundary and negative outside it. In 2023, researchers ran a controlled study with 758 consultants at Boston Consulting Group. On tasks the tool handled well, the AI group completed about 12% more tasks, about 25% faster, and were graded roughly 40% higher on quality. Then the researchers included a task deliberately designed to sit just outside the model's competence — one where a plausible-looking analysis was wrong. On that task, consultants using AI were about 19 percentage points less likely to reach the correct answer than those without it.
They called the boundary the jagged frontier, and the word jagged is the point. It is not a neat line with easy tasks inside and hard tasks outside. Two tasks that feel equally difficult to you can sit on opposite sides of it. Drafting a subtle diplomatic email: inside. Working out which of two nearly identical clauses governs: outside. Nothing in the interface tells you which one you are on.
And you cannot feel it. This is the result that should change how you work. In 2025, researchers at METR ran a randomised trial with sixteen experienced open-source developers on 246 real tasks in codebases they knew well. Before starting, the developers expected AI assistance to make them roughly a quarter faster. Afterwards, they reported that it had made them about 20% faster. Measured against the clock, they were about 19% slower.
Sit with that. Experienced people, real work, and their sense of their own speed was wrong by roughly forty percentage points — in the flattering direction.
Why the feeling lies
The mechanism is not mysterious. Waiting for a draft feels like progress in a way that staring at a blank page does not. Reading a fluent answer feels like understanding. And the time you spend correcting, re-prompting and re-reading is fragmented into thirty-second pieces that do not register as work, while the twenty minutes you would have spent writing registers as a solid block you believe you avoided.
There is a second reason, specific to expertise. On work you know deeply, your own first draft is already good. The model's first draft is average. Bringing average up to your standard can take longer than starting from your standard.
What to take from this
- Expect the largest gains where you are weakest. Writing in a second language. A document type you rarely produce. A software feature you use twice a year. The evidence is consistent on this, and it is good news for most people reading it.
- Expect the smallest gains, or losses, on your core craft. If you have written four hundred of these reports, you may be the slower half of the pair.
- Never trust your sense of speed. Time two of them. A phone stopwatch settles in ten minutes an argument that surveys cannot settle in a year.
- Assume a frontier exists and that you cannot see it. The practical form of this is a habit, not a belief: on anything consequential, ask yourself what the answer would look like if it were confidently wrong, then check that specific thing.
The number that matters is yours
None of these studies were run on your job. They tell you what kind of effect to look for and roughly how large it might be. They do not tell you whether the monthly variance report gets faster in your office.
That is why the previous lesson ended with a stopwatch. One task, ten timings, before and after, including checking. It is a trivially small piece of evidence and it beats every confident claim in this lesson, including mine, because it is about you.
A team that has measured one task honestly is ahead of a team that has adopted forty enthusiastically.
The one thing to keep
The measured gains are real, largest for the least experienced, and reverse on tasks just outside the tool's competence — which you cannot feel from inside the task, so you have to time it rather than trust your sense of speed.
Before you move on
A team lead reads that consultants using AI completed 25% more work, gives the tool to his eight analysts, and asks each of them after a month whether it made them faster. All eight say yes. What is the weakest part of this evidence?
Pick the one you would defend. Nobody sees your answer.