Addaly is in open beta. Things will change, and AI answers can be wrong — check anything that matters.

Making Things With AI

Images, video, voice and music — how they work, where they break, who owns them.

Lesson 76 of 849 min

Was the training lawful?

What everyone agrees on

The factual base is not in dispute. Large models were trained on very large quantities of copyrighted material — images, text, recordings — collected by web scraping or from existing datasets, in most cases without individual permission and without a workable route for creators to object beforehand.

The disagreement is entirely about whether that is lawful, and it takes different forms depending on where you are.

The argument for lawfulness

Training is analysis, not reproduction of expression. The model extracts statistical regularities. What it retains is not the works but abstractions from them, and copyright protects expression rather than facts, ideas or style.

Copies made during training are incidental to a transformative process, in the same family as the intermediate copying that search engines and text-and-data-mining research have long relied on.

The output is generally not substantially similar to any particular input, and where it is, that specific output is the infringement rather than the training.

The alternative is unworkable at scale. Licensing billions of works individually is not possible, so a rule requiring it forbids the technology rather than pricing it.

The argument against

Copying happened, at enormous scale. Every training run required reproductions of the works, whatever was done with them afterwards.

Commercial substitution. A model that produces illustration in a style competes with the illustrators whose work trained it. That market effect is precisely what several fair-dealing tests weigh most heavily.

Consent was never sought even where it could have been, and opt-out schemes arrived after the training.

Scale is not an excuse. "We could not have licensed it all" is an argument about business models, not about rights, and other industries with licensing problems solved them by building licensing systems.

Where courts have got to

Enough has happened to describe the terrain, without treating any of it as settled.

Courts in the United States have reached different conclusions on different facts. A ruling found no fair use where a system reproduced editorial content to compete directly in the same market. Another found that training a language model on lawfully acquired books was transformative fair use, while holding that building a library from pirated copies was a separate and serious matter, ultimately resolved by a very large settlement. Another granted judgment to a developer on the record before it while explicitly noting that a better-evidenced case on market harm might come out differently. The pattern is that how the material was acquired and what market effect is proved are doing most of the work.

In the United Kingdom, a major image case ended with the copyright claims largely falling away — partly on evidence about where training took place — and a limited trademark finding. The UK's text-and-data-mining exception covers non-commercial research only, and a government consultation on widening it, with an opt-out for rights holders, was contested by creators and remains unresolved.

The European Union has TDM exceptions with a machine-readable opt-out for commercial use, and the AI Act requires general-purpose model providers to publish a summary of training content and to respect those reservations.

Japan has one of the broadest permissions for machine learning, with limits where the use unreasonably prejudices the rights holder.

Other jurisdictions, India among them, have not tested the question in a decided case, and the applicable exceptions were written long before any of this.

What this means for you

Three practical positions that hold regardless of how it resolves.

Your exposure is on output, not training. You did not train the model. The risk you carry is publishing something that infringes, which is the melody and memorisation problem from earlier modules.

Prefer providers who say what they trained on, and read the answer carefully rather than accepting the word "licensed".

Do not repeat either side's confident version. The honest sentence is: it depends on the country, on how the material was obtained, on what market harm can be shown, and it has not been resolved.

The reason to know the argument rather than just the conclusion is that you will be asked, by clients, by collaborators, and by people who are angry about it for reasons that are entirely legitimate. Being able to set out both cases accurately is more useful, and more respectful to the people affected, than picking one.

The one thing to keep

Whether training on copyrighted work without permission is lawful is being answered differently in different countries and differently for different facts, so there is no single answer and any confident one is a jurisdiction or a sales position.

Before you move on

Why is "training on copyrighted work is fair use" an unsafe thing to state as settled?

Pick the one you would defend. Nobody sees your answer.

No ads. No data sale. No public scores on people. Ever.

© 2026 Addaly