Addaly is in open beta. Things will change, and AI answers can be wrong — check anything that matters.

Music, and whose music it learned from

Making Things With AI · lesson 7 of 8 · 8 min

How a generated song is made

Most commercial music generators are language models over audio tokens. A neural codec compresses a waveform into a stream of discrete tokens — instead of 44,100 numbers per second, perhaps 50 to 150 tokens per second. A transformer predicts the next tokens, conditioned on your style text and your lyrics. A decoder turns tokens back into sound. Some systems instead run diffusion over an audio latent, which is closer to how an image model works.

The token framing explains the artefacts you can hear. Transients go gluey, because the attack of a snare or a plucked string is exactly the fast detail a codec discards. Cymbals smear. Consonants in vocals dissolve into the mix. And structure drifts after two or three minutes, because the model has no score, only a sequence it is extending.

What it is good for

Background beds under video. Temp tracks so a client can hear pacing before you commission a composer. Mood pieces, meditation audio, loops for a game menu. Sketching an arrangement in ninety seconds to find out whether the idea works at all. Jingles nobody will hear twice.

What it is not good for

Mixing properly. Stem separation exists on most platforms, but it separates a finished mix after the fact. You do not get the tracked parts, and the separated stems carry artefacts.

Precise control. You cannot say "drop the second guitar in bar 17". You reroll and hope.

A hook that survives thirty listens. Generated music is usually pleasant and rarely memorable. Fine for a bed. Fatal for a single.

Anything where a musician's identity is the point. That is not a technical limitation and no model release will fix it.

The training-data question, answered commercially

In June 2024 the major record labels sued the two largest music generation companies, alleging their models were trained on copyrighted recordings at scale. It was set up to be the case that decided whether training on recordings without a licence is fair use.

It did not decide it. Through late 2025 those disputes moved toward settlements bundled with licensing deals: money changed hands, the companies agreed to licensed models, artist opt-ins and tighter download rules, and the underlying legal question went unanswered by any court. Separate claims from music publishers over compositions and lyrics have run on their own track.

That is worth sitting with. The largest open question in generative audio was resolved by commercial negotiation rather than by law. Which means the answer can move again with the next negotiation, and it gives you no precedent to rely on.

"In the style of" is three separate problems

Style is generally not protected by copyright anywhere. You cannot own "sounds like reggaeton", or even a particular producer's approach to drums.

A voice is protected in more and more places, as a personality or publicity right — see the previous lesson. A generated vocal that sounds like a specific living singer is the risky part, not the arrangement.

A melody is protected, and this is where people get caught. Substantial similarity is judged on the output. If the generated hook lands close enough to an existing hook, it infringes, and "a model made it" is not a defence — the model was trained on that catalogue, which makes the usual claim of independent creation hard to run.

Most services block artist names in prompts. Read that as a signal about where the companies think the risk sits, not as an assurance that you are clear.

Before you use a track commercially

Check which tier you are on. Most platforms grant commercial rights only on paid plans, and some retain rights in what free accounts generate. Check whether there is any indemnity; usually there is not. And whether you own the output at all is the next lesson, where the answer is less comfortable than the terms of service imply.

Before you move on

A studio uses a generated instrumental under an advertisement. A listener points out the hook is very close to a hit from the 1990s. What is the soundest analysis?

Pick the one you would defend. Nobody sees your answer.

No ads. No data sale. No public scores on people. Ever.

© 2026 Addaly