Addaly is in open beta. Things will change, and AI answers can be wrong — check anything that matters.

Making Things With AI

Images, video, voice and music — how they work, where they break, who owns them.

Lesson 60 of 848 min

Why music is harder than speech

Two problems speech does not have

Long-range structure. A song has form. A verse, a chorus that returns and is recognisably the same chorus, a bridge that departs and resolves, a harmonic plan that creates and discharges tension over three minutes. Almost none of that is local. A listener who hears the chorus return is comparing it with something they heard ninety seconds earlier.

That comparison happens outside any current model's attention window. The result is the most common complaint about generated music: it sounds convincing for fifteen seconds and then wanders. Sections appear and do not return. The energy stays flat. A key change happens for no reason and is not resolved.

This is the same window limitation as everywhere else in this course, and music is the medium where it is most audible, because form is the content.

What the window holds, and what a song needsInside a few seconds — reliableTimbre and the texture of an arrangementA groove that sits rightOne chord change into the nextA four-bar bed, a loop, a stingAtmosphereOver three minutes — nothing is tracking itA chorus recognisable as the same chorusTension created and then dischargedA cadence that ends rather than fadesAn energy shape across the whole pieceTwo independent lines that keep agreeingSpeech reached convincing quality first because its structure comes from the text you supply. Nothingsupplies a song's form, so the model must invent it over a duration far longer than anything it canattend to. A longer generation is six minutes of wandering rather than three.
What the window holds, and what a songneedsInside a few seconds — reliableTimbre and the texture of an arrangementA groove that sits rightOne chord change into the nextA four-bar bed, a loop, a stingAtmosphereOver three minutes — nothing is tracking itA chorus recognisable as the same chorusTension created and then dischargedA cadence that ends rather than fadesAn energy shape across the whole pieceTwo independent lines that keep agreeingSpeech reached convincing quality first because itsstructure comes from the text you supply. Nothingsupplies a song's form, so the model must invent itover a duration far longer than anything it canattend to. A longer generation is six minutes ofwandering rather than three.

Polyphony. Speech is one voice. Music is several simultaneous independent lines that must be consistent with each other harmonically and rhythmically while remaining distinguishable. The bass and the melody have to agree about the chord and disagree about the note.

A model generating a single token stream that decodes to a mixture is not tracking parts. It is producing something that sounds like a mixture, which works locally and produces the characteristic mush when the arrangement gets busy.

The failures to listen for

  • Structure that does not return. No recognisable repeat of a section.
  • Endings that do not end. The track fades or stops rather than resolving, because a cadence is a structural event.
  • Rhythmic drift over long durations, particularly at joins between generated chunks.
  • Harmonic wandering with no destination.
  • Lyrics that dissolve into vowel sounds, because intelligible sung words are the hardest case in the whole field: pitch, rhythm and phonetics simultaneously.
  • Instruments that morph — a guitar becoming a synth mid-phrase, the same failure as an object changing in a video.

What improves it

Conditioning on structure. Systems that accept a melody, a chord chart or a reference track produce far more coherent results than text alone, for the same reason a depth map beats a description of composition. If your tool accepts a hummed melody or a MIDI input, that is the strongest control it has.

Generating sections separately and arranging them. Generate a verse and a chorus, then build the arrangement in a free digital audio workstation. Structure supplied by you, texture supplied by the model. This is the same division of labour as compositing in the image module, and it is what people who get good results actually do.

Short loops. Anything under about fifteen seconds is inside the reliable zone. A great deal of practical use — beds, stings, loops, transitions — lives there.

Explicit form in the prompt, where supported: "intro, verse, chorus, verse, chorus, outro" helps some systems and is ignored by others. Test yours rather than assuming.

The comparison worth making

It is instructive that speech synthesis reached convincing quality before music did, given that speech carries meaning and music does not.

The reason is exactly this. Speech is one voice, and its structure is supplied by the text you provide. The model handles perhaps a sentence of context at a time and that is enough, because the coherence of a paragraph comes from the writing.

Music has no equivalent external source of structure. Nothing supplies the form. The model has to invent it, over a duration far longer than its window, with no scaffold.

That is why the strongest results come from putting the structure in yourself, and why the systems that improve fastest are the ones that accept structural conditioning rather than the ones that generate longer.

A test that shows you the ceiling

Generate a two-minute track and listen to it twice: once from the beginning, and once starting at the ninety-second mark. If the second listen tells you nothing about where you are in the piece — no sense that this is the end rather than the middle — the model has produced texture rather than form.

Do this deliberately rather than waiting to notice it. It takes four minutes, it is the clearest demonstration of the limitation in this whole module, and after you have heard it once you will stop expecting the next release to fix it by generating longer.

The honest limitation: this is not close to solved, and length is not the fix. A model that generates six minutes instead of three produces six minutes of wandering. What would fix it is an explicit representation of musical form, and that is a research direction rather than a product feature.

The one thing to keep

Music requires structure over minutes and simultaneous independent voices, both of which fall outside a model's attention window and neither of which local plausibility supplies.

Before you move on

Why does generated music tend to wander after fifteen seconds while generated speech holds together?

Pick the one you would defend. Nobody sees your answer.

No ads. No data sale. No public scores on people. Ever.

© 2026 Addaly