Why music is harder than speech
Two problems speech does not have
Long-range structure. A song has form. A verse, a chorus that returns and is recognisably the same chorus, a bridge that departs and resolves, a harmonic plan that creates and discharges tension over three minutes. Almost none of that is local. A listener who hears the chorus return is comparing it with something they heard ninety seconds earlier.
That comparison happens outside any current model's attention window. The result is the most common complaint about generated music: it sounds convincing for fifteen seconds and then wanders. Sections appear and do not return. The energy stays flat. A key change happens for no reason and is not resolved.
This is the same window limitation as everywhere else in this course, and music is the medium where it is most audible, because form is the content.
Polyphony. Speech is one voice. Music is several simultaneous independent lines that must be consistent with each other harmonically and rhythmically while remaining distinguishable. The bass and the melody have to agree about the chord and disagree about the note.
A model generating a single token stream that decodes to a mixture is not tracking parts. It is producing something that sounds like a mixture, which works locally and produces the characteristic mush when the arrangement gets busy.
The failures to listen for
- Structure that does not return. No recognisable repeat of a section.
- Endings that do not end. The track fades or stops rather than resolving, because a cadence is a structural event.
- Rhythmic drift over long durations, particularly at joins between generated chunks.
- Harmonic wandering with no destination.
- Lyrics that dissolve into vowel sounds, because intelligible sung words are the hardest case in the whole field: pitch, rhythm and phonetics simultaneously.
- Instruments that morph — a guitar becoming a synth mid-phrase, the same failure as an object changing in a video.
What improves it
Conditioning on structure. Systems that accept a melody, a chord chart or a reference track produce far more coherent results than text alone, for the same reason a depth map beats a description of composition. If your tool accepts a hummed melody or a MIDI input, that is the strongest control it has.
Generating sections separately and arranging them. Generate a verse and a chorus, then build the arrangement in a free digital audio workstation. Structure supplied by you, texture supplied by the model. This is the same division of labour as compositing in the image module, and it is what people who get good results actually do.
Short loops. Anything under about fifteen seconds is inside the reliable zone. A great deal of practical use — beds, stings, loops, transitions — lives there.
Explicit form in the prompt, where supported: "intro, verse, chorus, verse, chorus, outro" helps some systems and is ignored by others. Test yours rather than assuming.
The comparison worth making
It is instructive that speech synthesis reached convincing quality before music did, given that speech carries meaning and music does not.
The reason is exactly this. Speech is one voice, and its structure is supplied by the text you provide. The model handles perhaps a sentence of context at a time and that is enough, because the coherence of a paragraph comes from the writing.
Music has no equivalent external source of structure. Nothing supplies the form. The model has to invent it, over a duration far longer than its window, with no scaffold.
That is why the strongest results come from putting the structure in yourself, and why the systems that improve fastest are the ones that accept structural conditioning rather than the ones that generate longer.
A test that shows you the ceiling
Generate a two-minute track and listen to it twice: once from the beginning, and once starting at the ninety-second mark. If the second listen tells you nothing about where you are in the piece — no sense that this is the end rather than the middle — the model has produced texture rather than form.
Do this deliberately rather than waiting to notice it. It takes four minutes, it is the clearest demonstration of the limitation in this whole module, and after you have heard it once you will stop expecting the next release to fix it by generating longer.
The honest limitation: this is not close to solved, and length is not the fix. A model that generates six minutes instead of three produces six minutes of wandering. What would fix it is an explicit representation of musical form, and that is a research direction rather than a product feature.
The one thing to keep
Music requires structure over minutes and simultaneous independent voices, both of which fall outside a model's attention window and neither of which local plausibility supplies.
Before you move on
Why does generated music tend to wander after fifteen seconds while generated speech holds together?
Pick the one you would defend. Nobody sees your answer.