Pulling a mix apart
What it does
Give a separation model a stereo mix and it returns several tracks: vocals, drums, bass, and everything else. Better models split further — guitar, piano, strings.
The mechanism is a network trained on pairs of mixes and their known stems, learning to predict a mask over the time-frequency representation for each source. It is not unmixing in a mathematical sense; the information genuinely overlaps and cannot be perfectly recovered. It is a very good estimate.
Demucs is the best-known free implementation, released under a permissive licence, and it runs on ordinary hardware. Several free web tools wrap it. Quality on modern well-produced music is high enough that separated stems are used in professional workflows.
Where the artefacts are
Knowing them tells you when separation is fit for purpose.
- Reverb tails belong to whichever source they came from and get split badly, so a vocal stem carries a ghost of the room and the instrumental carries a hole where the reverb was.
- Overlapping frequency content — a snare and a vocal consonant in the same band — produces momentary bleed in both directions.
- Cymbals smear, because their content is broadband noise, which is exactly what the model cannot attribute confidently.
- Dense mixes separate worse than sparse ones. A solo voice with piano comes out nearly perfect; a loud rock mix does not.
- Older recordings with everything on two tracks separate poorly, because the sources were never independent.
The practical test is to separate, then sum the stems back together and compare with the original. What you hear missing or added is the error, and it tells you exactly how much you can trust the stems for your purpose.
What it makes possible
Remixing and mashups — the obvious use, and the one with the rights problem below.
Practice and study. Removing the guitar to play along with, isolating the bass line to learn it, slowing a solo down. This is transformative for anyone learning an instrument and it is free.
Restoration and repair. Removing a cough, fixing a bad note, rebalancing a live recording that was mixed badly.
Cleaning dialogue. Separating speech from background music in a recorded interview where music was playing — a common problem in documentary work.
Karaoke and rehearsal tracks, made in seconds rather than sourced.
Retiming and re-editing music to picture, which is far easier with stems than with a mix.
The part people get wrong
Separating a track does not give you rights to it. This is worth saying plainly because the technical ease creates a strong intuition otherwise.
A vocal stem extracted from a commercial recording is a derivative of that recording. Using it in something you publish involves the recording copyright, the composition copyright, and usually the performer's rights, exactly as sampling the original would. That the extraction was automatic changes nothing.
The uses that are uncomplicated: your own recordings, practice at home, material you have licensed, public-domain recordings, and anything where the rights holder has granted a remix licence. The uses that are not: publishing a remix of a commercial track without a licence, whatever the platform culture suggests, and building a cloned voice from a separated vocal, which adds a personality-rights problem to the copyright one.
What separation is doing to the field
Two consequences worth noticing.
Stems have become an expected deliverable. Clients increasingly ask for them because they can be re-edited, and the assumption that a mix is final has weakened.
And it has made unauthorised voice cloning of singers dramatically easier, since a clean isolated vocal is exactly what a cloning system wants. This is the single largest driver of the artist-voice disputes covered in the next lessons, and it followed directly from a tool that was built for entirely benign reasons and released freely. That pattern — a useful capability whose main effect turns out to be somewhere its authors did not look — is worth remembering generally.
One practical tip for the uses that are uncomplicated. Separation quality improves markedly if you feed the model the best available source. A lossless file separates better than a compressed one, and a compressed file separates better than audio captured from a video platform, because each lossy step has already discarded exactly the fine detail the model uses to attribute sound to sources. If a stem comes back watery, check what you gave it before blaming the tool.
The one thing to keep
Source separation splits a finished mix into instrument stems with free tools at usable quality, which enables real work and does not change the rights position of the recording at all.
Before you move on
A separated vocal stem from a commercial track sounds clean. What is the correct rights position?
Pick the one you would defend. Nobody sees your answer.