Studio · 8 min read
How to split a song into stems
Stem separation takes a finished stereo mix and pulls it back apart into its parts: the vocal on one track, drums on another, bass on a third. It is the fastest way to learn a part by ear, build a backing track for rehearsal, or find out what a guitarist is actually playing under a dense chorus. This guide covers how to do it in RigVerse Studio, which settings to choose, and — the part most guides skip — what to do when the result sounds wrong.
What stem separation actually does
A finished song is a single stereo file. Every instrument has been summed together, and mathematically that information is gone — there is no hidden multitrack inside an MP3 waiting to be unzipped.
What separation does instead is estimate. A model trained on many thousands of songs has learned what a snare looks like in a spectrogram, what a voice does across frequency and time, how a bass guitar behaves. It uses that to reconstruct each part as its own audio file. The result is a best guess, and understanding that explains almost every quirk you will hear: separation is very good at things the model has heard often, and less good at anything unusual.
That is why a modern pop record with a clear lead vocal separates beautifully, and a lo-fi live recording with a saxophone bleeding into three microphones does not.
Set up a workspace
Studio organises everything by workspace, and each one keeps its own sessions and separated stems. The workspace is what makes the work durable — close the tab, come back tomorrow, and your stems are still sitting there.
- Open Studio from the left rail and choose New workspace.
- Name it after the music, not the task. "Wildwood Flower — live take" will still mean something next month; "test 2" will not.
- Add the optional note if it helps: which take, which tuning, what you are trying to work out.
Choose Core stems or Full band
RigVerse offers two separation modes, and the choice matters more than people expect.
Core stems — vocals, drums, bass, other
Four tracks: the vocal, the drum kit, the bass, and everything else swept into "other". This is the recommended mode and the right default for most work.
Use it when you want to pull a vocal for a remix, mute the bass to practise along, or strip a song down to drums and bass to hear the arrangement underneath. Because the model only has to make four decisions instead of six, each one tends to be cleaner.
Full band — adds guitar and piano
The same four, plus guitar and piano pulled out of the "other" bucket. Useful when you specifically need the guitar part isolated, or when a piano is carrying the harmony you are trying to transcribe.
The trade-off is real: every additional split is another opportunity for the model to smear one instrument across two tracks. On a track with heavy overdriven guitar and organ, "guitar" and "other" can end up sharing material and neither sounds convincing on its own.
Which files work
Studio accepts MP3, WAV, M4A, FLAC and OGG, up to 15 MB per file.
Quality of the source matters more than the format. A 320 kbps MP3 separates about as well as a WAV; a 96 kbps file that has already been through two rounds of compression does not, because the encoder has already thrown away the high-frequency detail the model uses to tell a cymbal from a consonant.
If your file is over the limit, do not re-encode it at a lower bitrate to squeeze it in — that trades exactly the information separation needs. Trim it to the section you actually want instead. A cleanly separated 90-second chorus beats a degraded full album every time.
Run the job and come back later
Separation runs as a queued job rather than while you wait. That is a deliberate design decision: the work takes tens of seconds to minutes, and holding a browser request open that long is fragile — a refresh would lose the work and a retry would silently duplicate it.
Because it is queued, you can start a separation, close the tab, and open the workspace again later. The stems attach to the workspace when the job finishes.
When the result sounds wrong
Separation artefacts have recognisable causes, and most are fixable by changing the input rather than the settings.
The vocal sounds watery or underwater
This is the classic artefact of a heavily limited master — the loudness war problem. When everything is compressed to the same level, the model has less dynamic contrast to work with.
Try the least-compressed version you can find. A pre-master, a vinyl rip, or a streaming version at higher bitrate will often separate noticeably better than a loud radio master.
Another instrument is ghosting in the wrong stem
Usually a frequency overlap: a low male vocal sitting in the same range as a synth pad, or a floor tom overlapping a bass note. Switching from Full band to Core stems often helps, because the model stops trying to make a distinction it cannot confidently make.
Live recordings never separate cleanly
Microphone bleed is the reason. On a live recording the vocal mic has already captured the drums and the drum overheads have captured the guitar, so the "separate" sources were never separate to begin with. There is no setting that fixes this; expect a rougher result and treat it as a practice aid rather than a production asset.
Common questions
Can I remove vocals from a song to make a karaoke track?
Yes. Run Core stems, then mute the vocal track in the mixer. What remains is drums, bass and everything else — an instrumental backing track. Quality depends heavily on the source: a modern, well-produced record gives a very usable result, while a dense live recording will leave audible traces of the vocal.
What file formats does RigVerse Studio accept?
MP3, WAV, M4A, FLAC and OGG, up to 15 MB per file. Use the highest-quality source you have — a lossless WAV or FLAC separates more cleanly than a low-bitrate MP3, because heavy compression removes the detail the separation model relies on.
How long does stem separation take?
Typically tens of seconds to a few minutes depending on track length and queue depth. Because it runs as a background job rather than in your browser session, you can close the page and return — the stems attach to your workspace when the job completes.
Should I use four stems or six?
Start with four (Core stems: vocals, drums, bass, other). Six-stem Full band mode adds guitar and piano, but every extra split gives the model another chance to spread one instrument across two tracks. Only use Full band when you specifically need the guitar or piano isolated.
Why does my separated audio sound artificial?
Separation reconstructs each part rather than extracting it, so some artefacts are inherent. The most common causes are a heavily compressed master, a low-bitrate source file, or a live recording with microphone bleed. Using a less-compressed, higher-quality source usually improves the result more than any settings change.