>Drumsep - single percussion instruments separation
If you want to further separate single instruments from drums stem separated with e.g. MDX23 Colab, Mel-Roformer drums on x-minus.pro premium, MVSEP, or Demucs_ft (not necessarily BS-Roformer SW) into: hihat, cymbals, kick, snare and more, you might want to check below solutions. Sampling from such separated stems might be not the best idea due to the quality (see here for free Drumclone plugin allowing even different types of synthesised kicks from mixture; video). But e.g. it serves well for purposes of conducting new mixes/remasters of the same songs or separated instrumentals, e.g. when it's overlapped with better quality, previously separated drums stem. It might give interesting results when aligned with the original drums, and rebalanced with effects (drums stem might end up louder in the mix than separated percussion, as most likely it will still have better quality). Check out these drum replacers.
To potentially increase drumsep models separation quality, “try using a small pitch shift up or down, like +/- 1 or 2 semitones (...) can sometimes help bring out the lows or highs if they seem weak.” (CZ-84)
Also, consider using good instrumental model before using 4 stem model for drums (if it's not instrumental already), to enhance drums stem, and then to enhance drumsep result.
Some drumsep models might have a bug where “a small, but relevant portion of audio is being lost when the [drumsep] model is being used”
“The solution is to invert the phase on all the drum stems into the original file and save that as its own file, making your own "other" file”. It has been fixed on MVSEP.
Consider using Lew Apollo Universal on hi-hat and snare from drumsep 6 stems model by jarredou from drums model from e.g. flowersv10 inst model (I've noticed that most of the residual "Fuzz" is usually trapped in the snare or cymbals. - CC Karaoke)
Mel-Roformer MVSEP drumsep models
1) 4 stems v2 (kick, snare, toms, cymbals) - “It gives the best metrics with a big gap for kick, snare and cymbals.” - ZFTurbo. The old v1 below was removed.
(metrics; only toms are worse SDR-wise vs previous SCNet Drumsep models below)
1) 4 stems v1 removed (kick, snare, toms, cymbals) - average SDR of hihat ride, crash is 11,52 (but in one stem) and so far it’s the best SDR out of all models (even vs the previous ensemble consisting of three MDX23C and SCNet models).
2) 6 stems (kick, snare, toms, hihat, ride, crash) - average SDR of hihat ride, crash is 8.18 (but from separated stems), while
The snare in 1) has the best SDR out of all available models.
Kick and toms are still the best SDR-wise in the previous 3x MDX23C and SCNet ensemble (new ensemble with these new Mel-Roformers so far)
- The new models “are very great for ride/crash/hh. And overall, they have the best metrics for almost all stems.” - ZFTurbo
SDR/L1 Freq/bleedless/fullness chart of all models
Evaluations on new dataset (esp. check Log WMSE Results with “"bypass_filter" with torch_log_wms, ([good] at least for drums or anything rich in low frequency content)” - jarredou
Sometimes the newer jarredou’s drumsep 6 stems model below can serve to clean up “upper frequency range of snare hits” in the cymbals stems in “either MelRoFormer or SCNet-XL four-or-six stem DrumSep models” - Dyslexicon
“The core problem is that the main MVSep Drums model which is used by everything- including the Drumsep models- is not purely drums, it's mixed with other percussion which taints things.” - godzfire
SCNet MVSEP drumsep models
Better SDR than MDX23C and Demucs models above
- MVSEP 8 stems ensemble of all the 4 drumsep models below (along with MDX23C model, and besides Demucs model by Inagoy) metrics
- MVSEP’s SCNet 4 stem (kick, snare, toms, cymbals) out of following models, the best SDR for kick and similar to 6 stem below for toms - only -0.01 SDR difference)
- MVSEP’s SCNet 5 stem (cymbals, hi-hat, kick, snare, toms)
- MVSEP’s SCNet 6 stem model (ride, crash, hi-hat, kick, snare, toms) worse snare SDR
(newer) MDX23C 5 stem drumsep by jarredou
Model | yaml | MK Fixed Colab | UVR (not on MVSep) All SDR metrics are better than the previous 6 stem model below:
SDR: kick: 16.66, snare: 11.54, toms 12.34, hihat: 4.04, cymbals: 6.36 (all metrics).
Metric fullness for snare: 25.0361, bleedless for hh: 12.3470, log_wmse for snare: 13.8959
“it's more on the fullness side than bleedless” - from all the metrics, only bleedless for snare is worse than in the previous model:
26.8420 vs 30.4149
“Quite cleaner than the previous [6 stem] one”, “a lot noisier than other drumpsep models, but that's not necessarily a bad thing.”
Possible “UnpicklingError: "invalid load key, '\x0a'."” issue in UVR if you use the old 6 stem yaml.
Maybe if we separate just snare with the old MDX23C model below from an already separated drums stem, and mix/invert to get the rest, then pass it through the new model, the bleed would be gone.
For comparison, metrics of the old 6 stem jarredou/Aufr33 MDX23C model
(which has cymbals divided into ride and crash which are not in the evaluation dataset):
SDR: kick: 14.55, snare: 9.79, toms: 10.64, hihat: 3.20, cymbals: 6.08
Metric fullness for snare: 25.0361, bleedless for hh: 10.2765, log_wmse for snare: 12.4258
The model was trained with a lightweight config to train on a subpar T4 GPU on free Colabs and 10 accounts. The metrics do not surpass exclusive drumsep Mel-Roformer and SCNet models on MVSEP, but at least you can use this one locally.
“Depending on the quality tier of input source material, it can sometimes yield more accurate stem-to-stem separations than either MelRoFormer or SCNet-XL four-or-six stem DrumSep models. (...)
For example, I often find that MelRoFormer DrumSep can leave the upper frequency range of Snare hits and mis-assign them to the Cymbals stem. This is a common issue I have encountered with separating AUD recordings with MelRoFormer DrumSep.
MDX23c 5-stem drumsep is trained in such a way that it separates these snare remainders out of the Cymbals stem, which is extremely useful. " Dyslexicon”
(older) MDX23C 6 stem drumsep by jarredou/Aufr33
Use it on already separated drums.
Download
Model | yaml | MK Fixed Colab | UVR
Added on MVSEP and uvronline too.
(jarredou) “Drums Separation model trained by aufr33
(on my not-that-clean drums dataset)
Stems:
kick, snare, toms, hh, ride, crash
MVSEP dataset evaluation:
SDR: kick: 14.55, snare: 9.79, toms: 10.64, hihat: 3.20, cymbals: 6.08, hihat & cymbals: 6.77
To get potentially better results with the model “try using a small pitch shift up or down, like +/- 1 or 2 semitones, in the settings you use to extract the drum stem from the instrumental stem. (...) can sometimes help bring out the lows or highs if they seem weak.” (CZ-84)
Works better at debleeding than MVSep models above.
It can already be used, but training is not fully finished yet.
The config allows training on not so big GPUs [n_fft 2048 instead of 8096], it's open to anyone to resume/fine-tune it.
For now, it's struggling a bit to differentiate ride/hh/crash correctly, kick/snare/toms are more clean.
[“and has the usual issues with mdx23 models, but it’s an improvement over drumsep I think” - Dry Paint Dealer Undr]
- If you got an error while using jarredou’s Drumsep Colab (object is not subscriptable):
change to this on line 144 in inference.py:
if type(args.device_ids) != int:
model = nn.DataParallel(model, device_ids = args.device_ids)
(thx DJ NUO)
It works in UVR too. All models should be located in the following folder:
Ultimate Vocal Remover\models\MDX_Net_Models
Don't forget about copying the config file to: model_data\mdx_c_configs.
Once the model is detected, select the config in a new window, and that’s all.
The model achieved much better SDR on private jarredou's small evaluation dataset compared to the previous drumsep model by Inagoy which was based on a worse dataset and older Demucs 3 arch.
The dataset for further training is available in the drums section of Repository of stems/multitracks - you can potentially clean it further and/or expand the dataset so the results might be better after resuming the training from checkpoint. Using the current dataset, the SDR might stall for quite some amount of epochs or even decrease, but it usually increases later, so potentially training it further to 300-500-1000 epochs might be beneficial.
Attached config also includes necessary training parameters for training further using ZFTurbo repo.
Current model metrics (not MVSEP evaluation dataset):
“Instr SDR kick: 18.4312
Instr SDR snare: 13.6083
Instr SDR toms: 13.2693
Instr SDR hh: 6.6887
Instr SDR ride: 5.3227
Instr SDR crash: 7.5152
SDR Avg: 10.8059” Aufr33
And if evaluation dataset hasn't changed since then, the old Drumsep SDR:
“kick : 13.9216
snare : 8.2344
toms : 5.4471
(I can't compare cymbals score as it's different stem types)” - jarredou
After initial jarredou’s training in Colab, Aufr33 decided to train the model for additional 7 days, to at least above epoch 113 (perhaps around 150, it wasn't said precisely), while using the same config, but on a faster GPU (rented 2x RTX 4090).
Even epoch 5 trained on jarredou's dataset casually in slow and troublesome free Colab (which uses Tesla T4 15GB with performance of RTX 3050, but with more VRAM) with multiple Colab accounts and very light and fast training settings, already achieved better SDR than Drumsep using smaller dataset and older architecture. Colab epochs metrics:
“epoch 5:
Instr SDR kick: 13.9763
Instr SDR snare: 8.4376
Instr SDR toms: 6.7399
Instr SDR hh: 0.7277
Instr SDR ride: 0.8014
Instr SDR crash: 4.4053
SDR Avg: 5.8480
epoch 15:
Instr SDR kick: 15.3523
Instr SDR snare: 10.8604
Instr SDR toms: 10.3834
Instr SDR hh: 4.0184
Instr SDR ride: 2.7248
Instr SDR crash: 6.1663
SDR Avg: 8.2509”
Don't forget to use already well separated drums (e.g. from Mel-Roformer for premium users on x-minus or MVSEP Drums ensemble) from well separated instrumental as input for that model, or jarredou’s MDX23 Colab fork v. 2.5 or also for all stems - MVSEP 4/+ ensemble (premium).
Purely for drums separation from even instrumentals, the model might not give good results, hence it needs separated drums first. It was trained just on percussion sounds and not vocals or anything else.
Also, e.g. the kick and toms might have a bit weird looking spectrograms. It’s due to:
“mdx23c subbands splitting + unfinished training, these artifacts are [normally] reduced/removed along [further] training.” Examples
BS-Roformer MVSEP 4 stems
From “MVSep Mega 53 Stems” all-in-one or single models (less VRAM-hungry) - 1.27GB | splifft | MVSepless HF / HF CPU / Colab / newer (you can pick which stems you want there, but it’s in Russian, and HF doesn’t support auto-translate, but Colab might)
The model can be muddy, it's small. It was further trained on MVSep in the all-in-one model and returns only present stems.
('hh', 'kick', 'toms', 'snare'). It's further trained on MVSep. It will detect only these four there if you provide just drums stem as input, omitting all the rest, but it will be probably much muddier than all other models (maybe despite the imagoy’s), but also fast.
Older drumsep by Inagoy
Demucs 3 model. Just remember to use drums in one stem (e.g. with demucs_ft) from already good sounding instrumental or ensemble on MVSEP or MDX23 v. 2.4 Colab first, as use it as input (both are better for instrumental in most cases than just Demucs 4 - you can use various settings for ensembles to get better instrumentals, the better drums, the better results from drumsep)
- or Kubinka Colab (you can provide direct links there)
- Available on MVSEP.com (but you can use more intensive parameters in Colab for a bit better quality)
(Use these solutions instead of GitHub Colab as the model's GDrive link from OG GitHub Colab is currently deleted, so drumsep won’t work correctly, unless you replace GDrive link with model to the .th model reupload:
https://drive.google.com/file/d/1S79T3XlPFosbhXgVO8h3GeBJSu43Sk-O/view)
- Windows installation - execute the following:
demucs --repo "PATH_TO_DrumSep_MODEL_FOLDER" -n modelo_final "INPUT_FILE_PATH" -o "OUTPUT_FOLDER_PATH"
- You can also use drumsep in UVR 5 GUI
(so beside using fixed Colab or in CML):
Go to UVR settings and open application directory.
Find the folder "models" and go to "demucs models" then "v3_v4"
Copy and paste both the .th and .yaml files, and it's good to go.
Be aware that stems will be labelled wrong in the GUI using drumsep.
It's much more sensitive to shifts than overlap, where above 0.6-0.7 it can become placebo. Consider testing it with shifts 20.
But some people find using shifts 10 and overlap 0.99 better than shifts 20 and overlap 0.75.
Just be aware, that if you’re willing to wait, you can further increase shifts to 20 if you want the best of both worlds.
Also, consider testing it with -6 semitones e.g. in UVR 5.6/+, or with 31183Hz sample rate with changed (decreased) tempo and pitch.
-12 semitones from 44100Hz is 22050 and should be rather less usable in most cases, the same for tempo preservation, it should be off.
Be aware that sometimes it can “consistently put hi hats in snare stem” and can contain some artefacts, and results might not null with the source.
“From what I've tested (on drums already extracted with demucs4_ft from a live band recording from the output of the soundboard... so shitty sounding!), It is quite good at separating cymbals from shells, and kick from snare, but there are parts of kick or snare sounds that can go into the toms stem (...it's easy to fix manually in a DAW)”
"Ok I did test it.
- You're right, Drumsep is good if shifts are applied, this makes a HUGE difference, first time i did test it with 0 or 1 shift and results were meh. Shifts (from about 5/6/10 depending on source) clean it nicely.
Minuses: only 4 outputs. Not enough for a lot of drumtracks (but hey you can Regroove results, and this is what i will be doing probably from now) - It takes a long time with a lot of shifts, - it doesnt null with original tracks
- Regroove allows me more separations, especially when used multiple times, so as a producer it allows me to remove parts of kicks, parts of snares etc, noises etc. More deep control. Plus it nulls easily (it always adds the same space in front) so I can work more transparently.
But you're right, I will use drumsep in the Colab with a lot of shifts as a starting point in most cases now."
"It's trained with 7 hours of drum tracks that I made using sample-based drum software like Adictive Drums, trying to get as many different-sounding drums as I could. As everything was controlled with MIDI, I could export the isolated bodies: kick, snare, toms (all on one track), and cymbals (including hi-hat). So every dataset example is composed of kick, snare, toms, cymbals, and the mixture (the sum of all of them)." - said the author - Inagoy
From paid solutions for separating drums' sections there is mainly a paid FactorSynth and other alternatives are more problematic or less perfect.
Use free zero shot for separating single other instruments from e.g. others stem from Demucs or GSEP.
Moises.ai drumsep
(only for Pro)
- Kick, snare, toms, hi-hat, cymbals, other
It’s not well documented on their promotional materials, but the option is available after dragging your input file on the site, and then under drums button.
FactorSynth
Since version 3 available in a form of plugin for most DAWs. Demo runs for 20 minutes at a time. Exporting and component editing are disabled.
Till v. 2 it was Ableton-only compatible add-on. And (probably) could be used on free Ableton Live.
Also, not for separating drums from a full mix, but for separating your already separated drums into further layers like kick, snare, transients, cymbals, etc. from Demucs or GSEP (the latter usually has better shakers and at least hi-hats when they're in fast tempo).
[till v2 demo version limit was 8 seconds and no limit for full version]” “it’s amazing”.
It works the same way as Regroover VST (which may have some problems with creating a trial account).
It’s comparable or better quality (both better than zero shot for at least drums).
“Factorsynth has more granularity, but drumsep is easier to work with and gets less confused between toms and kicks.”
There’s a freeware prototype 0.4-0.1 versions from 2017 for Mac available to download:
https://www.jjburred.com/software/factorsynth/proto.html
Regroover
Regroover is only for 30 seconds chunks, and they require manual align due to phasing issues - additional silence is added in the beginning and ending.
“Get your 30-second drum clip, then drag and drop it into Regroover.
Make sure to de-select the Sync option, as it will time stretch it by default.
On the right-hand side, I recommend changing the split to 6 layers instead of 4, simply for flexibility.
Once it has processed that, you can choose export -> layers."
There was a report that probably newer versions might not be feasible for this task anymore.
In other words:
It’s much more hassle to use it than drumsep, but it’s very good “if you need particular sound and not about pattern etc.
1. separate drums from whole track (demucs)
2. Cut drum track into max 30 second cuts [regroover limits] and ideally cut right on transient, some space before kick helps,
2. You use regroover for the first time and for example try to separate to 4 tracks, just so overall separation.
3. Those separation sums exactly to that is given, sometimes it just need to be realigned few ms.
4. And if for example kick still has some not needed parts, you just regroove it once again.
If are looking overall fast and for patterns, drumsep. Regroover for painfull but precise job. Also in most cases hihats are trash, but snare's and kicks you often can find perfetclu usable ones. I'm not sure about metal but overall.”
UnMixingStation
"Very, very old and almost impossible to find, but the separations are 95% close to Regroover". The software is 13 years old, and their site is down, and the tool doesn’t seem to be available to buy anywhere.
LarsNet
Adden on MVSep. Colab. Source: https://github.com/polimi-ispl/larsnet
It separates previously separated drums into 5 stems: kick, snare, cymbals, toms, hihat.
It’s worse than Drumsep as it uses Spleeter-like architecture, but “at least they have an extra output, so they separate hihats and cymbals.”. Colab
“Baseline models don't seem better quality than drumsep, but the provided checkpoints are trained with oly 22 epochs, it doesn't seem much. (and STEMGMD dataset was limited by the only 10 drumkits), so it could probably be better with better dataset & training”
Similar situation as with Drumsep - you should provide drums separated from e.g. Demucs model.
There’s also Zynaptiq Unmix Drums, but it’s not exactly a separation tool, but to “Boost Or Attenuate Drums In Mixed Music”.
- For only kick and hi hat separation now free -
VirtualDJ 2023/Stems 2.0 (kick, hi-hat)
Probably using drums from Demucs 4 or GSEP first, will give better results but, it's not perfect. In many cases it may leave bleeding of snare a little bit, in both hi-hat and kick track. Sadly it sometimes confuses these elements of a mix.
"If you are not using it professionally, and do not use any professional equipment like a DJ controller, or a DJ mixer, then VirtualDJ is (now) FREE".
RipX DeepAudio (-||-) (6 stems [piano, guitar])
Popular tool. Decent results for specific drums' sections separation (but as for vocal/instrumental/4 stems separation, all the tools mentioned in at the very top of the document outperforms RipX, so use it only for specific drums’ section separation only, at best using Demucs 4 or GSEP for drums stem).
"It can separate a file into a buncha things into a lot more types of instruments than just the basic 4 stems (with varying degrees of success ofc).
Might be a case that old cracked versions of RipX don't allow separating drums sections well, or just the opposite - check both the newest version and Hit'n'Mix RipX DeepAudio v5.2.6, but probably the latter doesn't support separating single drums yet.
It’s basically UVR but with their custom models + SFX single stem
It's good for guitar, but not in all cases (possibly Demucs for 4 stems).
Piano and guitar models were added recently (somewhere in the January 2023)
- Hit 'n' Mix RipX DAW Pro 7 released. For GPU acceleration, min. requirement is 8GB VRAM and 10XX card or newer (mentioned by the official document are: 1070, 1080, 2070, 2080, 3070, 3080, 3090, 40XX). Additionally, for GPU acceleration to work, exactly Nvidia CUDA Toolkit v.11.0 is necessary. Occasionally during transition from some older versions, separation quality of harmonies can increase. Separation time with GPU acceleration can decrease from even 40 minutes on CPU to 2 minutes on decent GPU.
They say it uses Demucs.
We have reports about crashes, at least on certain audio files. There are various RipX versions uploaded on archive.org, maybe one will work, but some keys work only on versions from 2 and up.
Spectralayers 10
Received an update of an AI, and they no longer use Spleeter, but Demucs 4 (6s), and they now also good kick, snare, cymbals separation too. Good opinions so far. Compared to drumsep sometimes it's better, sometimes it's not. Versus MDX23 Colab V2, instrumentals sometimes sound much worse, so rather don’t bother for instrumentals.