Google Doc

9.703 model is UVR-MDX-NET 1, UVR-MDX-NET 2 is UVR_MDXNET_2_9682, NET 3 is 9662, all trained at 14.7kHz

Page 7 of 28 · Edit this page in Google Docs ↗

(instrumental based on processed phase inversion)

List of all (newer) available MDX models at the very top.

I think main was 438 in UVR 5 GUI at some point. At least now it's simply main_438 (if it wasn't from the beginning, but it was easy to confuse it with simply main model or even inst main)

(use MDX is a way to go now over VR) Generally use MDX when the results achieved with VR architecture are not satisfactory - e.g. too much vocal bleeding (e.g. in deep and low voices) or damaged instruments. If you only want acappella - it’s currently the best solution. Actually the best in most cases now.

MDX-UVR models are also great for cleaning artifacts from inverts (e.g. mixture (regular track) minus official instrumental or acappella).

(outdated) 9.682 might be better for instrumentals and inversion in some cases, while 9.7 for vocals, but better check already also newer models like 464 from KoD update (should be better in most cases) and also check Kim Model in GUI.

Generally on MVSEP's multisong dataset, these models received different SDR than on MDX21 dataset back in the days.

On MVSEP there’s 9.7 (NET 1) model, and it doesn't have any cutoff above training frequency for inverted instrumentals like currently GUI has. For (new) model it’s vocal 423 model and possibly with Demucs 2 enabled like in Colab, but it doesn’t have a specific jaggy spectrum above MDX training frequency which is specific to inverted vocal 4XX models from that period including Kim’s model.

Non-onnx version of voc_ft model in pth by MusicMan - 20x faster on MPS devices:

https://discord.com/channels/708579735583588363/887455924845944873/1204148534790852608 (roughly the same model size)

It won’t work in UVR. Inference code mirror: https://drive.google.com/file/d/1aSe0bwgIWhR7vvF1aoHQlCHpj39Kd-YK/view?usp=sharing

Mirror:

https://drive.google.com/drive/folders/16QbwuCBT0_w9nmNDg22m1niq0odtaZUP?usp=sharing

And the rest of MDX-Net v2 models: HQ_1-5, inst3, Kim inst, Kim Vocal 1-2, and older narrowband vocal and instrumental ones and Karaoke models.

(the old) Google Colab by HV

(with OG demucs 2 ensembling for vocal models)
https://colab.research.google.com/drive/189nHyAUfHIfTAXbm15Aj1Onlog2qcCp0?usp=sharing

Add separate cell as following, or else it won’t work

!pip install torch=1.13.1 (probably numpy 1.25 for this old Torch)

If you're still getting errors, delete whole MDX_Colab folder, terminate your session, make clean installation afterward, and don't forget to have this torch line executed after mounting (that might happen in case you manually replaced model.py with some of the ones below, and didn't restore the correct old one).

(The Colab to use MDX easily in Google’s cloud. Newer models not included, and it gives error if you add other models manually - custom models.py necessary, only 9.7 [NET 1-3] and karaoke models included above)

(In case of “RuntimeError: Error opening 'separated/(trackname)/vocals.wav': System error.” simply retry)

More MDX models explained in UVR section in the beginning of the document since they're a part of UVR GUI now.

Optionally, 423 model can be downloaded separately here (just in case, it’s main). It is on MVSEP as well.

(defunt) Upd. by KoD & DtN & Crusty Crab & jarredou, HV (12.06.23)
(probably now requires !pip install numpy==1.26 and restarting env)
It might have more models than above (e.g. some beta HQ ones)

____________________________________________________________________

The newest MDX Colabs - now with automatic models downloading (no more manual GDrive models installation as in older updates). Consider everything in the divided section later below as unnecessary.

https://colab.research.google.com/github/kae0-0/Colab-for-MDX_B/blob/main/MDX_Colab.ipynb (stable, lacks voc_ft batch process + also manual parameters loading per model like in the two above)

https://colab.research.google.com/github/jarredou/Colab-for-MDX_B/blob/main/MDX_Colab.ipynb (Beta. Might lack HQ_3 and voc_ft. It supports batch processing. Works with a folder as input and will process all files in it.

In "tracks_path" must be a folder containing (only) audio files (not the direct link to a file).

But the below might still work.)

https://colab.research.google.com/drive/1CO3KRvcFc1EuRh7YJea6DtMM6Tj8NHoB?usp=sharing (older revision with also auto models downloader, but with manual n_fft dim_f dim_t parameters setting like HV added)

and working one by HV linked at the top:

https://colab.research.google.com/github/NaJeongMo/Colab-for-MDX_B/blob/main/MDX-Net_Colab.ipynb

(new one by HV with community edits - 2025)

____________________________________________________________________

Old update from before model downloader implementation (May which year?)

MDX Colab with separate input for 3 models parameters, so you don’t need to change models.py every time you switch to some other model. Settings for all models listed in Colab. From now on, it uses reworked main.py and models.py downloaded automatically (made by jarredou). Don’t replace models.py from below packages with models from now on. Now denoiser also optionally added.

___________________________

(older Colab instruction)

To use more recent MDX-UVR models in Google Colab:

  1. Use and install this Colab (new) to GDrive at least once, run all the cells, nothing more - if you used MDX HV Colab (the one in the section above) on your specific Google Drive account before, ignore this step.
  2. Copy these files to onnx folder in MDX_Colab on your GDrive (inst1-3, 427) (down) https://drive.google.com/drive/folders/13SsV7b_kC6SqkICeX5wKhx-Z05uC8dLl (down)
  3. Overwrite models.py in MDX_Colab folder by provided below (not for new Colab)

(compatible with inst1-3, 427, Kim vocal and other)

https://cdn.discordapp.com/attachments/945913897033023559/1036947933536473159/models.py (completely different one with self.n_fft set to 7680 - incompatible with NET-1/9.x and 496 models)

  1. Use this notebook with added models

(the same as the link in point 1):

https://colab.research.google.com/drive/1zx7DQM-W9i7MJuEu6VTYz1xRG6lKRKVL?usp=sharing

  1. For Kims vocal model (poor instrumentals on Colab and no cutoff after inversion) copy vocals.onnx

(use the same models.py from point 3): https://drive.google.com/drive/folders/1exdP1CkpYHUuKsaz-gApS-0O1EtB0S82?usp=sharing

to onnx subfolder named "MDX-UVR-Kim Vocal Model (old)"

  1. For 496 inst model (inst main/MDX 2.1) go to the link below and put the model to onnx subfolder named “MDX-UVR Ins Model 496 - inst main-MDX 2.1” but you must replace attached models.py in the link in your GDrive (it’s from the OG HV Colab), and it is incompatible with the rest of the models in this new Colab - make a copy/rename the previous models.py in order to go back to it

(496 model is not as effective as 464/inst3 leaving more vocal residues in some cases, but might work well in specific scenarios). 496 is the only model requiring the old models.py from 9.7/NET1-3 models (attached below). https://drive.google.com/drive/folders/1iI_Zvc506xUv_58_GPHfVKpxmCIDfGhx?usp=share_link (if you place model in the wrong place, you’ll get missing vocals.onnx error [e.g. wrong folder structure or name] or “Got invalid dimensions for input: input for the following indices index: 2 Got: 3072 Expected: 2048.” [when having wrong models.py])

  1. Demucs turned on works only with default mixing algorithm and vocal models (or else you’ll get “ValueError: operands could not be broadcast together with shapes (8886272,2) (8886528,2)”). Also, chunks might have to be decreased.
  2. Be aware that after following these steps if you launch the old HV Colab above, it may overwrite models.py by the old one in point 6, which is compatible only with inst main/496 or full band models, so you'll need to repeat step 3 or 10 in case of invalid dimensions error or cutoff of full band model.
  3. In case of runtime error, to use Kim model decrease chunks from 55 to 50, and for Demucs on, decrease it to 40 (or respectively even lower)
  4. (beta) Full band beta 292 model (with new, only working for that model, models.py file with self.n_fft changed to 6144).

Go to the link below, copy model file to onnx subfolder called “MDX-UVR Ins Model Full Band 292” as in the link, and replace models.py (ideally make a backup/rename the old one in order to use previous models)

Thanks for help to Kim

https://drive.google.com/drive/folders/1CTJ6ctldr_avwudua1qJJMPAd7OrS2yO?usp=sharing

  1. (beta) Full band beta 403 model (with the same modified models.py for these two models)

Copy model file to:

Gdrive\MDX_Colab\onnx\MDX-UVR Ins Model Full Band 403\” as in the link below, and replace models.py in Gdrive\MDX_Colab

https://drive.google.com/drive/folders/1UXPxQMVAocpyDVb3agXu0Ho_vqFowHpA?usp=sharing

  1. (final) Full band 450/HQ_1 model (with the same modified models.py for the full band models)

Copy model file to:

Gdrive\MDX_Colab\onnx\MDX-UVR Ins Model Full Band 450 (HQ_1)\” as in the link below, and replace models.py in Gdrive\MDX_Colab (if you didn’t already for full band models)

https://drive.google.com/drive/folders/126ErYgKw7DwCl07WprAXWPD_uX6hUz-e?usp=sharing

  1. From now on, you’re forced to run separately newly added torch cell to fix PyTorch issues
  2. Newer full band 498/HQ_2 model (with the same modified models.py for the full band models)

Copy model file to:

Gdrive\MDX_Colab\onnx\MDX-UVR Ins Model Full Band 498 (HQ_2)\” as in the link below, and replace models.py in Gdrive\MDX_Colab (if you didn’t already for full band models)

https://drive.google.com/drive/folders/1O5b-uBbRTn_A9B2QkefklCT41YR9voMq?usp=sharing

  1. For full band models, use only modified models.py attached above, or you’ll get cutoff at 14.7kHz instead of 22kHz in spectrograms while using 427 models.py file.
  2. For Kim other FT instrumental model with cutoff but the highest SDR (even than inst3)

Copy both (vocals and other) model files to:

Gdrive\MDX_Colab\onnx\Kim ft other instrumental model\” as in the link below, and replace models.py in Gdrive\MDX_Colab (if you didn’t already for full band models)

https://drive.google.com/drive/folders/1v2Hy4AgFOJ9KysebGuOgn0rIveu510j6?usp=sharing (it will give only 1 stem output, models duplicated fixes errors in Colab, models.py is from inst3 model)

  1. If you use models.py from fullband model, it will output fullband for ft other model, but giving much more vocal residues (but it still might be even better in some busy mix parts than VR models, while having still less vocal residues only in those busy parts like chorus) - definitely use min_mag here.
  2. To fix the following error, make sure both vocals and invert vocals are always checked:

shell-init: error retrieving current directory: getcwd: cannot access parent directories: No such file or directory

Intel MKL FATAL ERROR: Cannot load /usr/local/lib/python3.9/dist-packages/torch/lib/libtorch_cpu.so.

Above error can also mean you need to terminate your session and start over. It randomly happens after using the Colab:

  1. I've reverted old "Karokee" and "Karokee_AGGR" models to use with the oldest HV’s models.py file, but these are old models (maybe they will do the trick, though).
  2. ModuleNotFoundError: No module named 'models'

Sometimes switching models.py doesn’t work correctly (especially during working on previously shared Colab folder with editing privileges) in that case, check Colab’s file manager if models.py is actually present after you’ve made a change on GDrive. If not, rename it to models.py (it might have been renamed to something else).

  1. Collection of all three models.py for all models for your comfort:

https://drive.google.com/drive/folders/1J35h9RYhPFk8dH-vShSW_AUharXY1YsN?usp=sharing

  1. Main_406 vocal model

https://mega.nz/file/dcREzKTR#PYKk3s1NPicC3mBBYH8ejC2rK_Im3sAj0p9xcOi1cpE

        "compensate": 1.075,

        "mdx_dim_f_set": 3072,

        "mdx_dim_t_set": 8,

        "mdx_n_fft_scale_set": 7680,

   

Models include here only: baseline, instrumental models: 415 (inst_1), 418 (inst_2), 464 (inst_3) trained on 17.7kHz, and vocal model 427, and Kim’s vocal model (old) (instrumental should be automatically made by inversion option, but it’s not a very good one for it) and 292 and 403 full band. If you want to use older 9.7 models, use old HV Colab above.

464/inst 3 should be the best in most cases for instrumentals and vocals than previous 9.x models, but depending, even in half of the cases, 418 can achieve better results, while full band 403 might give better results than inst3/464 in half of the cases.

Settings

max_mag is for vocals

min_mag for instrumentals

default

(deleted from the new HV Colab, still in Kae Colab above)

But "min mag solve some unwanted vocal soundings, but instrumental [is] more muffled and less detailed."

Also check out “default” setting (whatever is that, compare checksums if not one of these).

Chunks

As low as possible, or disabled.

Equivalent of min_mag in UVR is min_spec.

Be aware that UVR5, opposed to MDX Google Colab, applies cutoff to inverted output, matching the frequency of training frequency e.g. 17.7kHz for inst 1 and 3 models. It was to avoid some noise and vocal leftovers. You might have to apply it manually.

Also, you can uncomment visibility of compensation value in Colab, and change it to e.g. 1.08 to experiment.

Compensation value for 464 MDX-UVR inst. model is 1.0568175092136585

Default 1.03597672895 is for 9.7 model, and it also does the trick with at least Kim (old) model in GUI (where 1.08 had worse SDR).

Or check + 3.07 in DAW (it worked on Karokee model).

In Collab above, I also enabled visibility of max_mag for vocals and min_mag for instrumentals settings (mixing_algoritm).

Also, if you want to use Demucs option (ensemble) in Kae Colab, it uses stock Demucs 2, which in UVR5 was rewritten to use Demucs-UVR models with Demucs 3 or even currently better Demucs 4.

According to MVSEP SDR measurements, for ensemble Max Spec/Min Spec was better than Min Spec/Max Spec, but Avg/Avg was still better than these both.

Also for ensemble, Avg/Avg is better compared to e.g. Max Spec/Max Spec - it's 10.84 v 10.56 SDR in other result.

How denoiser work

It's not frequency based, it processes “the audio in 2 passes, one pass with inverted phase, then after processing the phase is restored on that pass, and both passes mixed together with gain * 0.5. So only the MDX noise is phase cancelling itself.”

Or the other way round:

“it's only processing the input 2 times, one time normal and one time phase inverted, then phased restored after separation, so when both passes are mixed back together only the noise in attenuated. There's no other processing involved”

Denoise serves to fix so called MDX noise existing in all inst/voc MDX-NET (v2) models.

______

Web version (32 bit float WAV as output for instrumentals, just use MDX-B for single MDX-Net models.

It was 9.682 MDX-UVR model in 2021, but in the end of 2022 it's probably inst 1 judging by SDR (not sure, as results are not exactly the same), then more models were added (e.g. HQ_3):

https://mvsep.com/

Web version (paid for MDX, lossless):

https://x-minus.pro/ 

In kae Colab, you can keep the option Demucs: off (ONNX only), it may provide better results in some cases even with the old MDX narrowband models (non-HQ).

In Colab you can change chunks to 10 if your track is below 5:00 minutes. It will take a bit more time, but the quality will be a bit cleaner, but more vocal residues can kick in (esp. short sudden ones).

Be aware that MDX Colabs for single models have 16 bit output.

And also noise cancellation implementation for MDX models in kae and HV Colab can differ a bit, plus there is also separate denoise method available as separate model.

Code for denoise method in HV Colab here.

As for any other settings, just use defaults since they're the best and updated.

Just for a vocal it’s one of the best free solutions on the market, very close to the result of paid and (partly) closed Audioshake service (#1 AI in a Sony separation contest; SDRs are from the contest evaluation based on private dataset). Very effective, high quality instrumental isolation AI and custom model (but the old models are trained at 14.7 kHz [NET-X a.k.a. 9.x] in comparison to VR models, and 17.7kHz in newer models like inst X and kim inst). 

In most cases MDX-UVR inverted models give less bleeding than VR (especially on bassy voice), while occasionally the result can be worse comparing to VR above, especially in terms of hi-end frequencies quality, but in general, MDX with UVR team models behaves the best for vocals and instrumentals.

Even instrumental from inverted vocals from vocal models gets less impaired than in VR, since vocal filtering is less aggressive, but with even more bleeding in some cases. Depends on a song.

You can support the creators of UVR and the newest MDX model is also available on https://www.patreon.com/uvr https://boosty.to/uvr to visit https://x-minus.pro/ to get an online version of MDX there as well (with exclusive paid models).

At least paid x-minus subscription allows you to use MDX HQ_2 498 (or HQ_3 already) instrumental model and for VR arch - 2_HP-UVR (HP-4BAND-V2_arch-124m), and Demucs 6s on their website. Feel free to listen and download lots of uploaded instrumentals on x-minus already. Dozens of instrumentals available.

Outdated

Alternatively you can experiment with 9662 model and ensemble it with the latest UVR 5's 4 band V2 with -a min_mag as Anjok suggested (but it was when new models weren't released yet).

Remotely I only know about old Colab which ensembles any two audio files, but it uses old algorithm if I'm not mistaken, so it is not as good (better use the ensemble Colab linked at the very top of the document):

https://colab.research.google.com/drive/1eK4h-13SmbjwYPecW2-PdMoEbJcpqzDt?usp=sharing

_____

Note

Don’t disable invert_vocals in Colab even if you only need vocal instead of instrumental, otherwise the Colab will end up with error.

MDX noise

There is a noise using all MDX-UVR inst/vocal models, and it’s model dependent (irc 4 stems don’t have it). It's fixed in Colabs using denoiser "however by using my method, conversions will be 2x slower as it needs to predict twice.

I see no quality degradation at all, and I can't believe it actually worked rofl" -HV

Also, UVR 5 GUI has the same noise filtering implemented (if not better, also with alternative model).

Current MDX Colab has normalization feature “normalizes all input at first and then changes the wave peak back to original. This makes the separation process better, also less noise. IDK if you guys have tried this, but if you split a quiet track, and normalize it after MDX inference the noise sounds more audible than normalizing it and changing the peak back to original.”

If you want to experiment with MDX sound, the Colab from before that change is below:

https://colab.research.google.com/drive/1EXlh--o34-rzAFNEKn8dAkqYqBvhVDsH?usp=sharing (might no longer work due to changes made by Google to Colab environment, the last maintained are kae and HV (new) Colabs)

Furthermore, you can also try manually mixing vocal with original track using phase inversion and add specific gain on vocal track (+1.03597672895 or +3.07) for 9.7 model (or other ones with different values), using both this and below Colab and save result as 32 bit float (but this might have more bleeding, but it uses 32 bit while chunking):

https://colab.research.google.com/drive/1R32s9M50tn_TRUGIkfnjNPYdbUvQOcfh?usp=sharing#scrollTo=lkTLtOvyBuxc

(for e.g. the best compensation value for 464 MDX-UVR inst. model is 1.0568175092136585

and it's not constant)

Also be aware that MVSEP uses 32 bit for MDX-UVR models for ready inversion of any model too.

If you look for eliminating the noise from MDX-UVR instrumentals, also the method described in Zero Shot below might work.

"I just run the MDX vocals thru UVR to remove any remaining buzz noises and synths, it works great so far" (probably meant one of VR models)

Average track in Colab is being processed in 1:00-1:30 minute using slower Tesla K80 (much faster than even UVR’s HP-4BAND-V2_arch-124m model).

If you want to get rid of some artifacts, you can further process output vocal track from MDX through Demucs 3.

Options in the old HV MDX Colab/or kae fork Colab (from the very top)

Demucs model in the older MDX-Net Colab

When it's enabled, it sounds better to me, used with the old narrowband 9.X and newer vocal models, as Demucs 2 model is fullband, but opinions on superiority of this option are divided, and MVSEP dev made some SDR calculation where it achieved worse results with Demucs enabled. But be aware, that inverted results from narrowband are still fullband despite the narrowband training frequency, as there’s no cutoff matching present in Colab, as it’s implemented in UVR GUI as a separate option. Using such cutoff matching training frequency (which can be observed in non-inverted stem) might lead to less noise and residues in the results. Demucs model will work correctly only with vocal models in Colabs (we didn’t have any MDX instrumental models back then, so naming scheme is reversed for these models, hence Demucs model with instrumental model produces distorted sound, it mixes vocals with instrumental in a weird way).

“The --shifts=SHIFTS performs multiple predictions with random shifts (a.k.a. the shift trick) of the input and average them. This makes prediction SHIFTS times slower but improves the accuracy of Demucs by 0.2 points of SDR. It has limited impact on Conv-Tasnet as the model is by nature almost time equivariant. The value of 10 was used on the original paper, although 5 yields mostly the same gain. It is deactivated by default, but it does make vocals a bit smoother.

The --overlap option controls the amount of overlap between prediction windows (for Demucs one window is 10 seconds). Default is 0.25 (i.e. 25%) which is probably fine.”

You can even try out 0.1, but for Demucs 4 it decreases SDR in ensemble if you’re trying to separate a track containing vocals. If it’s instrumental, then 0.1 is the best (e.g. for drums).

(outdated/for offline use/added to Colab)

Here's the new MDX-B Karokee model! https://mega.nz/file/iZgiURwL#jDKiAkGyG1Ru6sn21MkIwF90C-fGD0o-Ws58Mn3O7y8

The archive contains two versions: normal and aggressive. The second removes the lead vocals more. The model was trained using a dataset that I completely created from scratch. There are 610 songs in total. We ask that you please credit us if you decide to use these models in your projects (Anjok, aufr33).

__________________________________________________________________