UVR’s VR architecture models
(settings and recommendations;
mostly outdated arch for all vocals and instrumental models)
Available on Colab, HF, UVR, UVR old CLI, MVSEP
VR Colab by HV
Use this fixed notebook for now (04.04.23)
Sometimes Google Colab might break itself (e.g. error: No module named 'pathvalidate'), and then you can simply try to go to Environment and delete it entirely and start over, and then it might start working.
(since 17.03.23 the official link for HV Colab at the very top stopped working (librosa, and later pysound related issues with again YT links, but somehow fixed) “!pip install librosa==0.9.1” in OG Colab fixes the issue and is only necessary for both YT and local files and clean installation works too.)
- HV also made a new VR Colab which irc, now don’t clutter all your GDrive, but only downloads models which you use (but without VR ensemble) and probably might work without GDrive mounting.
(Google Colab in general allows separating on free virtual machine with decent Nvidia GPUs - it's for all those who don't want to use their personal computer for such GPU/CPU-intensive tasks, or don’t have Nvidia GPU or decent CPU, or you don’t want to use online services - e.g. frequently wait in queues, etc.)
Video tutorial how to use the VR Colab (it’s very easy to use): https://www.youtube.com/channel/UC0NiSV1jLMH-9E09wiDVFYw
You can use VR models in UVR5 GUI or
To use the above tool locally (old command line branch for VR models only):
https://github.com/Anjok07/ultimatevocalremovergui/tree/v5-beta-cml
Installation tutorial: https://www.youtube.com/watch?v=ps7GRvI1X80
In case of CUDA out memory error due to too long files, use Lossless-cut to divide your song into two parts,
or use this Colab which includes chunks option turned on by default (no ensemble feature here):
_________
Below, I'll explain Ultimate Vocal Remover 5 (VR architecture) models only (fork of vocal remover by tsurumeso).
For more information on VR arch, see here for official documentation and settings:
https://github.com/Anjok07/ultimatevocalremovergui/tree/v5-beta-cml
https://github.com/Anjok07/ultimatevocalremovergui
The best
VR settings
Explained in detail
Settings available in Colab and in CLI branch, and also UVR 5 GUI (but without at least mirroring2. mirroring in UVR5 GUI for VR arch got replaced entirely by High End Process (works as mirroring now, and not like original High End Process which was originally dedicated for very old 16kHz VR models only).
These VR models can be used in this 1) Colab or in 2) UVR5 GUI or on 3) mvsep.com (uses 512 windows size, aggressiveness option, various models) 4) x-minus.pro/uvronline.app (for free one UVR (unreleased) model without parameters ("lo-fi" option, mp3, 17,7 kHz cutoff) [Demucs 4 for registered users iirc (site by Aufr33 - one of the authors of UVR5)]
I had at least one report that results for just VR models are better using Colab above/old CLI branch instead of the newest UVR5 GUI, so be aware (besides both mirroring settings - only mirroring is working under high-end process - no mirroring2 [272 window size is added back as user input] all settings should be available in GUI). Interestingly, I received similar report for MDX models in UVR5 GUI comparing to Colab (be aware just in case). The problems might be also bound to VRAM, and don't exist on 11GB GPPUs and up or in CPU mode.
Before we start -
Issue with additional vocal residues when postprocess is enabled
- “postprocess option masks instrumental part based on the vocals volume to improve the separation quality." (from: https://github.com/tsurumeso/vocal-remover)
where in HV Colab it says: “Mute low volume vocals”. So, if it enhances separation quality, then maybe it should cancel some vocals residues ("low volume vocals") so that's maybe not too bad explanation.
But that setting enabled in at least Colab may leave some vocal residues:
(it’s fixed in UVR GUI "the very end bits of vocals don't bleed anymore no matter which threshold value is used")
Customizable postprocess settings (threshold, min range and fade size) in HV's Colab were deleted, and were last time available in this revision:
So change default 0.3 or 0.2 threshold value (depending on revision) and set it to 0.01 if you have the issue when using postprocess.
The threshold parameter set to 0.01 fixes the issue (so quiet the opposite thing happened using default settings than this option should serve to, I believe).
Also, default threshold values for postprocess changed from 0.3 to 0.2 in later revisions of the Colab.
- Window size option set to anything other than 512 somehow decrease SDR, although most people like lower values (at least 320, me even 272; 352 is also possible, but anything above changes the tone of sound more noticeably) - we don’t know yet why lower window sizes mess with SDR (similar situation like with GSEP) - 512 might be a good setting for ensemble with other models than VR ones or for further mastering. Sometimes compared to 512 windows size, 272 can lead to a bit more noticeable vocal residues. You might find bigger window sizes less noisy in general, but also more blurry for some people.
- Aggressiveness/Aggression - “A value of 10 is equivalent to 0.1 (10/100=0.1) in Colab”.
Strangely, the best SDR for aggressiveness using MSB2 instrumental model turned out to be 100 in GUI, 10 in Colab, while we usually used 0.3 for this model and 500m_x as well, while HP models usually behaves the best with lower values than HP2 models (0.09/10 in GUI).
- Mirroring turned out to enhance SDR. It adds to the spectrum e.g. above 20kHz for a base training frequency of VR model (all 4 bands).
none - No processing (default)
bypass - This copies the missing frequencies from the input.
mirroring - This algorithm is more advanced than correlation. It uses the high frequencies from the input and mirrored instrumental's frequencies. More aggressive.
mirroring2 - This version of mirroring is optimized for better performance.
--high_end_process - In the old CLI VR, this argument restored the high frequencies of the output audio. It was intended for models with a narrow bandwidth - 16 kHz and below (the oldest “lowend” and “32000” ones, none more). But now in UVR5 GUI, High-end process is counterpart of mirroring.
(current 500MB models don’t have full 22kHz coverage, but 20kHz, so consider using mirroring instead or none if you want fuller spectrum)
- Be aware, that even for VR arch, the same rule for GPUs with less than 8GB VRAM applies (inb4 - Colab T4 has 15GB) - separations on 6GB VRAM have worse quality with the same parameters. In order to work around the issue, you can split your audio into specific parts (e.g. for all chorus, verses etc).
VR models settings and list
For VR architecture models, you can start with these two fast models:
Model: HP_4BAND_3090_arch-124m (1_HP-UVR)
1) Fast and reliable. V2 below has more “polished” drums, while here they’re more aggressive and louder. Sometimes V2 might be safer and can fit in more cases where it’s not hip-hop and music is not drum oriented, but that one rarely harms some instruments more in certain cases with more busy mix with e.g. repeatable synth. You may want to isolate using these two models and pick the best results on even the same album.
Windows size: 272
Aggressiveness: 0.09 in Colab/CLI/MVSEP (9 in UVR 5.6.x)
TTA: ON (OFF if snare is too harsh)
Post-processing: OFF (at least for this model - it can get muffle instruments in background beside drums of the track in some cases, e.g. guitar)
"Mirroring" (Hi-end process in GUI) (rarely "Mirroring2" here, since the model itself is less smooth and usually have better drums, but it sometimes leads to overkill - in that case check mirroring2 in CLI or V2 model above)
Better yet, to increase the quality of the separation (when drums in e.g. hip-hop can be frequently damaged too much during the process) go now straight to the Demucs section and read the "Anjok's tip".
If you have too many vocal residues vs 500m_1 model, increase aggressiveness from 0.09 to 0.2 or even 0.3, but it’s destructive for some instruments (at least without Demucs trick above).
Model: HP-4BAND-V2_arch-124m (2_HP-UVR)
!) Fast and nice model, but sometimes gives lots of vocal residues comparing to above, but thanks to this, it may sometimes harm snare less in some cases (still 4 times faster than 500m_1) it’s ~55/45 which model is better and depends on the album even on the same genre:
Window size: 272 (the lowest possible; in some very rare cases it can spoil the result on 4 band models, then check 320)
Aggressiveness: 0.09 (9 in GUI)
TTA: ON (instr. separation of a better quality)
Postprocess: (sometimes on, it rather compliments to the sound of this model especially when the result sounds a bit too harsh, but it also can spoil drums in some places when e.g. strong synths suddenly appear in mix for short, probably misidentifying them as vocals, so be aware)
Mirroring (it fits pretty well to this model in comparison to mirroring2 which is not “aggressive” enough here) [mirroring doesn’t seem to be present in GUI so be aware)
Processing time for this model is 10 minutes using the weakest GPU in Colab (but currently you should be getting better Tesla T4).
(for users of x-minus) “slightly different models [than in GUI] are used for minimum aggressiveness. When we train models, we get many epochs. Some of these models differ in that they better preserve instruments such as the saxophone. These versions of the models don't get into the release, but are used exclusively on the XM website.”
Model: HP2-4BAND-3090_4band_arch-500m_1 (9_HP2-UVR)
3) Older good model, but resource heavy - check it if you get too many vocal residues, or in other cases - when your drums are too muffled - rarely there might be more bleeding and generally more spoiled other instruments in comparison to those above, it depends on a track. In some cases it bleeds vocal less than HP_4BAND_3090_arch-124m
Window size: 272
Aggressiveness: 0.3-0.32 (30-32 in GUI)
TTA: ON
Postprocess: (turned ON in most cases with exceptions (it’s polishing high-end), and the problem with muffling instruments using ppr doesn’t seem to exist in this model)
Mirroring2 (I find mirroring[1] too aggressive for this model, but with exceptions)
! Be aware these settings are very slow (40 minutes per track in Colab on the former default K80 GPU, but it's faster now) so just in case, you might want to experiment with 320/384, or at worse even 512 window size if you want to increase processing speed in cost of isolation precision.
Colab’s former default Tesla K80 processes slower than even GTX 1050 Ti, so if you have a decent Nvidia GPU, consider using UVR locally. Since May 2022 there is faster Tesla T4 available as default, so there shouldn't be any problem.
HP2-4BAND-3090_4band_arch-500m_2 (8_HP2-UVR)
was worse in I think every case I tested, but it’s good for a pair for ensemble (more about ensemble in section below).
Model: HP2-MAIN-MSB2-3BAND-3090_arch-500m (7_HP2-UVR.pth)
4) Last resort, e.g. when you have a lot of artifacts (heavily filtered vocal residues) some instruments spoiled, and no equal sound across the track. Last resort, because it’s 3 band, instead of 4 band, and it lacks some hi-end/clarity, but if your track is very demanding to filter out vocal residues, then it’s good choice. The best SDR among VR-arch models.
Window size: 272
Aggressiveness: 0.3
TTA: ON
Postprocess: ON
Mirroring
It’s similarly nightmarishly slow in Colab just like 500m_1/2 using these settings (1 hour for a track on K80) when you got accidentally slower Tesla K80 assigned in Colab instead of Tesla T4.
HighPrecison_4band_arch-124m_1
*)
May sometimes harm instruments less than HP_4BAND_3090_arch-124m, but may leak vocals more in many cases, but generally instrumentals lacks some clarity, but it sounds more neutral vs 500m_1 with mirroring (not always an upside). It’s not available in GUI by default due to its not fully satisfactory results vs models above.
Window size: 272
Aggressiveness: 0.2
TTA: ON
Postprocess: off
mirroring
SP in the GUI models stands for "Standard Precision". Those models use the least amount of computing resources of any other models in the application. HP on the other hand stands for "Higher Precision" those models use more resources but have better performance.
So, what's the best VR arch model?
I'd stick to HP_4BAND_3090_arch-124m (1_HP-UVR) if it only gives good result for your song (e.g. hip-hop). If you're forced to use any other VR model for a specific song due to unsatisfactory results with this model, then probably current MDX models will achieve better results.
Second most usable model for me was 500_m1(9_HP2), and then HP-4BAND-V2_arch-124m (2_HP-UVR) or something in between, but compared to MDX-UVR models, it might be not worth to use it anymore due to possibility of more vocal residues.
- 13/14_SP models (called 4-band beta 1/2 in the Colab) - less aggressive than above
(these are older UVR5 models by UVR team - less aggressive, give more vocal residues frequently’ the mid ones have less clarity, but might be less noisy - but they’re surpassed by MDX models)
- v4 models -
Even older models from times of previous VR codebase
"All the old v5 beta models that weren't part of the main package are compatible [with UVR] as well. Only thing is, you need to append the name of the model parameter to the end of the model name"
Also, V4 models are still compatible with UVR using this method.
“Main Models
MGM_MAIN_v4_sr44100_hl512_nf2048.pth -
This is the main model that does an excellent job removing vocals from most tracks.
MGM_LOWEND_A_v4_sr32000_hl512_nf2048.pth -
This model focuses a bit more on removing vocals from lower frequencies.
MGM_LOWEND_B_v4_sr33075_hl384_nf2048.pth -
This is also a model that focuses on lower end frequencies, but trained with different parameters.
MGM_LOWEND_C_v4_sr16000_hl512_nf2048.pth -
This is also a model that focuses on lower end frequencies, but trained on a very low sample rate.
MGM_HIGHEND_v4_sr44100_hl1024_nf2048.pth -
This model slightly focuses a bit more on higher end frequencies.
MODEL_BVKARAOKE_by_aufr33_v4_sr33075_hl384_nf1536.pth -
This is a beta model that removes main vocals while leaving background vocals intact.
Stacked Models
StackedMGM_MM_v4_sr44100_hl512_nf2048.pth -
This is a strong vocal artifact removal model. This model was made to run with MGM_MAIN_v4_sr44100_hl512_nf2048.pth -
However, any combination may yield a desired result.
StackedMGM_MLA_v4_sr32000_hl512_nf2048.pth -
This is a strong vocal artifact removal model. This model was made to run with MGM_MAIN_v4_sr44100_hl512_nf2048.pth -
However, any combination may yield a desired result.
StackedMGM_LL_v4_sr32000_hl512_nf2048.pth -
This is a strong vocal artifact removal model. This model was made to run with MGM_LOWEND_A_v4_sr32000_hl512_nf2048.pth -
However, any combination may yield a desired result.”
VR ensemble settings
As for VR architecture, ensemble is the most universal and versatile solution for lots of tracks. It delivers, when results achieved with single models fail - e.g. when snare is too muffled or distorted along with some instruments, but sometimes a single model can still provide more clarity, so it’s not universal for every track.
In most cases, ensemble of only VR models is dedicated for the tracks when in the most prevailing moments of busy mix in the track, you don’t have major bleeding using single VR model(s) because it rarely removes that well vocal residues from instrumentals better than current MDX models, or with high aggressiveness it becomes too destructive.
Order of models is crucial (at least in the Colab)! Set the model with the best results as the first one. Usually, using more than 4 models has a negative impact on the quality. Be aware that you cannot use postprocess in HV Colab in this mode, otherwise you’ll encounter an error. Please note that now UVR 5 GUI allows an ensemble of UVR and MDX models in the app exclusively, so feel free to check it too. Here you will find settings for “only” UVR models ensemble only.
- HP2-4BAND-3090_4band_arch-500m_1.pth (9_HP2-UVR)
- **HP2-4BAND-3090_4band_arch-500m_2.pth (8_HP2-UVR)
- HighPrecison_4band_arch-124m_1.pth (probably deleted from GUI, and you’d need to copy this model from here to your GUI folder manually - if it will only work)
- HP_4BAND_3090_arch-124m.pth (1_HP-UVR)
(order in Colab is important, keep it that way!)
Or for less bleeding, but a bit more muffled snare, use this one instead:
HP-4BAND-V2_arch-124m.pth (model available only in Colab, recommended
*on slower Tesla K80 you can run out of time due to runtime disconnection, but you should get faster Tesla T4 by default on first Colab connection on the account in 24h.
Aggressiveness: 0.1 (pretty universal in most cases, 0.09 rarely fits).
Or for more vivid snare if bleeding won’t kick in too much: 0.01 (in cases when it’s more singing than rapping - for the latter it can result in more unpleasant bleeding (or just in some parts of the track). Suggested very low aggressiveness here doesn’t leak as much as it could using the same settings on a single model, but it leaks more in general vs single models’ suggested settings).
0.05 is not good enough for anything AFAIK.
high_end_process: mirroring2 (just ON in GUI)
(for less vivid snare check “bypass”, (not “mirroring” for ensemble - for some reason both make the sound more muffled), be aware that bypass on ensemble results with less vocal leftovers)
ensembling_parameter: 4band_44100.json
TTA: ON
Window size: 272
FlipVocalModels: ON
Ensemble algorithm: default on Colab (min_mag for instrumentals)
Other ensemble settings
- For clap leftovers in vocal stem, check out this ensemble settings.
- For creaking sounds, process your separation output more than once till you get there with this setting
- Also reported clean instrumentals with this setting
Make sure you checked separated file after the process and file length agrees with original file. Occasionally, the result file can be cut in the middle, and you’ll need to start isolation again. Also, you can accidentally start isolation before uploading of source file is finished. In that case, it will be cut as well.
It takes 45 minutes using Tesla T4 (~RTX 3050 in CUDA benchmarks) for these 4 models settings. Change your songs for processing after finishing the task FAST, otherwise you’ll be disconnected from runtime when the notebook is idle for some time (it can even freeze in the middle).
In reality, Tesla T4 maybe has much more memory, but what takes 30 minutes on a real RTX 3050, here might take even more than 2 hours and sometimes slower or sometimes slightly faster (usually slower). So you're warned.
**Be aware that these 4 model ensemble setting with both 500m models in most cases won’t suffice for the slowest (and no longer available in 2023) Tesla K80 due to its time and performance limit to finish such a long operation which exceeds 2 hours (it takes around 02:25h). Certain tasks too much above 2 hours ends up with runtime disconnection, so you're warned.
Also be aware that the working icon of Colab on the opened tab sometimes doesn’t refresh when operation is done.
Furthermore, it can happen that the Colab will hang near 01:45-02:17h time of executing the operation. To proceed, you can click F5 and press cancel on prompt to whether to refresh. Now the site will be functional again, but the process will stop without any notice. It is most likely the same case when you suddenly stop connection to the internet, and the process will still run virtually till you reconnect to the session. But here, you just don’t have to click the reconnect button on the right top. Most likely you have very limited time to reestablish the connection till the process will stop permanently if you don't connect on connection lost (or eventually if progress tracker/Colab will stop responding). So in the worst case, you need to observe if the process is still working between 01:45-02:17h of processing. If you see that your GPU has 0.84GB instead of ~2GB, you’re too late and your process is permanently interrupted, and the result is gone. It’s harder to track how long it processes when you already used the workaround once, and the timer stopped, so you don't know how long it is separating already.
Limit for faster Tesla T4 is between 1:45 and 2:00h/+ (sometimes 2:25, but can disconnect sooner, so try not to exceed two hours) of constant batch operation, which suffice for 2 tracks being isolated using ensemble settings above with both 500m models (rarely 3 tracks).
HP2-4BAND-3090_4band_arch-500m_1 (9_HP2-UVR) - I think it tends to give the most consistent results for various songs (at least for songs when vocal residues are not too prevalent here)
HP-4BAND-V2_arch-124m (2_HP-UVR) - much faster and can give crisp results, but with too many vocal residues for some songs (like VR arch generally tends to)
HP_4BAND_3090_arch-124m (1_HP-UVR) - something between the two above, and can give the best results for some song too (out of other VR models)
HP2-MAIN-MSB2-3BAND-3090_arch-500m (7_HP2-UVR.pth) - tends to have the least vocal residues out of the VR models listed above, but in cost of instrumentals not sounding so "full"
HighPrecison_4band_arch-124m_1 (I think not available in UVR, you'd need to install it manually) - can be a good companion if you only have VR models for ensemble
HP2-4BAND-3090_4band_arch-500m_2 (9_HP2-UVR) - the same situation, I think it rarely gives any better results than 500m_1 (if in even any case) but it's good for purely VR ensemble
_______VR algorithms of ensemble _______
by サナ(Hv#3868)
“np_min takes the highest value out, np_max does vice versa
it's also similar to min_mag and max_mag
So the min_mag is better for instrumental as you could remove artefacts.
comb_norm simply mixes and normalizes the tracks. I use this for acapella as you won't lose any data this way”
Batch conversion on UVR Colab
There’s a “ConvertAll” batch option available in Colab. You can search for “idle check” in this document to prevent disconnections on long Colab sessions, but at least if you get the slowest K80 GPU, the limit is currently 2 hours of constant work, and it simply terminates the session with GPU limit error. The limit is enough for 5 tracks - 22 minutes with ~+/-3m17s overhead (HP_4BAND_3090_arch-124m/TTA/272ws/noppr/~2it/s) so better execute bigger operations in smaller instances using various accounts and/or after 3-5 attempts you can also finally hit on better GPU than K80.
To get faster GPU simply go to Runtime>Manage session>Close and connect and execute Colab till you get faster Tesla T4 (up to 5 times). But be aware, that 5 reconnections will reach the limit on your account, and you will need to change it. It’s easier to get T4 and not reach the limit reconnecting, around 12:00 CET in working days. 14:30 o’clock it was impossible to get T4, but probably it depended on a situation when I already used T4 this day since I received it immediately on another account.
For single files isolation instead of batch convert I think it took me 6-7 hours till the GPU limit was reached, and I processed 19 tracks using 272 ws in that session.
JFI: Even 5800X is slower than the slowest Colab GPU.
Shared UVR installation folder among various Google accounts
Since we no longer can use old Gdrive mounting method allowing mounting the same drive across various Colab sessions - to not clutter all of your accounts by UVR installation, simply share a folder with editing privileges and create a shortcut from it to your new account. Sadly the trick will work for one session at a time.
Firstly - sometimes you can have problems with opening the shared folder on proper account despite changing it after opening the link (it may leave you on old account anyway). In that case, you need to manually insert id of your account where you want to open your link to. E.g. https://drive.google.com/drive/u/9/folders/xxxxxxxx (where 9 is an example of your account ID which shows right after you switch your account on main Google Drive page).
After you opened the shared UVR link on your desired account, you need to add the shortcut to your disk (arrow near folder’s name) and when it’s done, create “track” and “separated” folder on your own - so delete/rename shared “tracks” and “separated” folder and create it manually, otherwise you will get error during separation. If you still get an error anyway, refresh file browser in the left of Colab and/or retry running separation three times till error disappears (from now on it shows error occasionally, and you need to retry from time to time and/or click refresh button in file manager view in the left or even navigate manually to tracks folder in order to refresh), Colab gets changes like moving files and folders on your disk with certain delay. And be aware that most likely such way of installing UVR will prevent you from any further updates from such account with shared UVR files, and on the account you shared the UVR files from, you need to repeat folder operations if you will use it back again on Colab.
Comparing 500m_1 and arch_124m above, in some cases you can notice that the snare is louder in the first, but you can easily make it up using mirroring instead of mirroring2. Downside of normal mirroring might be more pronounced vocal residues due to higher output frequency.
Also, in 500m_1 more instruments are damaged or muffled, though more aggressiveness in the default setting of 500m_1 sometimes makes an impression that more vocal residues are cancelled.
(evaluation tests window size 272 vs 320 -
it’s much slower, doesn’t give noticeable difference on all sound systems, 272 got slightly worse score, but based on my personal experience I insist on using 272 anyway)
(evaluation tests aggressiveness 0.3 vs 0.275 -
doesn’t apply for all models - e.g. MGM - 0.09)
(evaluation tests TTA ON vs OFF -
in some cases, people disable it)
5a) (haven’t tested thoroughly these aggressiveness parameters yet)
HP2-4BAND-3090_4band_arch-500m_1.pth
w 272 ag 0.01, TTA, Mirroring
5c)
HP2-4BAND-3090_4band_arch-500m_1.pth
w 272, ag 0.0, TTA, Mirroring 2
Low or 0.0 aggressiveness leaves more noise, sometimes it makes instrumental cleaner, if you don’t care for more vocal bleeding (it depends also on your sound system how you are able to catch them. E.g. whether you listen on headphones or speakers).
But be aware that:
“A 272 window size in v5 isn't recommended [in all cases]. Because of the differing bands. In some cases it can make conversions slightly worse. 272 is better for single band models (v4 models) and even then the difference is tiny” Anjok (developer)
(so on some tracks it might be better to use 320 and not below 352, but personally I haven’t found such case yet)
DeepExtraction is very destructive, and I wouldn’t recommend it with current good models.
Karokee V2 model for UVR v5 (MDX arch)
(leaves backing vocals, 4band, not in Colab yet, but available on MVSep)
Model:
https://mega.nz/file/yJIBXKxR#10vw6lRJmHRe3CMnab2-w6gAk-Htk1kEhIp_qQGCG3Y
Be sure to update your scripts (if you use older command line version instead of GUI):
https://github.com/Anjok07/ultimatevocalremovergui/tree/v5-beta-cml
Run:
python inference.py -g 0 -m modelparams\4band_v2_sn.json -P models\karokee_4band_v2_sn.pth -i <input>
5d) Web version for UVR/MDX/Demucs (alternative, no window size parameter for better quality):
How to use this free online stem splitter with a variety of quality algorithms -
1. Put your audio file in.
2. Choose an algorithm. Usually, you really only need to choose one of two algorithms:
- The best algorithm for getting clean vocals/instrumental is selecting Ultimate Vocal Remover. Once you selected Ultimate Vocal Remover, select HP-4BAND-V2 as the "Model type".
- The best algorithm for getting clean separate instrument tracks, like bass, drums and other, is Demucs 3 Model B.
3. Hit Separate, and mvsep will load it for you. This means you can do everything yourself, no need to ask for other people's isolations if you can't find them.
6) VR 3 band model (gives better results on some songs like K Pop)
(I think default aggresiveness was 0.3)
7) deprecated - in many cases lot of bleeding (not every time) but in some cases it hurts some instruments less than all above models (e.g. quiet claps).
MGM-v5-4Band-44100-BETA2/
(MGM-v5-4Band-44100-_arch-default-BETA2)
/BETA1
Agg 0.9, TTA, WS: 272
Sometimes I use Lossless-Cut to merge beta1 and beta2 certain fragments.
Models from point 4 surpasses ensemble of both BETA1 and BETA2 models.
(!) Interesting results (back in 2021)
“Whoever wants to know the HP1, HP2 plus v4 STACKED model method, I have a [...] group explaining it"
Long story short - you need to ensemble HP1 and HP2 models, then on top of it, apply stacked model from v4.
Be aware that ensemble with postprocessing in Colab doesn't work.
Instruction:
1 Open this link
https://colab.research.google.com/drive/189nHyAUfHIfTAXbm15Aj1Onlog2qcCp0?usp=sharing
2. Proceed all the steps
3. After mounting GDrive upload your, at best, lossless song to GDrive\MDX\tracks
4. Uncheck download as MP3, begin isolation step
5. Download the track from "separated" folder on your GDrive. You can use GDrive preview on the left.
1*. Alternatively, if you have a paid account here, upload your song to: https://x-minus.pro/ai?hp
Make sure you have "mdx" selected for the AI Model option. Wait for it to finish processing.
2*. Set the download format to "wav" then click "DL Music." Store the resulting file in the ROOT of your UVR installation.
6. Use a combination of UVR models to remove the vocals. Experiment to see what works with what. Here's a good starting point:
HP2-4BAND-3090_4band_arch-500m_1.pth
HP2-4BAND-3090_4band_arch-500m_2.pth
HP_4BAND_3090_arch-124m.pth
HP-4BAND-V2_arch-124m.pth
7. Store the resulting file in the ROOT of your UVR installation alongside your MDX result.
8. Finally, ensemble the two outputs together. cd into the root of your UVR installation and invoke spec_utils.py like so:
$ python lib/spec_utils.py -a crossover <input1> <input2>
the output will be stored in the ensembled folder
9* (optional). Ensemble the output from spec_utils with the output from UVR 4 stacked models using the same algorithm
Ensemble
spec_utils.py allowing ensemble is standalone, and doesn't require UVR installed in order to work. It accepts any of the audio files
mul - multiplies two spectrograms
crossover - mixes the high frequencies of one spectrogram with the low frequencies of another spectrogram
Default usage from aufr33:
python lib/spec_utils.py -o inst_co -a crossover UVR_inst.wav MDX_inst.wav
https://github.com/Anjok07/ultimatevocalremovergui/blob/v5-beta-cml/lib/spec_utils.py
Custom UVR Piano Model:
https://drive.google.com/file/d/1_GEEhvZj1qyIod1d1MX2lM6u65CTpbml/view?usp=s
______________
VR Colab troubleshooting
If you somehow can't mount GDrive in the VR Colab because you have errors or your separation fails:
- Use the same account for Colab and for mounting GDrive (or you’ll get an error)
- If you’re on mobile, you might be unable to use Colab without PC mode checked in your browser settings (although now it works without it in Chrome Android)
- In some cases, you won’t be able to write “Y” in empty box to continue on first mounting on some Google account. In that case, e.g. change browser to Chrome and check PC mode.
- In some cases, you won’t be able to paste text from clipboard into Colab if necessary, when being in PC mode on Android, if some opened on-screen applications will prevent the access - you’ll need to close them, or use mobile mode (PC mode unchecked)
- (probably fixed) If you started having problems with logging into Colabs.
> Actually, it doesn't show that you're logged in while the button says to log in.
So, it should respect redirections in Colab links to specific accounts, but if you're mounting to GDrive, and it fails with Colab error, simply click the button in the top right corner to log in. It will. Just won't show that you did that. Then Colab will start working.
- Don't use postprocess in ensemble, or you'll encounter error
- You can try checking force update in case of errors
- Go to runtime>manage sessions>terminate session and then try again with Trigger force update checked (ForceUpdate may not work before terminating session after Colab was launched already).
- Make sure you got 4.5GB free space on GDrive and mounting method is set to "new". You can try out "old" but it shouldn't work.
Try out a few times.
- If still nothing, delete VocalRemover5-COLAB_arch folder from GDrive, and retry without Trigger update.
On fresh installation, make sure you still have 4.5GB space on GDrive (empty recycle bin - automatic successful models installation will leave separate files there as well, so you can run out of space on cluttered GDrive easily)
- If still nothing (e.g. when models can’t be found on separation attempt), then download that thing, and extract that folder to the root (main) directory of Gdrive, so it looks like following: Gdrive\VocalRemover5-COLAB_arch and files are inside, like in the following link:
https://drive.google.com/drive/folders/1UnjwPlX1uc9yrqE-L64ofJ5EP_a8X407?usp=sharing
and then try again running the Colab:
https://colab.research.google.com/drive/16Q44VBJiIrXOgTINztVDVeb0XKhLKHwl
- if you cannot connect with GPU anymore and/or you exceeded your GPU limit
try to log into another Google account.
- Try not to exceed 1 hour when processing one file or one batch of files, otherwise you'll get disconnected.
- Always close the environment in Environment before you close the tab with the Colab.
That way, you will be able to connect to the Colab again after some time, even if you previously connected to the runtime and stopped using it. Not shutting down the runtime before exit, makes it wait in idle, and hitting timeout. Then the error of limit reached will appear after you'll try to connect to Colab again if it wasn't closed before. Then you'll need to wait up to 24h, or switch Colab account, while using the same Google account as for Colab in the mounting cell (otherwise, it will end up with error when you'll use different account for Colab and different for GDrive mounting).
- New layer models may not work with 272 window size causing following error:
“raise ValueError('h1_shape[3] must be greater than h2_shape[3]')
ValueError: h1_shape[3] must be greater than h2_shape[3]”
- (fixed) Sometimes on running mounting cell you can have short “~from Google Colab error” on startup. It will happen if you didn’t log into any account in the top right corner of the Colab. Sometimes it will show a blue “log in” button, but actually it’s logged in, and Colab will work.
- A network error occurred, and the request could not be completed.
GapiError: A network error occurred and the request could not be completed.
In order to fix these error in Colabs, go to hosts file in your c:\Windows\System32\Drivers\etc\hosts and check if you don’t have any lines looking like:
127.0.0.1 clients.google.com
127.0.0.1 clients1.google.com etc.
It can be introduced by RipX Pro DAW.
- Various Colabs might occasionally get unstable, and the environment disk might get unmounted, or you might get weird errors. In that case, simply kill the current environment and start over
- These are all the lines which fix problems in our VR Colabs since the beginning of the year when new versions of these dependencies became incompatible (but usually one Colab linked is forked when told and up-to-date with these necessary fixes applied already)
!pip install soundfile==0.11.0
!pip install librosa==0.9.1
!pip install torch==1.13.1
!pip install yt-dlp=2022.11.11
!pip install git+https://github.com/ytdl-org/ytdl-nightly.git@2023.08.07
Later in February 2024 we needed to switch to older Python 3.8 in order to make numpy work correctly with used deprecated functions. More details on these fixes and used lines below Similarity Extractor section (all those fixes should be already applied in the latest fixed Colab at the top).