deton24’s
Instrumental and vocal & stems separation & mastering
(UVR 5 GUI: VR/MDX-Net/MDX23C/Demucs 1-4, and BS/Mel-Roformer in beta / MSST/pymms
MVSEP-MDX23-Colab/Drumsep/SCNet/Apollo/MedleyVox
x-minus.pro (uvronline.app | nextgen)/mvsep.com/Colabs
Gaudio/Dango.ai/Audioshake/Iterative ens. Colab)
General reading advice | Discord (ask, or suggest edits only there) | Table of content (up-to-date in Options>Document outline in the app or on PC) | Training | Sep. Colab
Straight to the currently best models list
Instrumentals (ensembles) | Vocals (ensembles) | De-bleeding | Karaoke | BVs | Two singers | Harmonies | Speech | Various speakers | Phantom center | for RVC | 4-6 stems | Bass | Drums | Drumsep | Electric guitar | Acoustic guitar | Sep. both | Lead and rhythm guitar | Piano | Synths | Organs | Bells | Strings | Trumpets | SFX | Bird sounds | De-crowd | De-breath | De-reverb | De-noising | De-clippping | Fast models | Upscalers | Mastering
http://docs.google.com/document/d/17fjNvJzj8ZGSer7c7OFe_CNfUKbAxEh_OBv94ZdRG5c/export?format=pdf
___
Last updates and news
- (uvronlone) “Added BS-RoFormer Mag v2 vocal model: https://nextgen.uvronline.app/#model=bs_roformer_mag_v2_anvuew” - Aufr33
- New best open ensemble for instrumental was found by nextgen.uvronline.app user, and approved by dca:
Mel Inst_GaboxFv9 + BS SW (Max Spec)
Can be used there when post processing options to show the ensemble menu for logged in users is enabled (like described a few messages below).
- Also new best open vocal ensembles were added to the list.
- And even newer ones have been added since then (both inst and voc).
- So FV9 was also added to uvronline:
“According to user reviews, this model is better than the FV10 in some cases.” - Aufr33
- Becruily released his old unreleased Mel-Roformer model serving for cleaning acapella inverts (so when you use e.g. official instrumental with mixture inverted in order to get vocals and it gives some residues)
https://huggingface.co/becruily/invert-clean/tree/main | Colab
“Most regular models don't do the job (...) It removes peaks/pops, clicks, instrumental residue (...)
I didn't focus on things like fullness” - becruily
“Cool model, it's a little muddy on loud/clipped vocals but otherwise it works great (...)
it does work well for leftovers” - gabiigh
“It works great (...) the result is fantastic!” - billieoconnell.
“This has to be [becruily's] best model ever, so useful, sounds so great too... I'm shocked it truly is so good”, consider using it with overlap 4 - gilliaan
- Colab auto-cleaning by billieoconnell - uses Deux result acapella using the invert-clean model above
“Most of the acapellas come out fine, but in some, when I run them through the AI, there's a drop in the bass where the drums hit, making the vocals sound like they're on a bass drum. (...) when I run it through a model on other websites or similar platforms, the result is great.”
- AMD RX 400/500/GFX803/Polaris GPUs owners announcement. As you probably know, there's no native Windows support for ROCm on these older GPUs.
The latest GPU generation officially supported on Windows is with nightly builds here (and it's for GFX900 - Vega and later):
https://rocm.nightlies.amd.com/v2-staging/
a) I don't know if anyone has tried to make e.g. MSST work with ZLUDA with ROCm 5 and HIP 5.6 with modded libs linked in this repo on Windows:
(the option uses CUDA PyTorch instead of the ROCm one).
But it seems like there are two, much faster options than the potential ZLUDA option above, which already exist, worked, and also get much better separation times than UVR using DirectML, thanks to using Vulkan backend:
b) https://github.com/chenmozhijin/BSRoformer.cpp
(it uses MNN-Vulkan)
c) https://github.com/pymss-project/pymss-mnn
(uses MNN Vulkan, can be a bit slower, but still comparable; and optionally also CUDA).
Please report your findings on our Discord linked at the very top.
On Linux for ROCm 6 and 7 you can follow this instruction (similar, but surprisingly not necessarily faster performance than BSRoformer.cpp due to still having some CPU operations on audio. Requires WAV as input).
Optionally there's also https://github.com/nikhilunni/demucs-rs
which uses Vulkan backend (Linux/Windows) and Metal (MacOS; and for VST3 CLAP variant available), but generally Demucs models were outperformed by BS-Roformer SW in many cases.
- Two of the three currently recommended free model vocal ensembles by dca now can be used on nextgen.uvronline.app.
“Both free users and paid users can use it.” - squiddagod
Go to the site, log in, go to Advanced settings > Home Screen > enable “Show post-processing algorithm selection menu”> click “Music & vocals” on the main page>Test>pick vocfv7beta3 or anvuew ft1>pick Max Mag (SW)
- Our separation Colabs started being slow at installing dependencies. The process will finish, but ultimately it will take up to 15 minutes instead of a few like before.
> The issue was fixed by deleting ==2.2.2 from panda dependency (thanks Billie O’Connell) in the OG MK Colab and its fork based on older MSST with a bit more models.
- Ax3l-Ski3s/dr.y0nd3r released a Mel-Roformer vocal model called "Xeno".
https://huggingface.co/DrYond3r/Mel_band_rofo_XENO/tree/main | metrics
The "model is good for Instrumental, all of [the] vocal[s] on Instrumental stem is gone, no bleed" - daylightgay. It’s 3GB, but it gets smaller after cleaning the weight below.
For it to work in UVR, you need to preprocess the ckpt with this script first:
https://github.com/ZFTurbo/Music-Source-Separation-Training/blob/main/scripts/prepare_weights_for_inference.py (it works without it in MSST, but not in pymms Studio which needs modifying worker_models.py to be detected, edit. said to be fixed in 0.0.12 version;
the attached pkl and json files at the top won't help here).
Usage of the script (install Python, and on fresh installation execute:
“pip install torch” then execute the script):
python C:\path_of_the_script\prepare_weights_for_inference.py --checkpoint "path_to_the_input_model/model_to_clean.ckpt" --output_file "path_to_the_output_model/cleaned_model.ckpt”
“I experience vocal leakage on certain tracks, such as ABBA's "Dancing Queen"—an issue I do not encounter with the v1e plus model.” - aureyoboss
- Лямбда/necromunt released a new BandSplitPolarFormer Lead Vocal model | metrics
https://huggingface.co/Lambda001/BandSplitPolarFormers/tree/main | feedback
(“Yaml is not needed for this model. And it won't work in UVR.“, but some separators will still need it - then use this one)
If you have “unexpected keyword argument 'use_pope'” you must update MSST:
!rm -rf /content/Music-Source-Separation-Training
!git clone https://github.com/ZFTurbo/Music-Source-Separation-Training
“and you must reinstall the main branch's requirement.txt. (before it, edit requirements.txt to remove wxpython)” - Essid
wxpython is for GUI.
Alternatively instead of updating MSST, use this bs_roformer.py file and execute
pip install pope-pytorch
- You might want to check the new denoise method by n0blink/hitto1st called:
Phase-corrected de-noise (MVSep paid preset | workflow) -
iirc it uses the Gabox/Debleed, with phase from the Mel Aufr’s model aggressive variant.
Allegedly better than using just one of these on their own.
It might remove claps from live videos.
- (uvronline) “added the Max Mag (SW) ensemble for Gabox flowers v10 model.
It also includes phase correction.” - Aufr33 (probably for paid users)
- dca100fb8’s new ensembles for instrumentals, vocals and karaoke added
(order of the community inst. and voc. ones rearranged), one new inst #2 added.
- (MVSEP) Sharing preset links when you click the export button in https://mvsep.com/presets was added.
You can also make your presets public in the gallery:
https://mvsep.com/presets/gallery
Currently saving and using presets with more than two nodes require premium credits enabled.
E.g. you can check out two new dca's presets for instrumentals ensembles (more):
- Becruily 124 fullness phase fixed + SCNet XL IHF fullness (Avg): https://mvsep.com/presets?hash=QQoUWhQKDTtPhKDB
- Becruily 124 fullness phase fixed + Deux phase fixed (Avg): https://mvsep.com/presets?hash=i1x5OagBY3ajVa8I
Possible issue:
“I used the presets you sent and ran all 3 models in parallel at once, but when I checked my credit history, the algorithm name doesn't match the preset name.”
- (MVSEP) ”I added a new model in the BS Roformer (vocals, instrumental) section, created by becruily, with high instrumental fullness. It's available with name:
`becruily high instrum fullness 124 bands (SDR instrum: 18.47)`” - ZFTurbo
Vocals bleedless: 36.30, fullness: 21.38, SDR: 12.25
Instrumental bleedless: 44.29, fullness: 34.76, SDR: 18.47
“It’s dual model so vocals are also fullness, but whether they’re better or not than the original level 1-3 I don’t know”,
“It’s a fine-tune of [BS-Roformer 124 band] ZFTurbo's model, so unfortunately private” - becruily
It’s also paywalled like the previous model due to increased separation time from the model architecture. Iirc 1 credit per minute.
“This model is dope”, perfect mix of fullness and bleedless - dynamic64
Q: How does it compare to deux and hyperace, any insight?
A: “I like it more”, “I feel like it functionally replaces the bs rofo 124 fullness model” - -||-
For vocals it has much better bleedless metric than Deux, but less fullness.
“the instrumental fullness model makes cleaner acapellas than every other model
and the instrumentals sound so perfect its like they were phase inverted with the acapellas” - nymphi/nimfia
“For sure [it] sounds better than deux, but it has some difficulties keeping some instruments in instrumental compared to deux which is excellent at it (...) But yeah the model is amazing in how it sounds. Very high quality sounding especially with phase fix (...) [which] is optional.
I am still (...) going to use deux for some bands like Radiohead tho. Keeps more instruments as I said” - dca100fb8
“In my opinion brecruily high inst fullness vocal result seems to give the same result of BS Roformer fullness level 1” - mrmason347
“Fullness and bleedless are not very reliable. In my opinion it’s l1_freq and aura_mrstft both measure fullness and bleedless at once better (l1_freq has the edge because it also can show if model has phase issues like some fullness models)” - becruily
- Phase-corrected deux was added as a free preset in MVSep.
- You can now run presets from the API on MVSep (more)
- Phase fixer for less noise/vocal shells in instrumentals was added as a preset called "phase corrected instrumental” on MVSEP (the previous one called like that was wrong, the old one has its name changed).
This specific preset is currently only for premium because by default it uses HyperACE V2 inst with (currently paywalled) BS-Roformer 124 bands and by default avg algorithm is being used.
But you can just click “Editor” option in the presets menu, and pick “phase corrected instrumental” preset on a free account there, and on the displayed graph pick whatever models you want for phase fixing instead (e.g. becruily vocals (even instead of BS-Roformer 2025.07), which might work better for phase fixing than BS-Roformer 124 bands, and is free model), and maybe run the job from there. Solely presets functionality and the editor is not paid, but such a modified preset on your main MVSep page will be paywalled.
“If 2+ algorithm nodes, it is premium only”.
Currently you can only save your preset on your account and sharing is possible only by screenshotting or copying the graph code manually and sharing on your own.
- Algorithm used for ensembles is avg by default.
- Automatic phase fixer for chosen models (with inference) should be possible also in SESA Colab
- A new, faster MSST alternative for inference of audio source separation models came up.
In the Sucial’s MSST-WebUI repo we can read that “This project will soon stop being maintained. We recommend using pymss-desktop [now called [pymss-studio], which supports more comprehensive models, provides 50-100x faster inference than MSST-WebUI, and has a better-looking GUI.” And there is also pymss used as a base (CML/webUI).
It’s a new project by Baicai and Sucial (and contributors).
The Studio has installers pre-configured for CUDA, Mac (MLX) and CPU and Linux (pip install pymss)
But it turns out that pymss + WebUI (pymss-core was tested, we don't talk about the Studio with GUI) can also be managed to work with ROCm on AMD GPUs on Windows with great success (BS-Roformer-SW on a 6:20 minutes track took 1:24 minutes on RX 6800 XT. 36 seconds on the same song, with becruily's vocal model, although it's still slower than e.g. 4090 in UVR or MSST).
The below was tested on RX 6800 (6800 XT / 6900 XT / 6950 XT - so gfx1030-compatible) and the latest Adrenalin driver on Python 3.12. Use the following command to install ROCm on pymms (it will be detected as CUDA device):
pip install torch torchaudio --index-url https://rocm.nightlies.amd.com/v2-staging/gfx103X-dgpu/
(or check here if the link above becomes offline, or visit here to choose your GFX other GPUs than gfx103X, and paste the link in the above instead - untested; which GFX is which GPU; you can also scroll down MSST section to find older RX 7600 / XT guide which might be useful in case of issues with the method above, and even some RX 400/500 guides are later below there).
Be aware that configs for models trained with MSST which are not already in pymms use num_overlap instead of the newer overlap_size in pymms which gives the performance gains. Using the old num_overlap makes pymss fall back to 50% overlap (equivalent of 2 in UVR) which is several times slower than the new overlap_size. Models available in pymms (iirc non-core packages) are already configured for the new overlap. While adding a custom model, Claude suggested setting overlap_size to ~5 % of chunk_size (e.g. overlap_size: 24000 for chunk_size: 480000).
PolarFormer is currently unsupported.
Full CuƐ’s guide to create isolated environment for pymms + ROCm is here (with claudette pinpointing problems; thank you for bringing up the topic).
You can also use ROCm branch in non-Studio packages with compiled source on Linux or using WSL on Windows instead (but I can't tell which one of the three will be the fastest, but differences shouldn't be big).
There's also a separate pymsst package dedicated for training, and pymss-ara - VST3 plugin for separation inside your DAW in a separate process (currently Windows and CUDA supported).
- (MVSEP) “We have just released the best model currently available for vocal separation in terms of metrics: BS Roformer 124 bands”.
Vocals bleedless: 39.18, fullness: 18.11, SDR: 12.33'
Metrics with 3 fullness levels | here too
Link to the model option: https://mvsep.com/home?sep_type=40
Be aware that it only has 1 inst fullness stem.
“Because it is several times slower than the original BS Roformer (which is one of the most used models on the site), we have temporarily placed it behind a paywall [6 credits for 6 minutes song] to ensure our servers can handle the load. We plan to make it freely available later once we add more servers. The model can provide several levels of fullness for vocals and instrumentals out of the box (you need to choose it in the form).” - ZFTurbo
“level 3 is unusable (...) too much bleed of instrumental while there is singing. Particularly drums” - dca100fb8 (more)
Q: level 1 might be best A: Yeah I think so. For vocals - dca.
“Tested more on vocoder tracks (Daft Punk) and complex vocals such as yodelling (Deep Forest) and the model is good but still have some issues handling them”
“+0.4 SDR is not a small jump, I mostly check the spectrograms and frequencies appear quite more defined.
So more accurate in preserving the correct instruments overall, sound wise mostly similar to everything else non-fullness (...)
There might be something wrong with the fullness inst output, it appears to be muddier than [the] baseline.” - becruily
“So looking at this, Roformer level 1 fullness has about the same bleedless as polarformer (non fullness), but more fullness
and the polarformer fullness model is basically irrelevant now because of the extremely low bleedless score is compared to the fullness tradeoff (...) Metrics wise it looks like there is no good reason to use level 1 for instrumentals we gain 2 points of fullness and lose like 8 bleedless (...) I think though hyperace has its issues I prefer the sound of the fullness (...) oh my god fullness level 1 sounds IMMACULATE for vocals its like the vocals arent even TOUCHED by the drums - dynamic64
Q: Can the 124 BS-Roformer model make insts that have the fullness of deux/hyperace? I'm trying to avoid most of the muddiness. I don't have premium atm or I would just test myself lol
A: “No, not even close metrics wise tbh” - dynamic64
“SDR is good, but like Polarformer I am not really impressed with the occasional muffling. I think SDR by itself is not really a metric for how a separation model really sounds, and Deux is still my all time favorite model for that reason.” - analogspiderweb (more)
- (MVSEP) “We added 6 new algorithms:
- MVSep Clap (Class: Percussion)
Demo: https://mvsep.com/result/20260715164440-f0bb276157-mixture.wav
- MVSep Cowbell (Class: Percussion)
Demo: https://mvsep.com/result/20260715165426-f0bb276157-mixture.wav
- MVSep Vibraphone (Class: Keys)
Demo: https://mvsep.com/result/20260715165745-f0bb276157-mixture.wav
- MVSep Metal Bars (Class: Keys)
Demo: https://mvsep.com/result/20260715165821-f0bb276157-mixture.wav
- MVSep Rhodes (Class: Keys)
Demo: https://mvsep.com/result/20260715165856-f0bb276157-mixture.wav
- MVSep Whistle (Class: Wind)
Demo: https://mvsep.com/result/20260715165944-f0bb276157-mixture.wav” - ZFTurbo
- It looks like Huggingface acts weirdly in the last days affecting the functioning of our separation Colabs (errors like "Could not find [model name]” / “PytorchStreamReader failed” / model download stopped at 0 or more KB) and probably only Google’s machines/IPs.
In case of such issues, kill the environment in options, and start over. If still the same, maybe even clear cookies and/or use VPN, but it seems like everything might have got back to normal already.
- A new report about 403/Gateway timeout errors from HF has appeared.
Besides the solutions already mentioned,
you can download the model on your own machine, and upload it from your computer/smartphone to an already mounted and initialized/installed Colab, and drop it to ckpt folder replacing existing unfinished model files, but it will take some time when uploading via the file manager in the Colab. So you can upload it on your GDrive via web or app first, and then drag and drop it to ckpt folder using file manager in the Colab from the mounted GDrive folder there. But you will need to upload it every time the Colab is reinitialized, into GDrive that way.
So you must run mount to GDrive and Install cells first, before doing anything, then pubthe files to ckpt folder.
Once the script sees already existing files in the ckpt folder, it should skip downloading them from HF or anywhere it's hosted on.
You can start with inferencing 1297 model hosted on GH which should work, to make sure the files are in the right path with 1297 model files inside, then you can put your model files there, and keep doing it from now on without this sanity check when you know the right path.
Also, it turned out that at least one Colab has parameter “replace_exist=True” in the code, which means it will always download from source. Then you'll additionally need to set it to false.
Or you could reupload the weight to GitHub to the releases page of any repo, and paste the link to it into Colab code, or custom model import Colab. Just be aware that the yaml link from GH must be possessed from the RAW option when you open it from the repo and uploaded in the repo, not from the releases page, otherwise it used to error out the Colab.
- anvuew BS_Roformer_mag_v2 was released some time ago (17th June). DL | Colab (placed below deux and 1297)
Vocals bleedless: 27.50, fullness: 26.22, SDR: 10.74’
Mag so “specialized for magnitude spectrum accuracy” or RVC.
“Too noisy and SDR drops a lot, so I i just shared it on GDrive instead of HF” - anvuew
“generates a metallic noise” - adrielmz_
The older mag had 22.15 fullness, and it's the only metric which increased in the V2.
Still, it turned out to be a favourite model of one of our members (5b), so it probably needs some recognition and placing here anyway.
- (MVSEP) Editable beta presets for specific processing of models and creating complex audio pipelines added.
E.g. currently you can use:
- phase-corrected instrumental (it says it uses - phase inversion for vocal result of 2024.10; edit now it should work correctly, but with 124 band BS-Roformer - editable).
- Denoise instrumental (double pass with -6dB and polarity flip on one 2024.10 result)
- Karaoke pack (keep backing vocals; uses MVSep Karaoke)
- Clean Acapella (2024.10 denoised with Aufr33 standard model)
And there's also a presets editor to create completely new ones (or edit existing ones).
You need to be a registered user to use it.
"Algorithm nodes will take an ordinary amount of credits if you are premium" but other than that there's no credit cost for using the preset feature itself (unless it's specifically paywalled).
- (MVSEP) “I've added a new DeReverb model: "MVSep Team (2026.07, BSRoformer)". It's a universal model, so you can remove reverberation from any stem or vocal track. Currently, it gives the best results for both of our validation datasets.” ZFTurbo
Check out the images | Link to algo | Demo
Compared to anvuew room model it did nothing to the room echo - real_dyeu_kami
“I used it on an orchestral track (Ba'ku Village from Star Trek Insurrection), and it worked great. Lower volume reverb seems normal to me, as it's usually a distant reflection off walls etc.” - fal_2067
“i tried it on a full song at first just to see and the result was suuuuper quiet and not really accurate
so then i tried it on a vocal track just now
once again the reverb is like 36db quieter than the vocals
I can hear so much reverb in the dry part in my examples” - dynamic64
- gilliaan released SCNet CenterWide v2 by model | config (gilliaan_scnet_centerwide_v2) | Colab
SDR center: 18.17, wide: 18.05, full metrics.
This time based on other arch “Better SDR over the previous one, worse SDR than the MDX23C v2e”
But check it out in ensemble:
- Avg/Avg of SCNet V2 by gilliaan + gilliaan_mdx23c_centerwide_v2e
SDR center: 19.38, wide: 19.68, full metrics
“I find it really good in an ensemble (...) it adds back some fullness and sounds better”
Out of the four currently evaluated ensembles, it currently holds the highest center SDR (click).
“I think I've really pushed the models to the very limits of their capabilities; there's nothing else left to try. I hope you’ll like the models!
it's not a farewell, it's just a goodbye for now! (...) “SCNet v1 was a finetune of SCNet XL IHF” [so the v2 is further trained with probably some dataset changes]. There might be from scratch model on 5K dataset in the future - gilliaan
“tried the ensemble, the center channel is great” - isling
- unwa released a new BS-Roformer-Leap Xe vocal and instrumental models. It's 90-band.
https://huggingface.co/pcunwa/BS-Roformer-Leap/tree/main/Xe | Colab | feedback
Vocals bleedless: 37.39, fullness: 18.27, SDR: 11.76
“Doesn't sound full enough to me, and it is not catching some BVs (...) It sounds good tho” - dca100fb8
“It sounds good so far.” - rage313_
“Pretty great” - pixelrino
Inst bleedless: 41.46, fullness: 35, SDR: 17.53’
A bit better bleedless/fullness, but a bit worse SDR vs Deux (potentially a bit worse at recognizing instruments).
“Big issue (...) is BV bleed, sometimes (...) kept almost intact (...) Very good model in how it sounds like (...) Fullness is very good, sounds fuller than deux and hyperace, but less full than Flowers. Haven't found any tracks where it's muddy. (...) most of the tracks I tested it on, it really has almost no bothering noise (...) phase fix doesn't need to be always applied (...) quite song-dependent, which is not the case with deux (...) For keeping instruments, it's better than hyperace, but worse than deux. (...) Considering SFX vocals sometimes [they’re] kept and sometimes [they’re] removed, so I don't think it has been improved since previous models. Deux was also not good at it, while Flowers shines for this.” - dca100fb8
“sometimes leaves a bit of faint "crust" behind but still really fantastic stuff here” - pixelrino
“Mids sound a tiny bit muffled, but there's a lot of hf noise from about 7 kHz upward. I can hear what it's going for, it's trying to prevent the annoying kind of muddiness you get with a really bleedless model, but it's putting a lot of noise in a region I'm sensitive to.” - Musicalman
“a test of the new loss function.
It looks like it still needs some improvement.” - unwa
“leap: 62 bands
xe: 90 bands
In xe, I've increased the number of bands in the mid- to high-frequency ranges.”
- Jasper released a new 4 stem SCNet upmixer model
https://huggingface.co/jazzpear/audio_separation_models_trained_by_jazzpear/tree/main/surround | Colab (placed at the very end) | opinions
“Experimental model. It takes stereo music and cinematic content such movies, anime, TV shows, etc., and then separates the audio into 4 stems:
LR/Front stereo
S/Sides stereo
LFE/Sub frq
C/Center front channel”
“Mono 1-ch audio is NOT supported.
Dual-Mono 2-ch is also very much not supported as it'll mess up the placement and sound off to the ears. but like... it does do SOMETHING on it. Just not what it's meant to be LOL."
"Made a script for my surround model for merging the stems into a multichannel audio file!”
“Unlike those traditional stereo to surround methods this is trained on actual native 5.1 surround or 6-ch mixes segmented into training channels/stems.
Thus things like what types of sounds often bleed more or less between the space alongside the stereo field being taken into account what you end up with, is very good surround field separation and a proper center channel that isn't solely a mid side isolation muddied with bleed and tinny sound.”
More about the model usage:
“You take the stems and place them in a surround stem based project in any DAW that does surround mixing accordingly and then render. The result is a very easy upmix without the usual simple center extractor dilemma where you have mid side split files and just spread or expand the sides through the channels other than center.
Dry center-panned voices are center alongside mono SFX.
Wide stereo vocals and immersive SFX and instruments are placed in the front and back stereo pairs beautifully.
Once again this only gets the files ready for combining back into a multichannel file. You will have 4 stems. NOT a final surround multichannel wav. The combining of the stems back to a mix is up to you. But this speeds up mixing and takes manual sound panning out of the picture.” - Jaspear
Or you can just use the script for it.
The list of popular upmixers plugins (one is now free) - click
- (gilliaan) “New CenterWide model y'all!! MDX23c dual fullness model.
[model | config | MK Colab fork #2 | Metrics]
Yaml link fixed. “Trained using 2x more data over v1.”
The best overall, best SDR... best fullness (across SOTAs, not scripts), sounds like it's actually catching center/wide with more precision. It's also less noisy.
Even better SDR with this config on: batch_size: 16 - overlap: 16 - use_tta: true” metrics
“FANTASTIC improvement on this one. This has been one of the biggest jumps in objective quality I've ever seen from a model” - dynamic
- (MVSEP) “I've added a new BS PolarFormer model (124 bands, SDR: 12.02).
[Vocals bleedless: 39.26, fullness: 17.18, SDR: 12.02]
It currently achieves the highest SDR on the MultiSong dataset. It also introduces a new feature: generating two additional stems with higher fullness
[voc. bleedless 32.37, fullness 20.95, SDR 11.96].
I plan to add more fullness stems in the future. Also, it's very computationally heavy, so it might cause long queues. I will see how it goes.” - ZFTurbo
Model on MVSep | Metrics / fullness level 1 | Description | Demo 1 | Demo 2
“Can be lacking some fullness, it's tuned for low noise.” - rainboomdash
“Very good for lead vocals. Unfortunately, it doesn’t sound good on backing vocals or choir vocals” - myself044. Sometimes deux, sometimes this picks up BVs better - pezz23
“On the contrary for me it works best to remove all vocals” - duhh_its.chris
So it might be song-dependent.
“there's not a huge dramatic difference between it and [deux] (...) There is definitely a difference, but you probably wouldn't notice it based on only one listen if you were directly comparing them (...) results sound realistic” - pezz23
“the inst will always remain muddy no matter which fullness level it will be set to, since it's not the target instrument.
It sounds like phase remix demudder output, also what makes me think it’s a vocal model, is just how the inst sounds” - dca100fb8 (might be the case since the public PolarFormer is vocal target, indeed - squiddagod).
“This model can extract Daft Punk vocoder cleanly” - RME
- unwa released BS-Roformer-Leap vocal and instrumental model | metrics
Vocals bleedless: 38.90, fullness: 17.07, SDR: 11.72
MK Colab fork #2. The highest vocal SDR for the public model, lower than older MVSep PolarFormer 62 bands models, and higher metrics than the old Mel-Roformer 2024.10 (Bas Curtiz model fine-tuned by ZFTurbo, both only on MVSep.
“I conducted further training using 6000 Ada, which is relatively inexpensive (...) trained on a small dataset of less than 30GB.
The quality of the dataset is important.
I checked it carefully.
Now all that's left is to increase the volume.” - unwa
“BS Leap vocals feels a bit fuller than “[beta] resurrection v2, and it does a really good job picking up background vocals” - neoculture
The inst model “leaves a bit of a noise around 6 kHz or so” - Flangermaster
“it sounds like there's a ton of fullness in high freq but more filtering in mid freqs. Imo the signature is kinda the inverse of deux.” - Musicalman [sonically, not like you were to invert something on your own]
"inst model isn’t good" - gabox
"I found the vocals to be a bit muddy, but the instrumental is alright!" - gilliaan
"I think it['s] better on Vocal than Instrumental"
"Inst model doesn't sound right on my side (to the point VitLarge23 is better) - dca100fb8
- Temporary Colab dedicated for MDX23C models was fixed.
- BS-Roformer SW in the OG Makidanyee's Colab fixed.
- unwa new bs_resurrection_v2_quality_test (a.k.a. resurrection v2) vocal model in the testing phase. Please share your opinions about models in #model-sharing on our Discord server.
“It’s a really good vocal model. I compared results of a song to anvuew mag and it doesn't keep as much reverb but still very full result” - 5b
“the model is great for capturing side channels of the vocals.” - crazyboy4.8.8.1
“working well on songs I always test” - Gabox
“Sounds sometimes fuller than some models I tested, it needs a little bit of improvement in my opinion, but it is already really good. There is still extremely tiny non-vocal artifacts that i found in the separated vocal track sometimes” - dr.y0nd3r
Fine-tuned on rented RTX PRO 6000 WS.
- You can now find fully functional custom model import Colab here:
(no BigShifts support)
- You can now find fully functional manual ensemble Colab here:
https://colab.research.google.com/drive/147W6wty4sUkEXi6xHhMXDqW7M-xINjl1
Max FFT works correctly (thx squiddagod). GPU support can be probably disabled in the Environment options. Possible AttributeError: module 'gradio' has no attribute 'blocks' error. Then check whether this one works (v2). It has forced output 24-bit FLAC quality despite the chosen setting.
Or if you find the GUI irritating, use the oldschool one now functional here (none of these issues occur here).
- If you noticed that certain model links are offline now, here all are mirrored with configs:
https://huggingface.co/noblebarkrr/mvsepless_resources
https://huggingface.co/miercolesv/MVSEP-Central-Backup (it’s older by 5 months, but contains at least a few models the above lacks)
Use mvsepless Colab which is still fully functional:
Makidanyee's fork now has all models fixed:
(besides MDX23C ones; it also allows manual ensembling besides max FFT for now and model queuing during inference and overlap/chunk_size settings might not always work):
https://colab.research.google.com/drive/1IC6Q1hLF55_tK6mhky0SWYKGVF9T5WsY
Fork of Makidanyee's has MDX23C models fixed exclusively (older base, use only MDX23C [some other models will be offline]):
Custom model import Colab is offline for now. (see above)
All Colab and model links in the document outside of the news section were fixed.
- “My musicless sfx and voice isolator is on my Huggingface now” - Jasper
Q: It removes the music and leaves sfx and voc?
A: “Yes. Thus if you want no voice, you gotta remove voices first or after.” - Jasper
- “Added vocfv7beta3 model (to the Test section)
https://nextgen.uvronline.app/” - Aufr33
- The mvsepless Colab now has newer link:
https://colab.research.google.com/github/noblebarkrr/mvsepless/blob/dzeta/MVSepLess_Dzeta_Colab.ipynb
- (MVSEP) 1) New Guitar model MVSep Pedal Steel Guitar has been added.
Description: https://mvsep.com/algorithms/131?lang=en
Demo: https://mvsep.com/result/20260519131702-f0bb276157-mixture.wav
Use algo: https://mvsep.com/en/home?sep_type=124
2) New FX model MVSep Risers has been added.
Description: https://mvsep.com/algorithms/132?lang=en
Demo: https://mvsep.com/result/20260519132408-f0bb276157-mixture.wav
Use algo: https://mvsep.com/en/home?sep_type=125
3) New Mega 53 Stems Model has been added. It's located in the Experimental section.
Note 1: The model outputs only those instruments that were detected in the musical composition. Instruments that are not present in the track are not included in the output.
Note 2: Individual models for each instrument generally produce better results than this multimodel. Therefore, it is recommended to use this model to determine the set of stems first, and then extract them using separate instrument-specific models trained on narrower tasks.
Note 3: To reduce disk space usage, results are saved in all formats except WAV (FLAC is used instead of WAV).
Note 4: This model differs from our previously published open-source model and is an improved version of it. The comparison table is provided on the algorithm description page.
Description: https://mvsep.com/algorithms/135?lang=en
Demo: https://mvsep.com/result/20260519132948-53be20aa17-10seconds-song.wav
Use algo: https://mvsep.com/en/home?sep_type=126
- We figured out the fix for certain models not working on UVR’s RTX 5000 patch giving:
ModuleNotFoundError: No module named 'torch._dynamo'
You need to go to model’s yaml and edit the line:
“use_torch_checkpoint: True”
to False, or delete it.
To find out which yaml is associated with your model, scroll down the model list, and go to Edit Models Config>Change Parameters>Edit Model Param option, and the corresponding config name will be shown at the top.
Then you should be able to locate this yaml in: models\MDX_Net_Models\model_data\mdx_c_configs
Thanks to Kashi for testing.
- Validation for different types of reverb models was added on MVSEP. More
- neoculture released a fullness instrumental model called neo_inst_beta | yaml (fixed)
"(finetuned from BS Resurrection inst), trained on a small dataset (only 19 hours).
The vocal FX didn’t disappear completely, but the actual singer vocals are fully gone. (...) I tested the model again with a few more songs, and yeah, there’s still some vocal bleeding”. Might be noisier than Resurrection inst - Musicalman.
It was epoch 23, and the 21st was used.
“Lowering or increasing the chunk size value can help reduce vocal bleeding. I tested both 705600 and 882000. personally i prefer 882000 since it fits my taste more (you can try testing with your own chunk size values)
(...) I changed the yaml chunk size on my HuggingFace to 882000 (...) I thought only my model had bleeding issues, but Resurrection has them too” - neoculture
- If you occasionally use YouTube for separations of some rare music not available on other platforms, test out introC’s AAC/Opus combiner Python script with Colab (zipping feature added). The output is 44kHz (so Opus is downsampled).
If you find the results muddier than the OG Opus file, convert from 32-bit float to 24-bit.
Matches “QAAC -TVBR 86 quantization noise cancellation with the source audio”.
It works on top of yt-dlp, so on provided URLs, but the code can be made to work with downloaded files too (more).
(I must have confused the author's name with IncT before.).
- Dry Paint Dealer Undr released MelBand Roformer Duet model / mirror (Singer 1 and Singer_2)
“Only works on isolated vocals”. It doesn’t work with the RTX 5000 UVR patch.
“It's not perfect but it can work, although how well it works really varies from song to song.
I was originally going to hold off on releasing this to see if I could get it any better but I saw people wanted a model like this and thought I should probably just release it now.” - Dry
“Wow, this is surprisingly good. The model has slight bleed from the backing vocals but less than the karaoke models. Well, about the same. Still has quite a bit which is what i was expecting (...) it has potential” - isling
- (Uvronline) Added two models:
- Revive 3e by unwa
https://nextgen.uvronline.app (Test section)
- SFX Splitter by jasper (jazzpear)
https://nextgen.uvronline.app/experimental/” - Aufr33
- The old links to the testing version of the site now redirect above, and previously given credentials there are no longer necessary.
- The SFX model “separates SFX tracks into background and foreground SFX.
So say, you isolated SFX with DnR from a full mix or have a single all SFX track already. You run this and get separated version with BG and FG tracks for easier remastering of the mixed version SFX" - jazzpear
- IntroC released Iterative Ensemble Colab
https://colab.research.google.com/github/Qupci/Iterative-Ensemble/blob/main/Iterative_Ensemble.ipynb
By default it uses models from the end of 2025 (V1e, Resurrection, Revive3e, SCNet), but it's a subject for further development. It performs many iterations of separations with slowing down and mid-side processing for the ensemble. As the final step called "finisher" uses MVSEP's BS-Roformer model - you need to provide your MVSEP API key from here when you log in to the site to make it work (it can be a free account).
With 4 minute file, and 1 or 0 current MVSEP users in the queue (while usually 14 files normally being processed at the same time, e.g. 3:30 AM EU time), the whole Colab execution can take around 20 minutes.
First, you must create “output” and “input” folders on your GDrive manually.
In the current version, the output will be saved to "[filename]_mvsep_only_ensemble_side.wav" in "GDrive\output\mvsep_only" folder, and end up with "Finisher variants processing complete”. Progress cannot be tracked in real-time. MVSEP model processing starting with [DEBUG] is started currently near the end.
You should also abbreviate your input file name to a short title at times (if it has a lot of hyphens at least).
You could also try unchecking "amplify_masked_details" for less residues, although it makes e.g. snares more muddy, and not so punchy.
But maybe even ensembling both would be beneficial, it also has its own flavor.
And surprisingly, it left one quiet vocal residue the setting unchecked by default didn't have, despite drums being a bit muddier, but not even in all places, so cutting it and piecing together might work too.
Also, the option seems like to make all scratches in hip-hop and their vocal chops much quieter, so in such cases consider separating with both settings, and ensure the old result file won't be replaced (e.g. rename it at best).
“You should remove the checkpoints folder each time you want to process something new or if you want to resume (unless you already got a final iteration; it's not broken there)”
You can just open the file manager in the Colab, and drag and drop the checkpoints folder to any other folder (you cannot delete it from the manager).
It takes ~3.7GB for intermediate files called "checkpoints" on GDrive if you want to save them for testing purposes (the Colab makes many iterations of the same file for the same model).
It can also happen that the next task will cause an out of memory error, although it won't be stopped. Then stop it manually and wait a minute or two and retry.
If you're on the phone, you can't switch tabs or let the screen go to sleep during the process (change it in phone settings), otherwise the environment will be shut down.
Colab quota for previously unused accounts is 5 hours, so you should be fine with leaving the process overnight with your phone on the charger.
FYI, the Colab seems to consume more than 6GB VRAM (with possible spikes to 10+, can't remember). More.
- (MVSEP) “Glad to introduce a new model based on a modified BS Roformer → BS PolarFormer. It uses a different type of embedding based on polar coordinates, which allows it to handle long contexts like other MSS models (see the figure below).
It achieves strong metrics, though slightly below the best available BS Roformer: 11.76 vs. 11.89. Still, it is now the second-best architecture when ranked by maximum SDR quality.
Algorithm: https://mvsep.com/home?sep_type=123
Description: https://mvsep.com/algorithms/128?lang=en
Demos: https://mvsep.com/en/demo?algorithm_id=123” - ZFTurbo
Vocal bleedless 35.90, fullness 17.68, SDR 11.75.
So it appears to be further trained from the public weight with SDR 11.00 released before.
- One more model from mvsepless repository slipped through:
mel_band_roformer/mbr_denoise_yuluoye (yaml), and also:
denoise children phaedrus33 model | yaml (only mentioned on Discord once)
- New mdx23c model_sfxsplitter
model (yaml) by jazzpear/Jasper was released (224MB). It's new, and different from the MDX23C model (yaml) found previously in the mvsepless repository (1.3GB), which is a model from February 2026 already included in this doc before.
- (fliegenfuerst) I've made ChatGPT churn out a GUI to use the 53 stems selectively | DL/pic
- (MVSEP) “Three new models are now available:
1) In algo "DeNoise by aufr33 and gabox" new model "gabox"
2) In algo "Phantom Centre extraction" new model "Phantom Centre by gilliaan (mdx23c)"
3) In algo "MelBand Roformer (vocals, instrumental)" new model "gabox v10 flowers (SDR vocals: 10.67, SDR instrum: 16.97)"” - ZFTurbo
- iZotope released RX 12 with new features like Scene Rebalance, and also denoise, Dialogue Isolate, dereverb, and new models for source separation in Music Rebalance. It uses three algorithms: “real-time”, “better”, “best”. The last probably uses 48-band BS-Roformer judging by iZRX12Core.dll content. Despite it’s Roformer, it’s surprisingly fast, although still doesn’t use GPU (e.g. Ableton Live does). The quality is now better than in RX 11, even good for some songs, but “except for very low subs. They aren't being sep'd out of the rest” and also it’s still bad for some specific songs (CuƐ), so a hit or miss. SDR for vocals for the “best” model is similar to the old UVR-MDX-NET Main 427 vocal model (metrics, search). The real-time algorithm might be Spleeter (references to it returned to their licence files), and the “better” might be Demucs-based ~jarredou (thx!)
- dca100fb8 would like to present to a broader audience, an ensemble of two instrumental models, which so far, turned out to be the best out of many created so far here, and received good feedback. Read how to manually ensemble separations from MVSep here.
-> deux + FV10 (Avg)
The new best.
“Deux is great at everything and its only downside would be it's not full enough for some users (can sound "fuzzy"), so Flowers permit to fix that issue. This ensemble doesn't add bleed because Flowers has even less vocal bleed than deux.”
->Mel deux + MVSep SCNet XL IHF high fullness by becruily on MVSep (Average Spec/FFT Wave)
Former best
“Permit a cross arch ensemble, as diversity is better, and most of the artifacts are gone because with avg (50/50), if an artifact is present at one moment in one model, but not the other, in the ensemble the artifact will be heard twice time lower. Ensembling between Roformers create too much redundancy, even if the datasets between for example HyperACE and deux are different, with different loss etc but still.
I think it's a better ensemble than the others.” - dca
-> Mel deux + inst_gaboxFlowersV10 (Avg)
The #3 best now.
- Karaoke ensemble: BS Karaoke anvuew + Mel deux becruily (Max FFT)
The best for now
- bigshifts were added by jarredou in a newer jarredou’s inference Colab version here. It's based on a newer refactored MSST code. It might be potentially more future-proof, because more and more model trainers use this newer MSST code.
Be aware that this Colab currently falls behind in terms of models and model types support, with the makidanyee's Colab being supported more actively now.
Also the model payload with model links were moved to a separate json file.
The custom Colab (for comfy model links pasting) will have added bigshifts and this refactored code later.
- Essid found some new models we probably didn't know before, in a noblebarkrr public HF repo:
https://huggingface.co/noblebarkrr/mvsepless_resources/tree/main
It includes the archive of probably all existing public models released since 2021.
All these models can be used in:
https://huggingface.co/spaces/noblebarkrr/mvsepless_zero_gpu_long_quota
https://huggingface.co/spaces/noblebarkrr/mvsepless_zero_gpu
(you can only test one model in this one, it reaches quota fast with a timeout, but it can be bypassed with VPN).
It also has a Colab version:
(not sure how the models list is up to date here)
The issue is, it's all in Russian, and Google Translate doesn't work on the whole site unless you highlight fragments of the text and use context menu to translate or for copy-paste.
Also, judging by the GUI, it might still use dim_t instead of chunk_size, unless it reads it from the yaml if it's not forced.
dim_t to chunk_size conversion (every model has different optimal chunk_size, sometimes already in the config)
You can even use slower CPU HF versions in the meantime (click restart of necessary):
(the one above less models, and probably older ones)
The list of existing models receiving the first HF/Colab implementations above so far:
a) Single 53 stems MVSep models (and also one big single one - it might potentially work on Zero GPUs on HF, and single models everywhere). Exclusive DL location (only 77MB per model).
b) Sucial De-breath VR models v1/v2
c) PoPE a.k.a Polar Former ZFTurbo vocal model
(it could slip behind a radar for you)
"ngl that model sounds really good but would need some improvements i guess" - mesk
d) Aname's PoPE (fine-tune of the above)
("the 2 stem is decent, I tried the 4 stem and it was meh" - 5b
e) Single xlancelabs fine-tunes of the SW, with more instruments (percussion, synth, orchestra)
Potentially new model discoveries from the repo, which weren't known before
("mbr" in the model name = Mel Band Roformer)
bs_karaoke_3stem_giantailab (for use with this bs_roformer.py file)
demucs3_saxophone
mbr_guitar_chencfd
mbr_percussion_yolkispaliks
mbr_lead_rhythm_guitar_listra92
mbr_wsa (vocal)
bs_fnf_mrdense67 (vocal)
mbr_denoise_children_phaedrus33
Various 4 stems models.
Maybe some jazzpear’s are new.
Single DL links:
https://discord.com/channels/708579735583588363/708580573697933382/1497559415027535914
- frazer wrote mask estimator for the 53 stems MVSep model | DL (tech. explanation added)
“That thing should reduce it down by 50x memory.
So IDK how much memory the [53] separate mask estimators take up but this reduces them down to 1 (...) that tweak will make it smaller than the SW model”
- (jarredou) “Here's an edited bs_roformer.py file to run Mvsep mega 53 stems model using quite less VRAM (by processing each stem mask one by one instead of all at once), without affecting output quality.
You can tweak in config file:
model:
inference_stem_batch_size: 1
inference_offload_mask_estimators: true
“
https://drive.google.com/drive/folders/1x--32jXPZichCGeqyX0ugm0uebbfVj2V
The memory drop on 1080 Ti is from 11GB to over 5GB.
But if you still have memory issues:
"My edited bsrofo version is side-loading the unused masks to system memory instead of keeping everything on GPU memory. Maybe using inference_stem_batch_size: 2 or 4 could help balancing the stuff between RAM and VRAM [instead of 1; if you had such issue]” - jarredou
The Colab errors out on full songs with it, and without it, even with 88200 chunks and batch_size 1-4. Probably 16GB of RAM is necessary for it to work (free Colab has 12GB iirc)
Skipping stems during the process to decrease memory usage: “with a bit of tweaking, you can skip some of the stems mask estimators, so they will not be processed”
- All gillian’s models were moved to another repository. All the links are in Phantom Center dedicated section (the same for drums and bowed strings models)
- (gilliaan) Hi yall. I trained a bowed strings source separation model:
https://huggingface.co/gilliaaan/BS-Roformer-BowedStrings-Duality/tree/main (dead) | Colab
[Doesn't work in UVR correctly. It outputs only one stem and inverts it. Use MSST instead.]
it's a 2-stem model (strings vs. other) trained almost exclusively on commercial music
Pop, rock, metal, folk, funk, disco, Arabic music, Asian music, country... basically everything except orchestral (no symphony basically). That was a deliberate choice. There are already models that lean classical, and I wanted something that actually handles strings in other types of songs.
Trained on 4 cases:
Sustained, Pizzicato, Tremolo and Plucked strings.
Trained on Violin, Viola, Cello for all 4 cases, except for Bass I only trained on Sustained (because plucked bass for jazz might be confusing).
Best results at overlap 6. Overlap 1 is also worth trying in some cases! it captures more of the strings, but at the cost of some bleeding/cutting, so it depends on the track.
My notes: it definitely does the job. It's a dual model trained with a bias toward fullness on both stems, so neither side sounds too hollow.
Feedbacks are welcomed, especially on edge cases!!! (...)
Strings datasets are extremely hard to find, so I had to make it myself.
I found like 45 songs in the moisesdb, the rest is all me. (...)
your welcome! my secret is making 1 mins tracks from scratch and adding strings using kontakt sound banks! also got some stems from friends in the industry, and ofc moisesdb (...) [done] semi manually on FL Studio, but also there's a vst called tekno where you can randomize all the drums, i just found a way to change the key, eq, effect (wet, dry), did 400 songs like that by randomizing a lot of things” - gilliaan
“For Reaper users, there's this script to randomize plugins parameters with midi notes triggers” More - jarredou
“This strings model is great! I tried it on 2 tracks and it's pulling a lot of detail” - rage313_
- BS-Roformer “MVSep Mega 53 Stems” model was released
“First version (...) List of stems:
['accordion', 'acoustic-guitar', 'back-vocal', 'banjo', 'bass', 'bassoon', 'bells', 'bowed_strings', 'brass', 'cello', 'clarinet', 'congas', 'digital-piano', 'dobro', 'double-bass', 'drums', 'electric-guitar', 'flute', 'french-horn', 'glockenspiel', 'guitar', 'harmonica', 'harp', 'harpsichord', 'hh', 'keys', 'kick', 'lead-vocal', 'mandolin', 'marimba', 'oboe', 'organ', 'percussion', 'piano', 'saxophone', 'sitar', 'snare', 'strings', 'synth', 'tambourine', 'timpani', 'toms', 'triangle', 'trombone', 'trumpet', 'tuba', 'ukulele', 'viola', 'violin', 'vocal', 'wind', 'wind-chimes', 'woodwind']” - ZFTurbo
It has 4 stem drumsep ('hh', 'kick', 'toms', 'snare'), so a new, unavailable model on MVSep
It has vocal and lead vocal models
It has electric, acoustic and guitar models
Lacks Chello, Backpipes, Braam, Xylophone, Choir (choir/other), SATB Choir (soprano, alto, tenor, and bass) available as separate models on MVSep.
Be aware that stems won't invert with the mixture - it means that the stems have overlapping sounds or instruments. E.g. Vocals have all vocals, while BVs just BVs as intended.
At least free Colab with 15GB VRAM currently gives ^C error (from what jesse wrote), unless you use a 30-second song fragment with chunk_size 352800.
“Try with forcing batch_size=1.
The Colab is forcing batch_size=2, it was a workaround for a click issue, but not needed anymore since the issue was fixed at source a while back”
In the separation cell, edit:
data['inference']['batch_size'] = 2
to 1
- jarredou
Possibly you might force it with the "use_modelconf" option. Consider using 88200 or 112455, so, a much lower than in the OG config from the config (441000), but the results can be unpredictable.
“Managed to run inference on 4060 (8GB), but had to lower the chunk size to 88200 (...) still working fine with chunk size 112455” - neoculture
“I ran a 4 min song w default chunk size overlap 2 took 50 minutes 4070 super (12GB)” - 5b
It was rather using MSST and batch_size 1 from the yaml, (it's probably default in MSST).
The model also works in UVR, but because it probably forces batch_size 2, it might require more memory than MSST (older version of the MSST inference the UVR is based on had a bug where batch_size 1 caused skipping between segments, so it’s probably forced).
ZFTurbo announced he will further fine-tune the model to get the best possible metrics before uploading it on MVSep.
“most [stems] came out really muddy just bc of the song I used. I would test more if it didn't take so long to run the model fully” - 5b
“model definitely sounds rough in many ways, and I wouldn't use it for anything critical. However, it has its strong spots and imo has a lot of potential.” - musicalman
- A new vocal model was released:
Anvuew BS-Roformer FT1 12.55 | yaml
Vocal bleedless 35.17, fullness 19.88, SDR 11.54.
Works in UVR (at least on the non-RTX 5000 patch - the previous version of the model didn't work on RTX 5000 patch).
The metrics are better than the Unwa BS-Roformer voc HyperAce v2 and Gabox Mel-Roformer voc_fv7, and probably vs the last anvuew’s 12.45 model at least SDR-wise (if it was trained on the same validation dataset).
“I compared the model with big beta 7. both have pretty similar outputs. but for some reason, 12.55 has a bit of distortion, while big beta 7 sounds fine.
for the lead vocals, they’re quite similar, but for the backing vocals, 12.55 sounds slightly buried (you probably wouldn’t notice unless you listen really closely)” - neoculture
- gilliaan's drums model added to testing version of the uvronline, and it also works with demudder (“refreshing and re-applying the model to the same file made it work anyway”) - click for access
- If you still get weird results with fixed phase fixer Colab despite using FLOAT option instead of FLAC, A6 created manual phase fixer Colab based on the OG Michael’s one which Google broke by the changes in the env, which is deprived of the issue, but requires already separated files to work. Later Squid/squiddagod added optional FLOAT support.
It should sound similar to local scripts now.
- Aname released a new BS-PolarFormer duality model based on the new arch enhancement a.k.a. PoPE, added to the last versions of MSST (how to update it is explained in the bigshifts news right below - because you need the latest bs_roformer.py, later “open MSST folder, run CMD from it, and just ‘pip install pope-pytorch’” - lilihamer, 5b). It uses ZFTurbo’s model as a base, which already had a tad better SDR than the Mel Kim’s model. According to the Aname’s model card, it might excel well in vocal reverbs and in fullness of both stems. There’s also 4 stem variant.
https://huggingface.co/Aname-Tommy/BS-Roformers
Since the release, the model access has been restricted to only people providing contact information (although isling managed to download the model before, while Aname later left the server).
“Outputs were noisy and had higher levels of weird artifacts than the models I prefer. But some insts in particular were nice to listen to, so IDK. I believe these models have potential, but were undercooked/not optimally trained.” - Musicalman
“it wasn’t bad at all, the fullness was surprisingly good/clean (but can’t confirm if it’s the model itself or PoPE)” - becruily
“Seems to leave a metallic-ish reverb to vocals, almost like a metal plate being hit. Think it's bleed from the instrumental” - Pipedream
- Phantom Center MDX23C model by gilliaan is now fixed and can be used in this forked Colab
- Due to 1.02M character limit in the Google Documents, I was forced to move the whole training section to a separate document here (we’re only below 90 pages from reaching the limit with current formatting again, it’s 625 now, and in fact 1164326 characters, counting without spaces)
- Google added new GPU to paid the Google Colab plan called G4:
“Added support for Nvidia RTX Pro 6000 Blackwell Server Edition GPUs on G4 machines. With a peak rating of 960 BF16 TFLOPs (~50% more than the A100-80G) and 96 GB of VRAM (20% more than the A100-80G), it’s the most efficient high-performance GPU we’ve ever offered. Select "G4 GPU" in the runtime type menu.”
In late January, “H100 is being rolled out for more users” maybe paid ones, IDK.
- jarredou’s bigshifts trick from MDX23 Colab was merged with MSST lately. To use it, update your MSST (not necessarily WebUI) with:
!rm -rf /content/Music-Source-Separation-Training
!git clone https://github.com/ZFTurbo/Music-Source-Separation-Training
And you might have to reinstall the main branch's requirement.txt (not sure).
Then it can be set with --bigshifts 3 CLI argument.
Some people liked to increase BigShifts to 20 or even 30 with all other default settings (some songs might be less muddy that way), but default 3 is already a balanced value, but exceeding 5 or 7 may not give a noticeable difference, while increasing separation time severely.
BigShifts are “more efficient than high overlap value to improve the separation a bit, by doing a bunch of passes on the input audio (with some segments shifted in time for each pass, then restored, all passes are finally averaged).”
More about how BigShifts work can be found here.
- Actually the single phantom center MDX23C model below turned out to be better than the ensemble suggested below earlier. Recommended settings:
Overlap 16 with batch_size 16 and TTA
“it sounds extremely good” - gilliaan. Metrics.
UVR doesn't support batch_size and TTA, only MSST.
“You can run this with a 6GB [NVIDIA] VRAM card. I processed the 27 mins test set on a 6GB card using these parameters, only took me 10 mins” - gilliaan
In case of any errors with those models in UVR, use these modified configs (mirror) - thx vxsz.
“I did this for mixing/mastering purposes, but you could use it to extract center/wide and then process both files into a fullness instrumental model (I recommend resurrection inst by unwa), the instrumental sounds less noisy and cleaner (opinions based)
also only processing the center track with instrumental model is gonna give you a clean karaoke, but it's only for certain songs” - gilliaan
“if you use hyperace to process center wide output, there’ll be a bit of vocal reverb left in the instrumental. deux doesn’t do that, but it sounds a bit muddy”
“also, holy s*** the results are clean.
Only with Deux tho, I wish I could use hyperace in UVR tho” - vxsz
- (gilliaan) “Just dropped my new MDX23C Phantom Center model! | MK forked Colab (fixed again)
(dead) https://huggingface.co/gilliaaan/MDX23C-Phantom-Center-Wide-Extractor/tree/main
But honestly forget the single model, what matters is the ensemble. Lemme explain.
[For all of the models to appear in the Ensemble menu in UVR, you need to edit targets in the yaml to vocals/other]
Tested BSv1 + SCNet Dual XL + MDX23C [avg ensemble: metrics | examples] together and there are literally zero artifacts. invert both stems against each other and there's no muddiness whatsoever. Separation is at least 98% in my book, that's as close to perfect as I've heard from any of these. (...)
It exports what feels like it is in the center.
so all the effects like reverb, delay, wide synths, panned drums, they all go to the wide stem (...)
Every single model has its own artifact. but when you throw all three together they cancel each other out completely. Zero artifacts. None.
I consider using the solo models a waste of time so don't bother, just go for the ensemble, trust me, perceptual quality is through the roof.
Side note: this will probably be one of the best pre-processing tools for vocal extraction too.” - gilliaan
- CenterWide Phantom Extraction Dual Model trained with SCNet XL IHF was released by gilliaan
Model | Metrics | Leaderboard | MK Colab
“Sounds fuller, but also less precise than the BS Roformer version, also scores higher SDR than the BS version.
It's clear to me that an ensemble of these 2 models would give the best results (metrics)” - Gilliaan
In order to make this model work in UVR, you need to set target instruments to null in the yaml.
"I like doing that because when you average those 2 it removes bias from both" - dynamic64
"they both have different types of artefacting, averaging them gets you even closer to the "real" thing" - Gilliaan
“while [the SCNet] addresses the inconsistencies between the 2 previous models [Mel and BS], its consistency does not make neither of the 2 models completely obsolete yet. While I'm still in the early stages of testing with this model, I have already encountered one instance where the Mel-Band still preserves the center a little better. Despite that, this is still now my go to model that I will be pointing people towards when asking about Center Channel Extraciton.” - Vinctekan
I recommend using overlap 16 because SCNet is quite noisy.” - Gilliaan
“Phantom center single stem extractor trained on a custom synthetic dataset of 2000+ 4 mins songs. This model is trained to isolate only the correlated center content, so hard-panned signals stay in the wide output and don't leak into the center. This model was trained on center and wide, dual model. And is a fullness model.”
- dereverb_bs_roformer by anvuew
and BS-Roformer Phantom Center model by gilliaan added on MVSep
- gilliaan released drums model
BS-Roformer-DrumsOther-Duality
Available on testing version of uvronline.
It is a fine tune of the SW.
Incompatible with UVR - doesn't output both of the stems (only the drums stem without the other independent stem and not inversion). Use MSST instead.
The model is focused on fullness of both stems, not so much on accuracy (there are better models for it). E.g. mainly for a drummer wanting to play to the other stem without bleeding, and with also fullness preserved.
It can also be used as a pre-processor for vocal separation.
Or an ensemble for other drums models to enhance their fullness (e.g. SCNet or the SW) - more. Metrics are only available for the worse drums stem, because the other stem has vocals+bass+other.
“it really does not enjoy the sample rate trick” - Pipedream
- NVIDIA released a new upscaler called RE-USE - it's a speech enhancer.
It works on HF if you upload your audio as chopped 30s WAVs:
https://huggingface.co/spaces/nvidia/RE-USE
Compatibility:
NVIDIA Ampere (A100)
Preferred Operating System(s):
Linux
https://huggingface.co/nvidia/RE-USE#environment-setup
"Damn this is impressive! At least on my own random test recordings.” - Musicalman
“sounded really good” - mogwai_63846
- Looks like official implementations of BS/Mel Roformers were released by the OG authors from Bytedance team who released their papers for the archs back in 2023, which got later implemented by lucidrains publicly, and later reimplemented in ZFTurbo MSST repo.
https://github.com/asriverwang/BS-RoFormer/
- lucidrains account has been restored on GitHub with all the repositories:
- (fixed) MVSEP is currently undergoing technical issues with queues reaching 2K jobs expected to be solved in a few hours (16:30 UTC +1 25.03).
Update: it got better at least with prioritized users (logged in/paid).
Now probably everything got back to normal
- gilliaan released BS-Roformer model (yaml) for Phantom Center extraction (single stem) | MK Colab (fixed)
In order to make this model work in UVR, you need to delete the whole: mlp_expansion_facfor, skip_connection, use_torch_checkpoint lines from the yaml (or use modified yaml).
"This model was only trained on center. (...)
it's a lot better than the 2 previous Melband Roformers I've released. BS Roformer understood the task better. The residual stem (wide) sounds good as well, even though it's more on the bleedless side. For reference we went from 10.2 SDR up to 16.9 SDR. This is a big jump. (...) You can compare metrics here.
Ensemble Beta V2 + BSv1 (Average - Average (Avg)) sounds better" - gilliaan
Now also Phantom Center leaderboard for evaluation was added on MVSep.
- (MVSep) A new Special Effects model, MVSep FX (fx, other), has been added.
Demo: https://mvsep.com/result/20260318224517-f0bb276157-mixture.wav
Notes:
- It was trained on FX stems from songs.
- We didn't use DnR datasets for training, so the results should be quite different. It is mostly optimized for music rather than movies, etc.
- The model can also remove vinyl scratches (I remember someone asked for this).” - ZFTurbo
“Finally a model which can remove the cartoon sounds from the Rocky and Bullwinkle score! (...) works really well for removing sfx from old movies” - fal_2067
“works pretty well, I’m a fan” - dynamic64
Removes effects from music videos
FX is the effects, other is vocals and the music. - wancitte
- New audio upscaler’s training and inference code with models was released: UniverSR (by woongzip1)
Paper | Demo | GH | Colab #2 by Sir Joseph with chunking.
"Vocoder-free audio super-resolution model that upsamples 8/12/16/24 kHz → 48 kHz audio using flow matching in the complex STFT domain. Trained on speech, music, and sound effects." - Marko Ravich
There are two models: general and speech.
It seems like it's deprived of some nasty artefacts AudioSR and FlashSR can produce.
Sadly no demos on some dense mix.
“Chunking does not happen automatically, fails to run on an RTX 4090, chunking manually however worked and it works well, except it only takes in mono” - pipedream420
- And also FastWave by NikaIt
“Got it running, but not great results” - pipedream420
- gilliaan released a new dual phantom center model gilliaan_MonoStereo_Dual_Beta2 | yaml | MK Colab. It has better metrics than the beta 1. Newer models might be trained on other archs now which have better understanding of stereo mapping and spatialisation.
The model is not good concerning the center stem" - gilliaan
"Isolates the sides better than every other center extraction model I've tried" - pezz23
“it's f**** amazing” - knock2one
- The doc previously lacked SCNet difference/sides model by Dry Paint Dealer Undr and it was added in the relevant section.
https://drive.google.com/drive/folders/1ZSUw6ZuhJusv7HE5eMa-MORKA0XbSEht?usp=sharing
- gilliaan released leaderboard for phantom center with explanation on important metrics - click
- gilliaan (heauxdontlast) released:
gilliaan_MonoStereo_Dual_Beta1
(dead) https://huggingface.co/gilliaaan/Mel-Band-Roformer-MonoStereo-Duality
“Phantom center dual extractor trained on a custom synthetic dataset of 1000+ 4 mins songs. This model is trained to isolate only the correlated center content, so hard-panned signals stay in the side output and don't leak into the center.
Bleedless/Fullness are on the same level.
On my validation set it scores higher SDR than what's currently out there. Still training, this is a demo beta.
Feedbacks are welcomed
mid — SDR 10.41, aura_mrstft 16.79, bleedless 35.29
side — SDR 10.39, aura_mrstft 13.68, bleedless 27.95
avg — SDR 10.40, aura_mrstft 15.24, bleedless 31.62
I recommend overlap 8 (or more) for best quality.”
- anvuew released a new “vocal separation model specialized for magnitude spectrum accuracy.” called BS_RoFormer_mag | Colab
Bleedless 32.17, fullness 22.15, SDR 11.09'
Worked in UVR, but compatibility with RTX 5000 patch is not guaranteed.
“managed to pull out harmonies I knew were missing from my test track as well as not extracting some vocal sampled perc that had been in other models previously” - cristouk
“pretty noisy to me compared to fv7beta2
but I only tried it on two songs. def was keeping the harmonies at a more proper volume and less muddy, though. but I think it has the same problem as deux.. a lot of noise sometimes during quieter/silent parts” - rainboomdash
For comparison: Becruily “deux” for vocals:
voc. bleedless: 28.30, fullness: 23.25, SDR: 11.37’
All the bleedless, fullness, SDR metrics of the “mag” model are a bit better than Revive3e.
- drypaintdealerundr shared two previously unreleased Phantom Center/Similarity extraction models in BS-Roformer architecture, from some training sessions from September or October. DL
- unwa released BS-EXP-SiameseRoformer vocal model | Colab
“seems pretty high bleedless... muddying/breaking up the voice a bit” - rainboomdash
Incompatible with UVR. You need to replace that attached py file in your MSST installation.
Some training details in the repo.
- It looks like the model below doesn't work with the UVR's RTX 5000 patch. First, giving an inference error, then with torchdynamo when you use a fixed config probably already linked below, but I'll link here another working one by analogspiderweb just in case: DL.
Like always in case of torchdynamo issues, we recommend MSST instead for faster separation on RTX 5000 GPUs, instead of using UVR’s non-RTX 5000 patch, as you will have near zero GPU acceleration in such case on those GPUs.
- New Mel-Roformer small_karaoke_gaboxauf by Gabox and Aufr33 was released | yaml | Colab (fixed)
Compatible with UVR (but not RTX 5000 patch). Metrics.
“already noticing that it sounds much more fuller than other karaoke models i use [anvuew bs roformer, bs karaoke gabox, and becruily karaoke] and clears out bg vocal much more accurately” - baptizedinfear
“great! got to hear vocals I could not hear before.” - makesomenoiseyuh
“it's not great with duets btw” - Gabox
Lowering chunk_size to 352800 makes it muddier, but less noise, “maybe in between would be nice” - rainboomdash
- (uvronline.app) Model dereverb_bs_roformer_anvuew_sdr_22.5050 added
- Gabox released experimental last_bs_roformer. It “removes lead vocals focusing in bv”
“vocals has backing vocals and instrumental, instrumental has just lead vocals”
“some lead vocal leaked into instrumental so i'm cleaning the dataset for the fifth time” - Gabox
- (MVSEP) “New wind model MVSep Bagpipes (bagpipes, other) has been added.
Description: https://mvsep.com/algorithms/113
Demo: https://mvsep.com/result/20260221123255-f0bb276157-mixture.wav
New model MVSep Braam (braam, other) has been added. I put it in the "Effect" section.
Description: https://mvsep.com/algorithms/115
Demo: https://mvsep.com/result/20260221124005-f0bb276157-mixture.wav” - ZFTurbo
Braam is a loud, potentially distorted, low-sounding effect, most commonly known from the movie Inception (the composer used piano with 10x brass for it).
Bagpipes - “not too bad, you can still faintly hear the pipes though.” - fal_2067
- splifft now has support for most community models (BS/Mel-Roformers and MDX23 supported). More.
- unwa released Big Beta 7 Mel-Roformer vocal model | Mkd Colab
Bleedless 38.77, fullness 16.20, SDR 11.20.
“SDR is lower than the BS model, but personally I prefer this one. (...) Although not reflected in the metrics, noise has been reduced in sections without vocals.” - unwa
More bleedless model than 6X:
Bleedless: 35.16, fullness: 17.77, SDR: 11.12
- luidrains’ GitHub account has been suspended without a warning, and along with it, all the repos are gone. They have been moved here and here.
Dear Microsoft,
Have you fucking lost your mind?
- Anvuew released a new BS-Roformer 22.5050 de-reverb model
“Very interesting model, sounds like the Mel-Roformer one but even more aggressive, it's good”, it works on “reverb effect on vocals” at least - isling
“sounds cleaner than the `mono sdr 20.4029` one” - rainboomdash
“The dataset is the same as mel, all reverb is generated by VST plugins and includes waves IR1 presets. so it depends on whether you consider IR1 to be room reverb” - anvuew
- It turns out someone found a way to use GPU Conversion with Wine by using Bottles to use newer UVR Windows codebase with better support of the newest Roformers than using outdated native Linux code which is also convoluted to install.
Although it might be slower than MSST due to additional translation layer, so consider MSST instead, but it was still relatively fast on RTX 5060 using Demucs, although it was slow on 4090 in at least Roformers.
We received an independent confirmation from other users that it worked indeed:
"My way to easily run UVR on Linux:
I just downloaded Bottles
https://usebottles.com/ (which uses WINE) and used the provided .exe file from the repository's releases. I created a new bottle with a gaming profile (to utilize GPU) and moved the exe file into "drive_c" (otherwise it won't work), than just ran through the installer and it worked like a charm!" (src)
BTW. “Soda is bottles internal WINE fork”
- (it was called off for now) “Discord will require a face scan or ID for full access next month (...) Users who aren’t verified as adults will not be able to access age-restricted servers and channels, won’t be able to speak in Discord’s livestream-like “stage” channels, and will see content filters for any content Discord detects as graphic or sensitive. (...) However, some users may not have to go through either form of age verification. Discord is also rolling out an age inference model that analyzes metadata like the types of games a user plays, their activity on Discord, and behavioral signals like signs of working hours or the amount of time they spend on Discord.” - theverge.com
- In February 2026 it started to happen that Chrome on Windows was closing after opening this document. It helps to reopen it a few times when it closes. Before I also tried it along with uninstalling the offline GDoc extension (restart the browser afterwards) and later after going to: chrome://settings/content/all?sort=data-stored > Google > docs.google.com (I just did all of these), but the issue was recurring. Also, when you just disable the extension it was able to re-enable itself, but the issue might be bound more to another browser extension incompatibility too. Actually the same things even started after opening YT, and multiple reopenings of the browser helped in both cases ultimately. Few days later the same happened on Gmail.
You can also just download this document as PDF, or docx to preserve up-to-date table of content (although it will list all the headings used in the document in the docx, so it will be messy).
- “Github is facing issues since some hours
https://www.githubstatus.com/”. It results in at least some problems with cloning and errors 500 & 128 in Colabs randomly.
- MVSEP introduced a limit of 50 separations per day for free users (more). There were cases where certain free users were clogging the queue substantially.
“There won't be reset [time], it will just count all separations in the last 24 hours.”
- Jasper/jasperkt (a.k.a. jazzpear) made a an auxiliary/helper model for DNRv3 and some vocal model working for speech, serving for extracting foreground and background SFXs (it wasn't trained on voice and music; more info)
- Mel-Roformer for background music from a movie: model | yaml (stems: voice with SFX/BGM without singing voices), you might want to use some vocal model as a preprocessor here too.
Compatibility with UVR not guaranteed, consider using MSST in case of any issues.
- Gabox released voc_fv7 Mel-Roformer | yaml | makidanyee’s Colab
Bleedless: 33.85, fullness: 17.99, SDR: 11.16’
Beta 6X metrics for comparison
Bleedless: 35.16, fullness: 17.77, SDR: 11.12
HyperACE voc V2 for comparison:
Bleedless 34.08, fullness 19.10, SDR: 11.40
Sometimes results sound not as noisy as BigBeta5e (but it depends on a a song) and voc HyperACE (which might sound less muddy, but noisier), but it catches harmonies much better than BigBeta6X:
voc_fv7 is "a little noisy, quiet, but at least it's capturing [the harmonies] (...) a little more muddy than big beta 6x, big beta 6x is already muddy... so it's just.. meh...
definitely captures harmonies a LOT better than big beta 6x.
elsewhere from those, it's pretty mixed... sometimes big beta 6x was better, sometimes fv7. (...)
I tested a few more songs, a couple fv7 was fuller, but most songs it's more bleedless compared to big beta 6x.
As with all models, it varies drastically song to song on how full it is compared to other models
fv7 was also cutting the reverb really aggressively on one song compared to big beta 6x. I'd personally like if it was a little fuller, like in-between fv7beta2 and fv7.
Bleedless models usually have significant issues with backing vocals, they make the vocal way quieter than it's supposed to be to suppress noise (...) I am noticing fv7 frequently creates bursts of noise, but it's *usually* good.. it's doing that to capture something difficult
while big beta 6x doesn't do that as much
overall, fv7 does seem better than big beta 6x.. it might have a *slightly* higher chance of capturing instruments.. but far less than fv7beta did.. I would need to do more testing to really know. and it's pretty comparable to big beta 6x in noise level, so I'd say a good upgrade overall.. big beta 6x is usually good, but it falls apart sometimes, especially some backing vocals it just completely falls apart
voc hyperace is noisy in comparison (...) bleedless score is about the same between fv7 and voc hyperace.. but the fullness of voc hyperace is about 1 higher in metrics
which more would line up with how much noise I'm hearing.. it's def perceptually noisier (...) “def picks up more instruments than big beta 6x :/ but it picks up harmonies waaaay better... meh (...) it's picking up a lot of partial intact stuff from the instrumental
but the overall noise level is relatively low" - rainboomdash
- “I released example of Telegram bot: https://github.com/ZFTurbo/MVSep-Telegram-Bot-Example
It was mostly prepared by a student, so be [forgiving] please.
Also I put bot on the permanent run here: https://t.me/MVSep_com_bot
Report here what changes you want if any.” - ZFTurbo
“it's always 120bpm. if you want it to play at the speed of the original file, set your project bpm to 120” - isling
- Unwa released BS-Roformer-Large-Inst a.k.a. bs_large_v2_inst (there was no v1) | makidanyee’s Colab
Fullness: 32.06, bleedless: 43.95, SDR: 17.61
It uses a custom bs_roformer.py file attached - just replace it in the MSST installation (UVR is incompatible).
Training details:
"Instead of increasing the depth to 16, I added a four-layer TransformerBlock to the MaskEstimator." 238MB weight.
“sometimes was good and other times it completely missed the separations”, more noise than flowersv10 model - mesk (spectrogram example with lacking fragments missed by the model). “It leaves residue in some places” - fabio06844
“Just looking at metrics, fullness is ranked about the same as (...) FNO” - rainboomdash
- Gabox Flowers FV10 model added on uvronline.app while using following links:
https://uvronline.app/ai?discordtest
- link for free accounts
https://uvronline.app/ai?hp&test
- for premium ones
- Phase fixer Colab seems to be fixed here
- xlancelab who won the last raw stems restoration competition released pth models based on the BS-Roformer SW model along with dedicated code for chained separation (separation>dereverb>denoise).
In addition to the SW stems, they needed to train percussion and drums stems separately, and also orchestra and synthesizer. 9 stems (besides attached Aufr33 27.99 denoise and anvuew 19.17 dereverb models converted to single pth files too).
“Their training used L1 loss combined with multi-resolution STFT loss on MoisesDB and a manually cleaned version of the RawStems dataset (...) single H200 GPU for large scale training and a single 4090 GPU for small scale training [was used]” more info (thx jarrdou, becruily, anvuew)
“i tested vocals with one song, turns out it's supposed to give you dry lead vocals (didn't sound very good anyway)” - becruily
- (MVSEP) New model MVSep SATB Choir (soprano, alt, tenor, bass) has been added.
Description: https://mvsep.com/algorithms/104
Demo 1 vocals: https://mvsep.com/result/20260108154639-f0bb276157-mixture.wav
Demo 2 vocals: https://mvsep.com/result/20260108155023-f0bb276157-mixture.wav
Demo strings: https://mvsep.com/result/20260108154828-f0bb276157-mixture.wav
Very big thanks to Dry Paint Dealer Und for helping me to create this model.
P.S. Model works not only with vocals but with strings too" - ZFTurbo
It works also for instruments, piano layers "able to split chords into each individual layer (...) and such" - pitbulldale305
It's a BS-Roformer, and not a fine-tune of the previous model “Metrics [are] much better, so I'm not sure if it's reasonable to use the old model.”
BS Roformer 11.89 is currently used for the "Extract vocals first" option.
"it could be worth exploring running karaoke model first and then SATB, since SATB might combine the lead with some bvs (unless this is expected)" - becruily
"Im sure if you used a combination of a karaoke model to get a clean lead/backing stem, then put the backing through this model you would be so much better off than we have been" - dynamic64
"Works pretty well on both "inverted acapellas or official acapellas, official acapellas definitely would have cleaned BV".
Despite vocals, considering it was trained on DPDU’s dataset, it was also trained on MIDI strings, plus ZF usually never starts training from scratch from what he once said, but rather retrain on a model which already knows some patterns. It's also good for dubstep instrumental and "I think this is really useful for audio to midi" and sheet music transcriptions - pitbulldale305, dynamic64
Q: How close are we to actual harmony separation?
A: We're there. With a combo of karaoke and this SATB model. It's probably more manual than you'd like it to be though - dynamic64
Q: What's the difference then between Choir and Choir SATB models, ie: why would anyone choose choir if you can use Choir SATB? Is the accuracy numbers higher on choir?
A: Choir separates out only the choir, SATB splits the whole track into one of those stems.
If you want a choir out of a song separated, you need the choir model. - dynamic64
- “New model MVSep Choir (choir, other) has been added.
Demo: https://mvsep.com/result/20260107221631-f0bb276157-mixture.wav” - ZFTurbo
(works for e.g. choirs buried under main vocal in e.g. pop music)
- Gabox released inst_gaboxFlowersV10 (yaml) Mel-Roformer | makidanyee Colab
Inst. bleedless 36.68, fullness 37.12, SDR 16.95’
All the metrics are better than Inst_FV8b
(Inst. bleedless: 36.90, fullness: 35.05, SDR 16.59)
Consider phase fixing with Becruily voc model instrumental result - prodbyluke/ghostofmiami_
Incompatible with RTX 5000 UVR patch (users will encounter “ModuleNotFoundError: "No module named 'torch._dynamo.polyfills.fx'")
Chunk size for the flowers model in the Colab should be 352800.
“it was going to be "inst_gabox_deuxnt”
it is deux, just with "other" target (...)
some pop/rock n'roll added” to the dataset - Gabox.
The model is planned for training further in future.
Be aware that the yaml is new, and you cannot use any old one for that model.
“a huge improvement over what Gabox has accustomed us to!!!! I got amazing results. It truly achieves a stability that previous models lacked; the clarity is much greater in these.” - billieoconnell.
“sounds pretty good tbh (at least for my songs)” - neoculture
More sensitive to keep some tiny sounds over gaboxflowersv10 - sakkuhantano
- (MVSEP) “Three new models have been added.
In BS Roformer (vocals, instrumental):
1) unwa BS Roformer HyperACE v2 instrum (SDR instrum: 17.40)
2) unwa BS Roformer HyperACE v2 vocals (SDR vocals: 11.39)
In MelBand Roformer (vocals, instrumental):
1) becruily deux (SDR vocals: 11.35, SDR instrum: 17.66)”
When you search for “deux”, it somehow doesn't appear, you need to go to the MelBand section and find it near the end of the models list.
- jarredou evaluated different chunk_size settings and how they affect different evaluation parameters on the deux model. Vocal stem was tested, but he said the inst stem has similar curves.
Best SDR: 705600 (possibly the least crossbleeding)
Best fullness: 1102500 (requires lots of VRAM)
Best bleedless: 441000 (still not low enough for 4GB VRAM on AMD and Intel GPUs in UVR)
Default in the yaml: 573300 (becruily’s recommendation till at least now)
“I also did the same test as jarredou, before releasing and inded 705600 gave the highest SDR but it also added a bit more noise (for instrumentals it might be good, for vocals not so much)
so I found 573300 to be a good middle ground between fullness/bleedness/SDR” - becruily
“882000 - seems the maximum viable one if you aim for best fullness before such diminishing returns” - makidanyee
“I do recommend trying higher chunk sizes for instrumentals with deux
I find 661500 works for a lot of songs, 749700 for a good amount of others
higher=more fullness, less bleedless” - rainboomdash
- (uvronline) “I updated the pre/post processing for the Male/Female and Mel-RoFormer Lead/Back models. Now the Deux model is used, so the instrumentals will be less muddy.” - Aufr33
- (uvronline) The new Deux model by becruily added. It uses vocals stem and instrumental is inversion, not dedicated instrumental stem, and that's how becruily recommends to use it. It was fixed, and was actually a mistake.
“Another improvement: since the deux model creates two stems, phase correction is now applied immediately. You don't need the correct phase button “- Aufr33. At least before, phase fixer was only for premium.
“First, phase inversion is performed, resulting in a muddy instrumental. But the result is only used for phase correction.The second stem, which is available for download/listening, is more full.”
- (MVSEP) New model MVSep Celesta (celesta, other) has been added.
Demo: https://mvsep.com/result/20251230133507-f0bb276157-mixture.wav
- Thanks to Ari/arxynr, we rewritten phase fixer section, so it no longer has the mistakes preventing you from getting correct results as intended.
- Becruily released dual Mel-Roformer model called “deux” for vocal:
Voc. bleedless: 28.30, fullness: 23.25, SDR: 11.37
and instrumental separation (two stem model):
Inst. bleedless 41.36, fullness 34.25, SDR 17.55
Compatible with UVR Roformer patch, including the RTX 5000 one | Colab
“Unfortunately it won't null to the mixture perfectly, this applies to any multi stem model” - becruily
“[for instrumental] the best fullness model and does not even need phase fix. SOTA even (...) my fave fullness model” - dca100fb8
“[for inst] more bleedless than resurrection inst (...) on a song with mostly piano, sounds considerably less muddy than resurrection inst... maybe overall slightly more noise... but the noise isn't bothersome” Resurrection inst seems to be fuller at times, but also noisier, and when the song has the noise in it intentionally, ress inst tend to pick it up. - rainboomdash. Continues -
“deux just seems to make some things super muddy and/or break them up... like an instrument here is just really wobbly with deux, probably because it doesn't have the fullness required.
hyperace sounds better in areas here, but noisy in others.. I guess you really gotta use all 3 models if you want a great result”
“Love the new model, besides it being on par with v1e in terms of fullness (on some tracks it's even fuller) without the noise tradeoff, it's also noticeably better in terms of bleedlessness as well, it managed to completely remove some faint choir/vocals pretty much all other models weren't able to remove. Most definitely my favorite model so far.” - Shintaro
Deux doesn't work well with VHS recordings - MarsEverythingTech
“vocal model picks up some things super well, like certain backing vocals or there was yelling in the background that it picked up that fv7beta didn't (...) generally slightly less full than fv7beta3 from the songs I tested, but has less instrumental bleed and slightly less noise.” - rainboomdash
“[voc] sounds great! I don't hear the dips in the vocals from some songs that are overly compressed” - Rage313
“deux vocal stem has more backing vocals than gabox fv7 beta 1-3 and other models I've ever tried, but it may be rather noisy on silent parts or fadeouts” - makidanyee
“I exported it using fp16 just for the smaller size [only 432 MB, the] quality is the same.” - becruily. Trained “almost” from scratch on a rented GPU.
Gabox tried to convert it to just an instrumental model, but the quality got worse.
- (MVSEP) “New model MVSep Xylophone (xylophone, other) has been added. Demo: https://mvsep.com/result/20251223210226-f0bb276157-mixture.wav
- Inst_GaboxFv9 Mel-Roformer has been released
Inst. bleedless 36.20, fullness 37.19, SDR 16.56
- unwa released bs_roformer_inst_hyperacev2 model (so alongside the voc v2)
incompatible with UVR | use MSST | makidanyee’s Colab
Inst. bleedless 37.87, fullness 38.03, SDR 17.40
“HyperACE is def way better than v1e+. but I have also noticed it's more noisy” - rainboomdash
“The model files (bs_roformer.py) for v2_voc and v2_inst are the same. (...)
The aura_mrstft score was further improved, and the SDR also increased. (...)
Incidentally, this new model outperformed v1e+ on all metrics. (...) holds the highest aura_mrstft score on the instrumental side of the Multisong dataset.” - unwa
https://mvsep.com/quality_checker/entry/9475
“kept the chants in [one song], resurrection inst isn't really much better, though...
it takes a high vocal fullness model to extract these (like fv7beta1, even fv7beta2 isn't enough), maybe that's why the instrumental models are getting confused... Surprised it's still very much there with resurrection inst…
I'm surprised; hyperacev2 fixed vocal bleed on one song compared to the previous version. Does seem like vocal bleed with hyperace v2 is a bit better
it's not removed, just quieter in areas where it did happen” - rainboomdash
“picks up vocals better than the first model, but along with the vocals, it muffles other instruments, such as a synthesizer.” - Halif
- Previously released FNO inst model by unwa now has a separate Colab #3 to run it
- Along with the inst variant, unwa released BS-Roformer HyperAce vocal model v2
separate Colab #2 | makidanyee’s Colab | (incompatible with UVR, use MSST)
Voc. bleedless 34.08, fullness 19.10, SDR 11.40’ (there was no v1 variant of HyperAce)
“The model files (bs_roformer.py) for v2_voc and v2_inst are the same.”
“aims to outperform Big Beta 6X overall, and I believe it has achieved that goal.
Scores for all metrics except bleedless and log_wmse have improved compared to 6X.” - unwa
Is it better than Beta 6X?
“hard to say, imo, tbh
it has some added noise, especially during some quieter parts that big beta 6x doesn't have. Noisier in those areas a lot of times than even fuller models. unwa thinks it's overall better” - rainboomdash
“Compared to the Resurrection vocal model, this new model achieves higher scores across the board except for the bleedless score. (...) This model aims to outperform Big Beta 6X overall, and I believe it has achieved that goal.
Scores for all metrics except bleedless and log_wmse have improved compared to 6X.” - unwa
“I feel like most songs hyperace gets a good bleedless result but ill probably stick w voc fv7 beta3 or revive e3” - 5b
“It seems to be like.. trying really hard to pull certain stuff out that is causing noise, I think
it doesn't sound that bleedless to me, due to that reason.
(...) there's some noise during some quieter parts but I'll def be adding this to the list of models I use. I was testing against fv7beta2, during louder parts there seemed to be less noise, but during quieter parts there seemed to be more noise (...) I'll prob use fv7beta2 for the most part still, but I'll try adding hyperace vocal model to the mix of models I use. lol, soon I'll be using 10 models throughout one song” - Rainboomdash
Note: It uses its own inference script (bs_roformer.py) different from the previous one, and it’s also incompatible with UVR. “You can use this model by replacing the MSST repository's models/bs_roformer.py with the repository's bs_roformer.py.”
To not affect functionality of other BS-Roformer models by that file, so older BS-Roformers will still work, you can add it as new model_type by editing utils/settings.py and models/bs_roformer/init.py here (thx anvuew).
For error while installing the py file for HyperACE model in Sucial’s WebUI:
from models.bs_roformer.attend import Attend
ModuleNotFoundError: No module named 'models'"
The fix: “SUC-DriverOld/MSST-WebUI use the name "modules" and ZFTurbo/Music-Source-Separation-Training use the name "models". And Unwa's bs_roformer.py that you replace with, also use "models". So you'll have to do some coding and symlink to make it work.” - fjordfish
Training details: “In v2, the TFC-TDF module used in models like the MDX23C has been added to the FreqPixelShuffle module.
Additionally, frequency-domain downsampling is now performed downstream of the Backbone module.”
Some components in the SegmModel module were implemented based on this paper: https://arxiv.org/abs/2506.17733
“HyperACE is the core of YOLOv13 (and when i asked about that, he replied with this graph [from here])”
- (MVSEP) “I added a new Crowd removal model based on BSRoformer architecture. It's available in "MVSep Crowd removal (crowd, other)" with name "BS Roformer (SDR crowd: 7.21)". SDR increased from 6.27 up to 7.21.” - ZFTurbo
Some people have issues with it:
"The 6.27 one removed all kinds of crowd noises, sound effects and general noise, this one only removes random bits of music!"
- HyperACE model added to uvronline.app
- Gabox uploaded an experimental “test” model “instmel.ckpt” loosely on some hosting “(muddy, no losses)” - Gabox
(it won’t work in the custom import Colab as it doesn’t support direct linking):
https://gofile.io/d/jJbBtm
Also, a small clarification for the old models was provided:
fv7b - uses mse loss
fv7z - uses l1 loss
fv7 - uses stft loss
- cyatarow with unwa found a way to use MSST natively on Windows on RX 9060 using ROCm and released PyTorch for ROCm for Windows without having to use WSL - click
- Instruction for MSST-WebUI has been made too - click
- Our user had some success in porting our Colabs to Keggle. Inference Colab and Apollo ones have been made so far, but the porting process seems to be rather straightforward.
Be aware that it might gradually fall behind with any newer models published in later periods.
https://www.kaggle.com/code/zzryndm/music-source-separation-training-inference-webui
https://www.kaggle.com/code/zzryndm/apollo-colab-inference-i-fucking-give-up-with-this
How it was done -
“I didn't even have to debug anything. Pasting the Colab's code without changes worked miraculously.
Kaggle doesn't support markdown IN code cells but I found a workaround using U+200C for the variable names.
Also set the acceleration to GPU P100. It's better than t4x2. I learned that through the "patience is a virtue" route.
And trust me when i say it was NOT fun” - ryn (xxml)
- Gabox voc_fv7 beta 3 added to the Colab
- (MVSEP) “Four new models have been added:
1) MVSep Percussion (percussion, other) Demo: https://mvsep.com/result/20251128141738-f0bb276157-mixture.wav
2) MVSep Keys (keys, other) Demo: https://mvsep.com/result/20251128142835-f0bb276157-mixture.wav
3) MVSep Brass (brass, other) Demo: https://mvsep.com/result/20251128142905-f0bb276157-mixture.wav
4) MVSep Woodwind (woodwind, other) Demo: https://mvsep.com/result/20251128143157-f0bb276157-mixture.wav
- Unwa BS-Roformer HyperACE inst model has been added to a separate Colab
- Gabox released two new models (vocal and instrumental):
Mel vocfv7beta3 | yaml
Voc. fullness 21.82 | bleedless 30.83 | SDR 10.80
“beta 1 and 2... eh, pretty close to same instrumental bleed,
but beta 3 def a step up from the two songs I compared (...)
most songs so far, fv7beta3 is fuller than fv7beta1,
def less robotic sounding at times (when a voice gets quiet/hard to capture, and it just fails).
Just had another song where fv7beta1 was fuller than fv7beta3, but it was also a lot noisier
large majority of the songs I tested, fv7beta3 was fuller... I think fv7beta3 is usually a bit noisier than fv7beta1? But also sounds fuller in those cases, I'd say it's generally worth it
instrumental bleed, usually worse with fv7beta3 versus fv7beta1, but it depends
fv7beta2 is always less full/less noise, but only slightly less instrumental bleed than fv7beta1” - rainboomdash
Inst. fullness 27.07 | bleedless 47.49 | SDR 16.71
“this may be the last beta before the final model” - Gabox
The highest bleedless metric out of all instrumental models so far. But fullness is worse than even most vocal Mel-Roformers (including BS-RoFormer SW and Mel Kim OG model).
“On the fuller side, somewhere around inst v1e+, maybe a tiny bit below. The main thing I notice is it captures more instruments than v1e+, but isn't muddy like [inst Resurrection] (which also captures more instruments) (...) It can add a lot of crackling noise, though, more than v1e+ (...) can be a little on the noisy side sometimes... but it at least isn't muddy and sounds natural (...) I'd still ensemble if you want the noise reduced - rainboomdash
Vs fv7z slightly more buzzing, but fuller and better transients, the b doesn’t eat hihats in sparse mix.
(src)
- Unwa released BS-Roformer-HyperACE instrumental model | separate Colab
Inst. fullness 36.91 | bleedless 38.77 | SDR 17.27
(less fullness than v1e+: 37.89, but more bleedless: 36.53, SDR: 16.65)
Note: It uses its own inference script. “You can use this model by replacing the MSST repository's models/bs_roformer.py with the repository's bs_roformer.py.”
To not affect functionality of other BS-Roformer models by it, you can add it as new model_type by editing utils/settings.py and models/bs_roformer/init.py here (thx anvuew).
For error while installing the py file for HyperACE model in Sucial’s WebUI:
from models.bs_roformer.attend import Attend
ModuleNotFoundError: No module named 'models'"
The fix: “SUC-DriverOld/MSST-WebUI use the name "modules" and ZFTurbo/Music-Source-Separation-Training use the name "models". And Unwa's bs_roformer.py that you replace with, also use "models". So you'll have to do some coding and symlink to make it work.” - fjordfish
“Currently, this model holds the highest aura_mrstft score on the instrumental side of the Multisong dataset. (...)
Some components in the SegmModel module were implemented based on this paper: https://arxiv.org/abs/2506.17733
Simply put, it's a module that utilizes hypergraphs to capture global relationships and standard convolutions to capture local relationships, thereby generating the final “Correlation-Enhanced” feature map.
This weight is based on the following weights. Thank you, anvuew!
https://huggingface.co/anvuew/BS-RoFormer” - unwa
- The inference file got updated to fix error
Consider changing overlap from default 4 to 2 in the yaml of the model. The difference won’t be really noticeable for most people, but it will be faster.
“thirty minutes of audio [on 4090]:
FNO was 71.38 seconds, HyperACE was 120.37
so HyperACE is about 2x longer than FNO (...) Does seem like HyperACE is picking up more instruments than v1e+
does seem like slightly worse vocal bleed overall (still need to test this more, though)... haven't encountered the super tinny vocal bleed like v1e+, at least
still fails to pick up that brass instrument on one song... Not really any worse than v1e+, though (...) resurrection inst does sound more muddy, but also a lot less noise... which makes sense... IDK, a little muddy for my tastes.
I did find one song/spot and resurrection inst was on par with HyperACE in picking up the wind instrument, v1e+ lost it for a bit.
I have found in the past that resurrection inst generally picks up more instruments than v1e+ (...) fullness of HyperACE is much closer to v1e+ than resurrection inst (...) it gets pretty staticy compared to v1e+ [on some drums] (...) v1e+ does this to a lot less extent
it's not super common, though… (...) I'm very confident in saying HyperACE picks up more stuff than v1e+.
Resurrection inst does pick it up much better than v1e+, but I think it's still too quiet
resurrection inst really does just pick up so much more instruments, despite having a lot less fullness” - rainboomdash
“fullness that is comparable to v1e+, but has significant more vocal crossbleeding in instrumental than BS Roformer Resurrection Inst, but still less than v1e+ and v1e” - dca100fb8
“the best instrumental model ive ever heard
Unbelievable how realistic it sounds
especially with bass and piano - PezZHasACat/pezz23
Might have problems with flute in specific songs - Hen
- (MVSEP) “We have released a new model 'MVSep Lead/Rhythm Guitar (lead-guitar, rhythm-guitar) '. It has two variants:
1) Two-stage model (SDR: 9.21) - Best guitar model applied, and then 2-stem model is used which can separate lead/rhythm guitar.
2) One-stage model (SDR: 9.02) - Single model is applied, which was trained on a 3 stem dataset.
They can give pretty different results, so worth trying both.
Demo: https://mvsep.com/result/20251120090832-f0bb276157-mixture.wav
- (MVSEP) We have released the "MVSep Plucked Strings (plucked-strings, other)" model.
Demo: https://mvsep.com/result/20251120092757-f0bb276157-mixture.wav” - ZFTurbo
- fr4z49 reported that they managed to use MSST with ROCm 7 and 6 on Linux and AMD RX 7600 for fast separations. Officially, it’s not supported by AMD, but works, although your mileage might vary from GFX to GFX (range of GPU models inside various generations/archs).
- Probably, 5700 XT with some older ROCm versions (e.g. older than 6.33) might work too (e.g. HIP 5.7 and ROCm around 5.2.* - src, although you can try out 6.21 or 6.2.x to ensure, as it could happen that some earlier 6.x wasn’t supporting RX 5700 XT correctly, while e.g. for RX 6000 ROCm 6.24 worked in some apps at some point, but more up-to-date information might be found in some ZLUDA guides, as it needs ROCm too - some suggestions here).
- The most performance gains on ROCm 7 might be potentially observed on officially supported GPUs like Instinct MI350 CDNA 4, providing even 3-7x performance gains over 6.0 in some applications (more).
- Official support for RX 400/500 (a.k.a. Polaris/GCN 4/GFX803) GPUs support was dropped, but you can follow this repo for unofficial ROCm 6 support.
Or for ROCm 5, this Ubuntu guide (it might even potentially work from Windows using WSL [if using at least Ubuntu 22.04 LTS] with almost no GPU performance overhead). There seemed to be some issues building Torch (UtilsAVX512.cc/tensorpipe) on Python 3.13, fixed on Python 3.10, and maybe 3.11.9.
Also, there seems to be some Arch Linux community package to install Pytorch still compatible for these GPUs (click).
Or also might be potentially supported with some other specific versions of ROC, e.g. 5.7.2 and also described above:
export ROC_ENABLE_PRE_VEGA=1 (deprecated in ROCm 6; might help for lacking dependencies or wheel building issues). Or check out also this:
https://github.com/AUTOMATIC1111/stable-diffusion-webui/issues/10435#issuecomment-1555399844, or alternatively follow below instructions:
https://pytorch.org/get-started/locally/ and then execute:
pip3 install torch torchvision torchaudio --index-url https://download.pytorch.org/whl/rocm5.4.2
- Since then also ROCm 6.4.4 allowing using PyTorch natively on Linux and Windows on RX 7000 and 9000 was released (more), but it wasn’t tested yet (DL). You might get the error using at least WebUI.
- Also, you might want to experiment with using ZLUDA in UVR (CUDA>ROCm translation layer - some suggestions here).
>fr4z49 ROCm report:
- “I managed to make MSST-WebUI work [on Linux] with:
Torch 2.10.0.dev20251110+rocm7.0
on RX 7600
(...) it seems like ROCm 7.0 is about a second faster [than 6.x]”
(probably by adding just pip install before it)
Turns out that if you do:
export TORCH_ROCM_AOTRITON_ENABLE_EXPERIMENTAL=1
it uses waay less VRAM and processes even faster.
inst_V1e_plus batch_size=2 overlap=3 chunk_size= 485100, 51.78s/it [3:50 of audio in 61 seconds]
For ROCm 6.x (a tad slower, might work on more GPUs) use:
torch 2.9.0+rocm6.3 torchvision0.24.0+rocm6.3 [--index-url https://download.pytorch.org/whl/rocm6.3]
Or older version suggested above.
Thanks, fr4z49.
- yxlllc’s harmonic noise separation VR (6 or 5.x model, unsure) if someone was interested:
https://github.com/yxlllc/vocal-remover/releases/tag/hnsep_240512 (July 2025)
- Gabox released beta 2 of vocfv7 Mel-Roformer “fullness went down a little bit”
Voc. bleedless: 31.55, fullness: 20.44, SDR: 10.87
https://huggingface.co/GaboxR67/MelBandRoformers/blob/main/melbandroformers/experimental/vocfv7beta2.ckpt | yaml | Colab (TL;DR in the vocals section)
“still quite a bit fuller than big beta 6x, but has less noise than even fv4 (also a bit less fullness, of course)” at least when the instruments are loud, fv7beta2 is usually quite a bit less noisy than fv4, while still maintaining a decent amount of fullness... it is a bit less, but not too much (...) both are pretty noisy with fv4 (...)
still gonna have an issue with backing vocals compared to fv7beta1 sometimes... (makes sense, it's a less full model). (...) Fv7beta2 has still been significantly better with BV than fv4, despite quite a bit less noise” but “significant issues on one song, while fv6/fv7beta1 didn't (...) Def an improvement over fv4. I'm really liking the balance of fullness and noise for most songs. fv4 and fv6/fv7beta1 are usually pretty noisy... this is less noisy, but still has a good amount of fullness.” “Where the noise was undesirable and I ensembled fv4/fv6/fv7beta1 with big beta 6x, now I can just use this instead”.
“Fv7beta2 has still been significantly better with BV than fv4, despite quite a bit less noise” but “significant issues on one song, while fv6/fv7beta1 didn't”
“If the noise isn't an issue, and you just want fullness, fv6/fv7beta1 are still the best models. I'd say fv6 and fv7beta1 are better models than fv4, fullness/noise aside. It depends with fv7beta1 versus fv7beta2, sometimes the noise can be pretty significant with fv7beta1, and fv7beta2 may have the fullness you desire.
fv6 is usually more noisy/full than fv7beta1, but it just depends... I've had instances where it's less noisy/full than fv7beta1. But if you really want high fullness, fv6 and fv7beta1 are the choices. Sometimes fv6 can be quite a bit more noisy and the gain in fullness isn't worth it”.
- rainboomdash (thx).
“vocals sound very robotic with those models, however. Compared to fv4” - pipedream
- Anvuew released a new BS-Roformer vocal model:
https://huggingface.co/anvuew/BS-RoFormer
It also doesn't work on the UVR’s RTX 5000 patch - then use MSST instead.
On an M1 Mac, you will probably need to decrease chunk_size in the yaml a bit.
"On one song it was on par with the 2025.07 BSRoformer model on MVSep, at least to my ears (Tentative by System Of A Down)
The other song [Linkin Park - Part of Me] has some background vocals that are hard to get for a lot of voc/inst models, 2025.07 manages to get them while this model doesn't.
The instrumentals and vocals seem pretty good other than that" - ryanz48
“Guessing it's not high enough fullness
the l o s t is extremely muddied, and the other chanting is just gone.
The lower harmony at 0:49-0:53 is mostly gone, making it sound very thin
and a lot of the vocals just sound like they are breaking apart.
Hmm, big beta 6x is significantly better, it's def a fullness issue.. probably too high of a bleedless model for my tastes
big beta 6x still isn't super great here, the one I used for the other one I posted was fv7beta1, which is a fullness model.
yeah, big beta 6x seems more balanced, it's a bit fuller but not noisy to my ears, either
but I'm not using headphones, so I won't hear any minor noise easily.
yoooo, it properly doesn't capture the instrument here.
even FT2 bleedless gets tricked by this part, but this does just fine.
Maybe I'll try it out for this song.. most I'll still use higher fullness models” - rainboomdash
- Anvuew’s BS-Roformer Karaoke Model added to the inference Colab
- (MVSEP) “Eleven new models have been added:
1) MVSep Triangle (triangle, other) Demo: https://mvsep.com/result/20251104082053-f0bb276157-mixture.wav
2) MVSep Sitar (sitar, other) Demo: https://mvsep.com/result/20251104082317-f0bb276157-mixture.wav
3) MVSep Harpsichord (harpsichord, other)
Demo: https://mvsep.com/result/20251104082809-f0bb276157-mixture.wav
4) MVSep Tuba (tuba, other) Demo: https://mvsep.com/result/20251104082845-f0bb276157-mixture.wav
5) MVSep Bassoon (bassoon, other) Demo: https://mvsep.com/result/20251104083211-f0bb276157-mixture.wav
6) MVSep Congas (congas, other) Demo: https://mvsep.com/result/20251104083239-f0bb276157-mixture.wav
7) MVSep Bells (bells, other) Demo: https://mvsep.com/result/20251104083305-f0bb276157-mixture.wav
8) MVSep Ukulele (ukulele, other) Demo: https://mvsep.com/result/20251104083332-f0bb276157-mixture.wav
9) MVSep Dobro (dobro, other) Demo: https://mvsep.com/result/20251104115825-f0bb276157-mixture.wav
10) MVSep Wind Chimes (wind-chimes, other) Demo: https://mvsep.com/result/20251104115849-f0bb276157-mixture.wav
11) MVSep Accordion (accordion, other)
Demo: https://mvsep.com/result/20251104115916-f0bb276157-mixture.wav” - ZFTurbo
- “Yeah, I just tried [the bells] on a drum loop sample with sleigh bells I wanted to isolate, and I got a rude awakening lol”
- “That's why that model name is confusing, lol. What they mean is tubular Bells or chimes. There's currently no sleigh bells model, but the Tambourine model may work”
- “It worked (...) I used drumsep on it before”
“I finally was able to extract the lead guitar in this song using the dobro model, but i noticed how the bass synth is leaking in the dobro stem”
“So I thought, what if I just remove the bass in the source and try again? I did, and now it doesn't pick the lead guitar anymore”
- (uvronline) “Added two new models:
BS-RoFormer Kar (anvuew)
De-reverb Room (anvuew)” - Aufr33
The latter is also added to the Colab.
- (MVSEP) “We added MVSep Synth (synth, other) model. Synth is included the following stems: Synth, Synthesizer, Synth Pad, Synth Bass, Synth Vocals, Synth Strings, Synth Percussion, Synth FX, Synth Keys, Synth Brass, Synth Guitar, Synth Flute, Synth Ambiant.
Demo: https://mvsep.com/result/20251031214429-f0bb276157-mixture.wav” - ZFTurbo
“So, my initial thoughts. The model works great for certain kinds of sounds e.g. leads, pads, plucks. But it's tricky to predict what it'll do, so might be safer to get rid of the other stems first, and then using synth on what's left if you need further separation.
Some examples:
It isn't picking up synth basses for me, so use BS-Roformer SW for that.
It also sort of picks up synth brass, but wind model is catching that better, at least on the stuff I ran. The same could possibly be said for synth guitars.” - Musicalman
“I tested it on a front channel rip from a cue from the TV show CHUCK, and it ripped the synth lead. But it did not isolate the synth "effect" at the end of it” - fal_2067
“is good, but sometimes it can't detect some bass synths for some reason, still not a big problem since an SW does the work almost every time. And sometimes it picks the strumming of guitars. And also it seems to fail if they are vocals or harmonies.” - smilewasfound
“Most of the time gets more out of the song rather than taking out guitar, keys, bass separately until synths are left over. Sometimes very few misses or stem bleeds through, but overall very impressive!! It also picks up Vibraphone” - Tobias51
“Synth stem seems way too muddy on full songs, (...) But much better than I was expecting, I'll be honest. It messes up a lot on full songs, it seems. It seems like it's more like a stem remover than an isolator to me. The no synth stems sounds very clean. Tried it on Dolby stems, same thing, the no synth stem was very clean, synth stem sounded a bit muddy though. That's fine though. Gets very confused with bass though.
Seems to also have a lot of the classic MVSep phase issue, where for some reason half the stem is in the synth stem and the other half is in the no synth stem
and inverting it cancels it out (literally almost all models have this issue on MVSep it's very strange). It's much less than the other models, but yeah, it still happens. (...) Using any other model in UVR or uvronline don't have this issue. (...) I put a bass guitar stem from Fortnite festival to test and it has the phasing issue, it inverts” - Isling
“I put it the synth model through a really packed song with no synth to see if it would get tripped up, it didn't other than some bass at the end.
Which actually didn't get picked up by the bass model, so even that is a win” - dynamic64
- “Several SATB (soprano, alto, tenor, bass) choir models I trained ages ago, currently only a scnet_masked model is available, but I did have Demucs, MDX23C, and standard SCNet models that I will upload to this link when/if I find them, although I'm pretty sure the scnet_masked model was the best in the end:
https://drive.google.com/drive/folders/1BpPgtlDk0yqrlArmrq9vnYErixb8I8zJ?usp=sharing” - Dry Paint Dealer
Treat it as proof of concept.
“I've tried the SCNet one, it's really noisy, and it has a lot of bleed, it kinda works. I can see the potential on this kind of model ngl.” - smilewasfound
“You can’t install [the VR ones] into UVR since that only supports VR v5 [and 5.1] not VR v6”
- (MVSEP) Seven new models have been added:
1) MVSep Electric Guitar (electric-guitar, other). Demo: https://mvsep.com/result/20251031064813-f0bb276157-mixture.wav
2) MVSep French Horn (french-horn, other). Demo: https://mvsep.com/result/20251031072529-f0bb276157-mixture.wav
3) MVSep Banjo (banjo, other). Demo: https://mvsep.com/result/20251031095934-f0bb276157-mixture.wav
4) MVSep Marimba (marimba, other). Demo: https://mvsep.com/result/20251031100024-f0bb276157-mixture.wav
5) MVSep Glockenspiel (glockenspiel, other). Demo: https://mvsep.com/result/20251031100134-f0bb276157-mixture.wav
6) MVSep Timpani (timpani, other). Demo: https://mvsep.com/result/20251031100232-f0bb276157-mixture.wav
7) MVSep Harmonica (harmonica, other). Demo: https://mvsep.com/result/20251031100508-f0bb276157-mixture.wav” - ZFTurbo
“The Harmonica model is hit or miss.” - musicbybrooks
“Wow, the electric guitar model is really neat. One thing I noticed is that it seems to be better than other models at picking up midi/synth lead guitars. At least on stuff I tried.
I think it also gets tripped up a bit more by weird FX and synth sounds being partially flagged as guitar. An interesting model, though, for sure.” - Musicalman
- (MVSEP) “The karaoke model by anvuew has been added under the algorithm "MVSep Karaoke (lead/back vocals)". It is available as the option "BS Roformer by anvuew (SDR: 10.22)" - ZFTurbo
For some reason it seems to give worse results than the ckpt anvuew shared.
- Dear friends at Apple Music. Please stop harassing labels and their sound engineers for making Atmos mixes using our and yours awesome AI models for audio separation. The artificial artifacts you're solely looking for in spectrograms in separate channels are inaudible in the entire tracks. The tracks are well mixed and accepted by major labels, but rejected by your lazy ass incompetent bullshit. The quality of Atmos mixes got better since the very beginning, and either your employees, or your algorithms, or both, do a lazy job without even hearing the shit on their own, while still rejecting stuff without sensible reason! You're making things nasty difficult for artists who lost their multitracks for certain legacy songs, rendering re-releasing of their albums in Atmos potentially impossible. Bring it up with the executives. Get your shit together, for fuck sakes!
- Full release of mesk’s rifforge Mel-Roformer model focused on inst/voc separation for metal music
“The model can have some quirks (just like most models) but it's all around clean for me to release.”
Training details:
“Characteristics:
This is a dimension 512 depth 24 model (so fairly large file size at 1.9 GB!), with an SDR of 14.2436.
It's finetuned from an older Melband Roformer checkpoint with an SDR of 13.7.”
- Gabox released experimental BS-Roformer karaoke model | metrics
It gives the same error for RTX 5000 UVR patch users as the avuew’s model.
- New ensemble (avg) of anvuew’s and becruily & frazer karaoke models was evaluated on the leaderboard (metrics lower than BSkarfrazerBecruily+BSkarMVSEP+MBkarGaboxV2 SDR-wise). Probably you could make a fusion model out of the two to save on inference time in cost of slight SDR decrease (both use the same config so it might work).
- erosunica found out that BS-Roformer SW drums is “really good to remove some SFX and foley, way better than DnR v3”
- Gabox released voc_fv7 beta 1 Mel-Roformer model | yaml | Colab
Voc. fullness: 21.21, bleedless: 30.81, SDR: 10.96
"Just a better fv4 it seems, better bleedless" (fullness: 21.33, bleedless: 29.07, SDR 10.58)
vs voc_fv4 "It is noisier.. Kinda closer to beta 5e?” “It's slightly less noise and fullness than beta 5e but picking up the backing vocals REALLY well, significantly better than beta 5e”
But it's pulling the backing vocals out even better than 5e” “the backing vocals are so good!
“it does have significant synth bleed, too... it at least wasn't coming through at full volume
when I say fullness, I specifically mean how muddy it sounds” - Raiboom Dash
- Anvuew released a new Karaoke BS-Roformer model
https://huggingface.co/anvuew/karaoke_bs_roformer
https://mvsep.com/quality_checker/entry/9180
UVR users will encounter “ModuleNotFoundError: "No module named 'torch._dynamo.polyfills.fx'" with this model (consider MSST instead). Maybe users of RTX 5000 won’t encounter that issue due to newer PyTorch in dedicated patch. Sadly not - even with CPU only. Even more, the issue seems to exist only on RTX 5000 patch.
“karaoke anvuew extracts lead vocals a bit better than karaoke becruily frazer, and in some parts, the lead vocals from karaoke anvuew still sound brighter compared to karaoke becruily frazer, which sounds a bit more compressed. oh, and for some reason, the becruily frazer model doesn’t detect vocals with radio effects, while anvuew’s model handles them just fine” - neoculture
“lead vocals leak into instrumental (...) Mel Becruily and Frazer’s BS don’t have this problem”
In that case, maybe “isolate the acapella first in almost all cases of using a karaoke model”
Demo:
https://pillows.su/f/671f60ffdd615eb2613c78dca70319fe
https://pillows.su/f/391ec7ba8353a086989c9c0934321260”
- (MVSEP) “I added a new DeReverb model https://huggingface.co/anvuew/dereverb_room
by avuew. It's available in Reverb Removal (noreverb) [by choosing] DeReverb room by anvuew (BSRoformer). It works only for vocals. Since it is a mono model, it processes 2 stereo channels independently.” - ZFTurbo
Demo: https://mvsep.com/result/20251017064532-53be20aa17-10seconds-song.wav
- (MVSEP) Four new models have been added:
1) MVSep Tambourine (tambourine, other). Demo: https://mvsep.com/result/20251015221411-f0bb276157-mixture.wav
2) MVSep Oboe (oboe, other). Demo: https://mvsep.com/result/20251015221618-f0bb276157-mixture.wav
3) MVSep Clarinet (clarinet, other). Demo: https://mvsep.com/result/20251015221718-f0bb276157-mixture.wav
4) MVSep Digital Piano (digital-piano, other). Demo: https://mvsep.com/result/20251015221944-f0bb276157-mixture.wav
Sometimes it can also work better for normal piano and make even better work then SW if it works.
“Absolutely fantastic for an epiano. i just put your example song through mvsep with the bs sw piano model as a comparison and bs sw did terribly. Ofc your epiano model picked up all of the epiano in a clean way”
“Pretty impressive. Besides it being more full than other piano models (in most cases), it's also by far the only piano model that doesn’t mistakenly pick up other instruments like tubular bells as piano.”
“From what i tested is more for midi piano, i tested with some tracks with that kind of midi sound and it worked way better that SW.”
- (training) Becruily made a modification of dTTnet arch working in MSST (DL).
“They report very good performance on vocals with low parameters” - Kim
Back in the end of 2023, one indie pop song from multisong dataset (of the two there) received the best SDR - Bas Curtiz
“Better than SCNet imo, remains to see if it can beat rofos” - Becruily
“Not fast to train. I'm back with vanilla mdx23c. Trying a config to train model with less than 4GB VRAM, (...) with my 1080Ti and batch_size=1, chunk_size is around 1.5sec)” - jarreou
Installation instruction:
“In the latest MSST [at least for 13.10.25]
add the ddtnet folder to "models" and replace your settings file in utils with this”
The mod breaks compatibility with the authors' checkpoint.
“The weird thing is, it sounds like a fullness model despite not being one, I barely can find dips in instrumentals. ddtnet vs kim melband, if anyone is curious” - bcr
“Also keep in mind authors trained with l1 loss only, default in MSST is masked loss”
“l1 loss when dataset is noisy, mse loss when dataset is clean”
“the loss is defined from msst, but in the original dttnet it was in the code itself
you can just --loss l1_loss”
@jarredou “I copied your tfc and tfc_tdf classes to my files (and used that latest stft/istft I sent) - and seems to be better, just like the og dttnet
the tfc/tdf fixed the nan issue for me (...)
Keep in mind, ddtnet was trained only with musdb and has 10-20x less params while being comparable in quality”
“the authors checkpoints had 16khz cutoff because dim_f was smaller than nfft/2
if you want to train model with cutoff it's fine, if you want fullband then dim_f must be half of nfft + 1” - becruily
Hit our #dev-talk for more.
- New sites added to Site and rippers (deezmate.com and tidal.qqdl.site).
Qubuz remains or defunct/problematic for now.
- We have numerous reports about some models like Unwa Resurrection inst having problems on AMD (and probably Intel GPUs) in UVR, returning “Invalid parameter” error. In that case, uncheck GPU Conversion (but it will be slower). If you find a fix, please let us know on the Discord (link at the top of the doc).
- if you deal with slow separation times on becruily & Frazer karaoke model, decrease chunk_size to 160000 on 8GB GPUs. As long as decreasing chunk_size on CUDA (NVIDIA) doesn't seem to affect separation times, it's not the case with DirectML (AMD/Intel), if you're exceeding your VRAM, but it still doesn’t crash.
- Added anvuew BS-Roformer Dereverb Room (mono) model to the inference Colab
- Sir Joseph released a Colab for A2SB: Audio Restoration NVIDIA’s upscaler.
It’s very slow - on 4070 Super it was slow already, and in free Colab we got Tesla T4 with RTX 3050 performance with 12GB of VRAM instead of 8 - memory issues might occasionally occur, Colab Pro recommended.
Only inpainting doesn’t work (feature for filling silences if exist or missing parts) - “I couldn't fix the error. If anyone solves it, I'd be glad if they let me know so I can update it too.”. A2SB should rather surpass AudioSR, Apollo and FlashSR (it does at least metrically).
https://colab.research.google.com/drive/1ThenZDCRTJKV1I_ax17XGWmkB1qoKrFs?usp=sharing
- Gabox released inst_fv4 model. Don’t confuse it with inst_fv4noise - the regular variant was never released before (and with voc_fv4).
https://huggingface.co/GaboxR67/MelBandRoformers/blob/main/melbandroformers/instrumental/inst_Fv4.ckpt (yaml) | Colab
“Seems to be erasing a xylophone instrument. Does sound not too noisy and not muddy, I like it. (...) A little noisy with piano (I split the song up and process with resurrection inst there). (...) Does have some issues that resurrection inst doesn't have, but it doesn't sound muddy! It usually works great. (...) In my opinion, fv4 still has vocal traces, I don't know if in all of its songs and v1e plus doesn't have them, but the noise can bother you even though it's not much. Does have more vocal bleed at times. I think a lot of what I thought was vocal bleed was a synth, it did a pretty good job... There was one segment on a song where it caught vocal residues, though” - rainboomdash
- neoculture released a Mel-Roformer instrumental model focused on preserving vocal chops
Inst. fullness 39.88, bleedless: 32.56, SDR: 14.35
https://huggingface.co/natanworkspace/melband_roformer/blob/main/Neo_InstVFX.ckpt (yaml) | Colab
“great model (at least for K-pop it achieved the clarity and quality that no other model managed to have) it should be noted that it has a bit of noise even in its latest update, its stability is impressive, how it captures vocal chops, in blank spaces it does not leave a vocal record, sometimes the voice on certain occasions tries to eliminate them confusing them with noise, but in general it was a model that impressed me. It captures the instruments very clearly” - billieoconnell.
“NOISY AF, this is probably the dumbest idea ever had for an instrumental model. Don’t use it as your main one, some vocals will leak because I added tracks with vocal chops to the dataset. Just use this model for songs that have vocal chops” - neoculture
It was trained on only RTX 4060 8GB.
- Aname Mel trained a Mel-Roformer model called Full Scratch
Inst. fullness: 25.10, bleedless: 37.13, SDR: 14.32
Voc. fullness: 13.24, bleedless: 30.75, SDR: 8.01
(“trained from scratch on a custom-built dataset targeting vocals. It can be used as a base model or for direct inference. Estimated Training cost: ~$100”)
For state_dict error, update MSST to the last repo version:
!rm -rf /content/Music-Source-Separation-Training
!git clone https://github.com/ZFTurbo/Music-Source-Separation-Training
“and you must reinstall main branch's requirement.txt. (before it, edit requirements.txt to remove wxpython)” - Essid
Kim Mel model for reference:
Inst. fullness 27.44, bleedless 46.56, SDR: 17.32
Voc. bleedless: 36.75, fullness: 16.26, SDR: 11.07
- (MVSEP) “1) Model by baicai1145 was added in Apollo Enhancers (by JusperLee, Lew, baicai1145) with name Universal Super Resolution (by baicai1145) (...)
2) New option added for Apollo Enhancers (by JusperLee, Lew, baicai1145) - Cutoff (Hz). Sometimes it can be useful to cut higher frequencies before applying model.” - ZFTurbo
- baicai1145 released their own Apollo vocal restoration model, which surpassed Lew’s vocal V2 model metrically.
“with a 92-hour high-quality vocal dataset trained for 1 million steps.”
https://huggingface.co/baicai1145/Apollo-vocal-msst/tree/main
https://mvsep.com/quality_checker/entry/9105
21.24 vs 13.09 Aura MR STFT
(thx Essid)
ReminderL Apollo arch support was added to UVR too (acceleration work with NVIDIA GPUs only). Installing the model should be possible also there, although currently UVR seems to be incompatible with the model outputting following error:
KeyError: "'infos'"
- (MVSEP) “Two new algorithms have been added:
1) MVSep Mandolin (mandolin, other). Demo: https://mvsep.com/result/20250927132339-f0bb276157-mixture.wav
2) MVSep Trombone (trombone, other). Demo: https://mvsep.com/result/20250927132547-f0bb276157-mixture.wav” - ZFTurbo
- ROCm 6.4.4 now allows using PyTorch natively on Linux and Windows on RX 7000 and 9000, so you don’t need WSL with them - link
- BS-Roformer 6 stems added on uvronline
- “I added new version of MVSep Organ (organ, other) model: "BS Roformer (SDR organ: 5.08)". SDR increased from 3.05 to 5.08.
Demo: https://mvsep.com/result/20250924223759-0efd607228-song-organ-000-mixture.wav” - ZFTurbo
“I think the model is remarkable improvement” - totalmentenormal
“the result is great, no bleed from what I tested it on” - dynamic64
“much better isolation of the Hammond organs compared to the previous model. In places where the organ sound was not picked up before, it is now separated in the track” - lukasz2286
- Google Colab now allows pinning your environment to specific version having the same versions of packages, so maybe your notebook won't break in the future due to changes in the environment introduced by Google in the Colab with package updates.
For now there is only 2025.07 and the latest environment to choose from, and it's hard to tell if e g. 2025.07 environment will be gradually replaced along the time while new changes to the latest Colab environment will be made:
https://developers.googleblog.com/en/google-colab-adds-more-back-to-school-improvements/
To use it, go to Environment>Change environment type>Environment type version and choose 2025.07 option
- Essid reevaluated GAudio (a.k.a. GSEP) for the leaderboard.
https://mvsep.com/quality_checker/entry/9095
Inst fullness: 28.83, bleedless: 31.18, SDR: 12.59
The result would rather cover my observations that instrumentals rather have gotten worse over the years (at least since the last 2023 Bas' evaluation or even earlier, at least for certain songs). But it appears that the vocals might got better.
Despite the fact the metrics are worse than even the least bleedless free community models like even V1e, for specific songs where bleeding doesn't occur so badly, GSEP might be still interesting too try out to some limited extend, being a different architecture, sounding maybe less filtered. Also, mixdown of multi stem extraction instead, should rather have bigger bleedless metric, but since the appearance of instrumental Roformers, GSEP relevance for separation is rather faded.
- "Ensemble of 3 [karaoke] models "Mvsep + gabox + frazer/becruily" gives 10.6 SDR on leaderboard. I didn't upload it yet, but I had local testing.” - ZFTurbo
- fabio06844 shared his method for “very clean and full” instrumental lately.
1) Go to MVSep and separate your song with the latest Karaoke BS-Roformer by MVSep Team
2) On its instrumental stem use DEBLEED-MelBand-Roformer (by unwa/97chris)
Despite the fact that “the MVSep Team Karaoke uses the MVSep BS model to extract/remove vocals, then applies [the] karaoke model to that”, it was told to be not enough to just use BS 2025.07 model instead, leaving a little more residues.
- Aname released Mel-Roformer duality model.
“it's odd why the model is named duality, but it has a single target (and the file size of the ckpt confirms it further)” - becruily
It’s focused more on bleedless than fullness metric contrary to the unwa’s duality v2 model, but with bigger SDR.
Inst. fullness 24.36, bleedless 46.52, SDR: 17.15
“instrumental is really muddy” - Gabox
For comparison -
Mel Duality v2 by unwa
Inst. fullness 28.03, bleedless 44.16, SDR: 16.69
MelBand Roformer vocals by Kim
Inst. fullness 27.44, bleedless 46.56, SDR: 17.39
Instrumental public models with the biggest fullness metric -
Gabox Mel Roformer Inst_GaboxFv7z
Inst. fullness: 29.96, bleedless: 44.61, SDR: 16.62
Unwa BS-Roformer-Inst-FNO
Inst. fullness: 32.03, bleedless: 42.87, SDR: 17.60
- (MVSEP) “I added new SCNet vocals model: SCNet XL IHF (high instrum fullness by becruily). It's high fullness version for instrumental prepared by becruily.”
Inst. fullness 32.31, inst. bleedless 38.15, SDR 17.20
“One of my favorite instrumental models, Roformer-like quality.
For busy songs it works great, for trap/acoustic etc. Roformer is better due to SCNet bleed” - becruily
“It's better than BS Roformer (mvsep 2025.07 and inst Resurrection) at low frequencies, but bad at highs due the bleeding. I think it has better phase understanding, because it keeps the harmonics that were masked behind vocals cleaner (but it might not necessary be true to the source, it might just interpret/make up the harmonics instead of actual unmasking)” - IntroC
- Dry Paint Dealer Undr released Melband Roformer and Demucs (“or at least I think this is the correct model file”) Lead and Rhythm guitar models.
“My own very mediocre model for it that I never shared. It does work but has issues that I imagine any better executed model won't.”
“Wait, I think it separated doubles in vocals” - isling
Demucs model doesn't work in UVR, as it was trained on MSST, and not the OG code (I tried to workaround the bag_num issue before, and failed)
- (MVSEP) “Two new instrumental models have been added:
MVSep Harp (harp, other)
Demo: https://mvsep.com/result/20250921131108-f0bb276157-mixture.wav
MVSep Double Bass (double-bass, other)
Demo: https://mvsep.com/result/20250921131129-f0bb276157-mixture.wav” - ZFTurbo
“The BS-Roformer SW bass model should probably be used first to extract the double bass. Creates a better sound. This does not apply to bowed double bass.
Bowed double bass doesn't get picked up by BS-Roformer and therefore needs the double bass model. Good news is that bowed double bass is picked up in the strings stem so if you run the strings model you're good either way.” - dynamic64
- (MVSEP) “A new saxophone model based on BSRoformer was added. It has a much better metric compared to the previous [model]. SDR grew from 7.13 up to 9.77.
It's available in "MVSep Saxophone (saxophone, other)" with option "BS Roformer (SDR saxophone: 9.77)
Demo: https://mvsep.com/result/20250920151232-f0bb276157-mixture.wav” - ZFTurbo
“after testing this on a song where trumpet and sax play in unison, doing the trumpet model is cleaner than doing the sax model” - dynamic64
“Amazing. Tested it on one song, it got every single Saxophone Part from the song it seems lit. Can hear one small little bitty part of it where it tries to come in off the Sax part, however I can barely hear it” - cali_tay98
- GAudio (a.k.a. GSEP) announced their SFX (DnR) model in their API:
“DME Separation (Dialogue, Music, Effects)”
So far it’s not available for everyone on their regular site:
But the link on their Discord redirects to the site with a form to write an inquiry:
https://www.gaudiolab.com/developers
Shortly after entering the one or both of the links and logging on the first, you might get an email that $20 of free credits to access their API have been added to your account, and link to the API documentation:
https://www.gaudiolab.com/docs
- New Dango Karaoke model released
https://tuanziai.com/en-US/blog/68ca20c87c8c85686c1b4511
A lot of problems when songs don't have lead vocals in the center.
- (MVSEP) “I added new Karaoke model: "BS Roformer by MVSep Team (SDR: 10.41)" it's available under option "MVSep MelBand Karaoke (lead/back vocals)". Metrics.
In contrast with other Karaoke models, it returns 3 stems: "lead", "back" and "instrumental".
Example: https://mvsep.com/result/20250915192251-53be20aa17-10seconds-song.wav” - ZFTurbo
“If I had to compare it to any of the models, it is similar to the frazer and becruily model. Sometimes it does not detect the lead vocals specially if there's some heavy hard panning, but when it does, there is almost no bleed, and it works very well with heavy harmonies in mono from what I tested.” - smilewasfound
“becruily & frazer is better a little when the main voice is stereo” - daylightgay
“On tracks I tested, harmony preservation was better in becruily & frazer (...) the new model isn't worse, I ended up finding examples like Chan Chan by Buena Vista Social Club or The Way I Are by Timbaland where it is better than the previous kar model. The thing is, with the Kar models, it's just track per track. Difficult to find a model for batch processing as it's really different from one track to another” - dca100fb8
“I also found the new model to not keep some BGVs, mainly mono/low octave ones, despite higher SDR” - becruily
“I think I've found a solution for people who don't like the new model.
If you put an audio file through the karaoke model and then put the lead vocal result through that, it usually picks up doubles.
Which you can then put in your BGV stem if you'd like” - dynamic64
“it's definitely not as good as the one by frazer and becruily. SDR can be misleading sometimes” - ryanz48
becruily [“our model] uses 11.9 SDR vocal model as a base”
ZFTurbo “I started from SW weights”
“I've had fantastic results with it so far. Much MUCH better at holding the 'S' & 'T' sounds than the Rofo oke (for backing vox). Generally seems to provide fuller results .. but also the typical 'ghost' residue from the main vox can end up in the backing vox sometimes, but it's usually not enough to be an issue. I won't go so far as so say that it's replacing the other backing vox models for me entirely .. but it feels like the best of both worlds that Rofo and UVR2 provide.” - CC Karaoke
- (MVSEP) “We’ve added a mirror of MVSep (big thanks to okhostok): https://mirror.mvsep.com
If you have a problem with upload/download speed or can't reach the main site then try the mirror.
Report please if it helped you to speed things up.” - ZFTurbo
Some issues with being unable to click separate button for some users were fixed.
- Gabox released BS_ResurrectioN model | yaml
“It is a finetune of BS Roformer Resurrection Inst but with higher fullness (like v1e for example), it needs [MVSEP’s] BS 2025.07 (as a source/reference) phase fix [so you “should process the instrumental result using BS 2025.07 then put [it] as source in UVR GUI phase fix tool”]. I requested it because I found some songs where Resur Inst was producing muddy instrum results (...) I requested it not just for me because I saw other people were looking for something like v1e++” - dca
- anvuew released BS-Roformer Dereverb Room model | Colab
“specifically for mono vocal room reverb.” as most are recorded in mono.
Not that long inference compared to other Roformers.
“Really liking the fullness in the noreverb stem. Virtually all dereverb roformers I've tried sound muddy, but this one is just the opposite. (...) Other noises may interfere, and in my experience, makes the model underestimate the reverb. [The previous anvuew’s mono model] is way different [from] this one in every way. So, like I say, worth a shot.” - Musicalman. “WOAH this is insane. This would go viral if someone implemented in a plugin” - heuhew
We have reports about errors in UVR while using this model. Consider using MSST instead.
If you have stereo errors using MSST on stereo files, update MSST (git clone and git pull commands) or:
(it might work in your current version and not only in the linked repo too, but potentially the code will be located in a different line, the change will be pushed there later)
“Edit inference.py from my repo line 59:
Replace :
# Convert mono to stereo if needed
if len(mix.shape) == 1:
mix = np.stack([mix, mix], axis=0)
by :
# If mono audio we must adjust it depending on model
if len(mix.shape) == 1:
mix = np.expand_dims(mix, axis=0)
if 'num_channels' in config.audio:
if config.audio['num_channels'] == 2:
print(f'Convert mono track to stereo...')
mix = np.concatenate([mix, mix], axis=0)”
- jarredou
- BS-Roformer Karaoke model by becruily & frazer released | MVSEP | uvronline
Metrics better than even fused model gabox + aufr33/viperx and SCNet IHF below).
Make sure you don’t have the option “Vocals only” checked in UVR.
“After dozens of tests I can tell this (...) is the best (better harmony detection, better differentiation between LVs and BVs, sounds fuller, less background roformer bleed, better uncommon panning handling etc)” - dca
“it also can detect the double vocals” - black_as_night
It works the best for some previously difficult songs. Aufr33 and viperx model seems more consistent, but the new BS is still the best in overall - Musicalman
“my og Mel also catches some of the FX/drums, I guess quite a difficult one due to how it’s mixed” - becruily
“it does do better on mono than previous
sometimes confuses which voice should be the lead, but all models do that on mono in the exact use-case I normally test” - Dry Paint Dealer Undr
“In my opinion, this model is in no way inferior to the ViperX (Play da Segunda) — it's really very good. (...) I noticed that in the separation, the first voice still appears mixed with the second. The second voice, however, stands out more, but not completely isolated—in some passages, it still appears alongside the first. In short: the model better separates the second voice, but still presents some mixing between them.” - fabio5284
“The new karaoke model doesn't actually differentiate between lvs & bvs and there's some lead vocal bleeding in the instrumental stem” - scdxtherevolution
Fixes and expansion to the dataset and retrain of the model possible in the future.
“The dataset isn't correctly labelled, so in some training examples it was literally training the model to treat the backing vocal as the lead” - frazer
VS the newer BS-Roformer MVSEP team model above: “sound isn't as clear, but it does an infinitely better job at telling lead/BGV apart”
Becruily:
“I want to remind something regarding my (and the frazer) models
they're made to separate true lead vocals, meaning either all of the main singer's vocals, or if it's multiple singers - theirs too
this means if the main singer has stuff like adlibs on top of the main vocals, these are considered lead vocals too - they go together
if there are multiple singers singing on top of each other, including harmonise each other, and if there are additional background vocals behind those - all the singers will be separated as one main lead vocal, leaving only the true background vocals”
think of them like concert ready models - the output instrumentals will be ready to play in cases where all main vocalists are going to sing on top of the karaoke instrumental
ps: and yes, double/stereo lead vocals are still lead vocals, they're not bgvs (only in rare cases)
ps 2: if there are two singers singing the same melody and they don't harmonise each other - the model will most likely consider both singers as one lead vocal (again in rare cases one singer could be left) ”
- (MVSEP) “New Karaoke model based on SCNet XL IHF was added on site in "MVSep MelBand Karaoke (lead/back vocals)". Name of model "SCNet XL IHF by becruily (SDR: 9.53, metrics)". It has slightly worse metrics than the top Roformer model, but since it's different architecture it can give better results in some cases where the Rofo failed.
Demo: https://mvsep.com/result/20250908072226-f0bb276157-mixture.wav” - ZFTurbo
Iirc it's BVE or IHF unpublic ZFTurbo model retrain, and ckpt won't be public till further notice, as becruily said.
“SCNet is more bleedy in general despite me trying to reduce the leakage
it's recommended for busy songs, often captures proper lead vocals better than Roformer. Another use case is to ensemble it with Roformer to improve fullness” - becruily
“Oh, might be related to the lead vocals panning, it seems this model doesn't like when it's not center (...) I'm indeed noticing this model works really great on some songs that the Mel Rofo Karaoke had trouble with (...) I noticed that, this model, instead of creating crossbleeding between LVs and BVs, make them both quieter. I prefer that compared to previous models Plus, it handle songs which have lead vocals in the sides and BVs also in the sides better”
To fix bleed in back-instrum stem, use “Extract vocals first, but, “I noticed a pattern that if you hear the lead vocals in the back-instrum track already (SCNet bleed), dont try to use Extract vocals first because there will be even more lead vocal bleed” - dca
“Separates lead vocals better than Mel-Roformer karaoke becruily. It's not perfectly clean, sometimes a bit of the backing vocals slips through, but for now, scent karaoke model still the most reliable for lead vocals separation (imo)
https://pillows.su/f/df8c1791bceba5fe3ef6b16d310ec123
https://pillows.su/f/e1272a02c56e3d3eb7ba4007bbb0c4bd” - neoculture.
“the model seems to handle mono vocals better than melband but isn't as clean, lot of bleed” (extract vocals first was also used to test this) - Dry Paint Dealer Undr
Since the Mel Kar Becruily's model, the dataset is “larger” now, but still not “great”, and it might get eventually fixed, becruily said.
- (MVSEP) Four “new models for independent instruments were added:
1) MVSep Viola (viola, other) Demo: https://mvsep.com/result/20250907234931-f0bb276157-mixture.wav
2) MVSep Cello (cello, other) Demo: https://mvsep.com/result/20250907235225-f0bb276157-mixture.wav
“quite impressive” - dynamic64
3) MVSep Trumpet (trumpet, other) Demo: https://mvsep.com/result/20250907235543-f0bb276157-mixture.wav
“I can't get over how good the trumpet model is, it's so cleannn” - Shintaro
“trumpet struggles a bit on muted trumpet” - dynamic64
4) MVSEP Strings BS-Roformer (strings, other)
Demo: https://mvsep.com/result/20250907225920-f0bb276157-mixture.wav
The SDR has increased significantly compared to the previous MDX23C model, from 3.84 to 5.41. It is currently the best model on the leaderboard: https://mvsep.com/quality_checker/leaderboard/strings/?sort=strings” - ZFTurbo
“From some quick testing, it does not disappoint. Still playing with it, but atm it's exactly what I hoped for.” - Musicalman
“Yeah, I’m running some tests too with a few tracks that were really hard to separate, mostly ones with cellos or vocals that were too blended with the strings to isolate even with the latest inst/voc models, and it’s been working out surprisingly well.”
- anvuew released experimental BS-Roformer vocal model (nfft 4096, stft_hop_length 1024 “so not that large”) with 12 SDR measured on musdb18hq dataset. Might be worth checking: download (dead; newer model was released since then).
11.60 SDR on the same test set was previously achieved by one of the first Mel-Roformers trained by Bytedance on musdbhq + 500 songs (paper), although it wasn't nfft 4096.
It uses a very high 1024000 chunk_size in the yaml, so consider decreasing it when having memory issues, 500MB ckpt size.
- introC released a python script to get rid of vocal leakage in v1e+ model
- iZotope released Ozone 12. Separation still has Spleeter-like quality, but “it's unclear what they use” - Spleeter references disappeared from their readme (jarredou).
“the stems sound very bleedy and not at all usable” - becruily.
A notable new feature working competitively is their Delimiter.
- Ableton received its own stem separation feature in Live 12.3. It’s made in cooperation with Moises.ai. https://www.youtube.com/watch?v=uSahY-HGKt4
“doesn’t even sound good” - isling
It doesn’t use GPU, and has a slower High Quality setting too (single model for each stem opposing to default multi stem), but it can take even 20 minutes for 1 minute file on a slower CPU. At least default sounds more similar to Demucs than Roformers or SCNet archs, although files look like BS-Roformer judging by memory dump (model files are encrypted). Here are the low default model stems metrics - e.g. vocals only 8.71 SDR, but HQ option has bigger SDR than public SCNet weights released by ZFTurbo in MSST repo, but they're on a bleedless metric side, fullness is lower than in the public undertrained SCNet XL 4 stem model. Average SDR of the first 12 songs in the multisong dataset vs public SCNet XL: drums: 11.58 vs 11.22, bass: 12.25 vs 11.27 (thx jarredou).
“The boring thing is that you have to launch separation for each file manually (no batch processing). To nice things is that the separated stems are automatically saved individually in folder (no need to export them individually through Live's rendering and all issue that this can produce; different length...)” - jarredou
It seems like GPU support was later added or wasn’t noticed.
- It seems like we’ve received a step-by-step tutorial how to install the new Nvidia’s upscaler: click (thanks Pipedream)
- “I added BS Roformer flute model. It's available in "MVSep Flute (flute, other)". It superior comparing to SCNet version. SDR: 9.45 vs 6.27. More than 3 SDR difference.
Example: https://mvsep.com/result/20250830211041-f0bb276157-mixture.wav” ZFTurbo
- Thanks to Essid, metrics for following instrumental models were added to the models list:
INSTV7N, inst_fv8 (v2), inst_gabox3, Rifforge model, older mesk’s metal model, FVX, Bv1, Bv2 (b - bleedless, v - for version)
- “New Wind model based on BS Roformer has been added in MVSep Wind (wind, other):
Demo: https://mvsep.com/result/20250829230056-f0bb276157-mixture.wav
Results on quality checker: https://mvsep.com/quality_checker/entry/8933
It increased SDR +2.5 comparing to previous best model.” - ZFTurbo
“this one does not disappoint. At least not with the stuff I've tried so far. (...) the improvement is most noticeable with orchestral music. In heavy mixes eg. with lots of strings, the old models trip out. [The] new one is a lot more robust.” - Musicalman
“the model is not only cleaner but also detects some wind instruments that the previous one couldn't (specially baritone saxophones, I need to test it a bit more)” - smilewasfound
“the bs roformer wind model does really well with the other result and the violin model really is quite useful” - dio7500, dynamic64
- Suno now has stem separation feature “t's generative, so the separation isn't exact. Also, you apparently can't use it on like famous songs because they'll get flagged.” - Musicalman
”it sounds like shit tbh, tried it out” - dynamic64
- Gabox released experimental inst Mel-Roformer model (yaml) called just “fullness”.
“this isn't called fullness.ckpt for nothing.” - Musicalman
Inst. fullness: 37.66, bleedless: 35.53, SDR: 15.91 (thx Essid)
- (MVSEP) “We added 2 new algorithms for Acoustic Guitar (based on BS Roformer) and for Flute (based on SCNet XL)
1) `MVSep Acoustic Guitar (acoustic-guitar, other)` Example: https://mvsep.com/result/20250825095613-f0bb276157-mixture.wav”
“excellent, it's separating acoustic from electric very well, even in fuzzy, lo-fi recordings” - Input Output (A5)
“outperforms moises' model like crazy” - Sausum
“2) `MVSep Flute (fulte, other)` Example: https://mvsep.com/result/20250825095856-f0bb276157-mixture.wav” - ZFTurbo
“I tried the fulte model on stairway to heaven and it was so disappointing” - santilli_
- We have the first lucky person on the server who succeeded to actually use the new NVIDIA’s upscaler, and their messy AF code on Windows using Docker.
The output is mono, so you need to process each channel manually.
Also, it's extremely slow, even on 4070 Super, but results are “impressive”. More (don't expect step-by-step tutorial for now because the guy is “not tech support”).
- Unwa released BS-Roformer-Inst-FNO model (incompatible with UVR, use MSST and read special model installation instruction below).
inst. bleedless: 42.87, fullness: 32.03, SDR: 17.60
“very small amount of noise compared to other fullness inst models, while keeping enough fullness IMO. I don't even know if phase fix is needed. Maybe it's still needed a little bit.” dca
“seems less full than resurrection, which I would expect given the MVSEP [metric] results. (...) I'd say it's roughly comparable to gabox inst v7”
“I replaced the MLP of the BS-Roformer mask estimator with FNO1d [Fourier Neural Operator], froze everything except the mask estimator, and trained it, which yielded good results. (...) While MLP is a universal function approximator, FNO learns mappings (operators) on function spaces.”
“(The base weight is Resurrection Inst)”
Installing the model - instructions:
1. For Pytorch newer than 2.6, replace in utils folder by this models_utils.py (neoculture), or edit it manually:
“I had many errors with torch.load and load_state_dict, but I managed to solve them.
PyTorch 2.6 and later have improved security when loading checkpoints, which causes the problem. torch._C_.nn.gelu must be set to exception”
> “Add the following line above torch.load (at utils/model_utils.py line 479; 531/532 in updated MSST - old one doesn’t have utils folder and that py file):
with torch.serialization.safe_globals([torch._C._nn.gelu])
- unwa
> Or use PyTorch older than 2.6.
For old MSST without utils/model_utils.py replace that inference.py in the root MSST directory.
2. (linked model card for reference).
Replace this bs_roformer.py in models\bs_roformer folder, or edit it manually:
“You need to replace the entire "MaskEstimator" class in original bs_roformer.py from ZFTurbo (in models/bs_roformer folder) with the code provided by unwa [indention error fixed].
3. Also install this lib https://pypi.org/project/neuraloperator/” so:
“pip install neuraloperator==1.0.2”.
“Please note that since FNO1d appears to have been removed in the new version of neuraloperator, you will need to install an older version. [so not current 2.x]” - unwa
4. “Errors may also occur when using load_state_dict. In such cases, specify strict=False as an argument.(at utils/model_utils.py line 532)”
6*. To not affect functionality of other BS-Roformer models by that file, so older BS-Roformers will still work, you can add it as new model_type by editing utils/settings.py and models/bs_roformer/init.py here (thx anvuew).
For error while installing the bs_roformer.py file in Sucial’s WebUI:
from models.bs_roformer.attend import Attend
ModuleNotFoundError: No module named 'models'"
The fix: “SUC-DriverOld/MSST-WebUI use the name "modules" and ZFTurbo/Music-Source-Separation-Training use the name "models". And Unwa's bs_roformer.py that you replace with, also use "models". So you'll have to do some coding and symlink to make it work.” - fjordfish
7*. Seems like MSST might have some issues with GPUs other than corresponding archs to RTX 5000, 4000, 3000, H100, H200 or maybe using ROCm, resulting in SageAttention error, forcing slower CPU separation.
In that case, ensure you have compatible CUDA/torch/torchvision/torchaudio installed:
Compatible CUDA version requirement for GTX 1660 is 10 (e.g. on GTX 1060, Torch 2.5.1+cu121 can be used), but pip doesn’t find such package of Torch. To fix it:
*a) Check out index-url method described below:
pip install torch torchvision torchaudio --index-url https://download.pytorch.org/whl/cu118
or
pip install torch==2.3.0+cu118 torchvision torchaudio —-extra-index-url https://download.pytorch.org/whl/cu118
or
pip install torch==2.3.0+cu118 --extra-index-url https://download.pytorch.org/whl/cu118
and
pip install torchaudio==2.3.0+cu118 --extra-index-url https://download.pytorch.org/whl/cu118
Replacing cu118 with newer cu121 or even 129 seems to give proper working URL too.
Maybe replacing 2.3.0 with 2.3.1 will work too.
*a2) Alternatively, you can try to install it from here from wheels by the following command:
“pip install SomePackage-1.0-py2.py3-none-any.whl” - providing full path with the file name should do the trick. Just for the location with spaces, you also need " ".
On GTX 1660 and Turing GPUs, you might seek for e.g. cu121/torch-2.3.1" and those various CP wheels (there are no newer versions).
JFYI, the official PyTorch page: https://pytorch.org/get-started/previous-versions/
lacks links for CUDA 10 compatible versions for older GPUs other than v1.12.1 (which is pretty old, and might be a bit slower if even compatible at all), so the only way to install newer versions for CUDA 10 is the --extra-index-url trick, as executing normally “pip install torch==2.3.0+cu118” will end up with the version not found error.
*b) You might still have SageAttention not found error. Perform the following:
“Had to replace cufft64_10.dll from C:\Users\user\AppData\Local\Programs\Python\Python313\Lib\site-packages\torch\lib
by the one from C:\Program Files\NVIDIA GPU Computing Toolkit\CUDA\v10.0\bin”
It is even compatible with the newest Torch 2.8.0 (if you followed the instruction to fix the dict issue above) if you grab that apparently “ fixed version of cufft64_10.dll from CUDA v10.0” - dca
“I guess it's possible to use it with the Colab inference custom model one, you run install cell, install the neuralop thing with "!pip" in a cell code (on Colab, system command needs the "!" before them), then edit the existing Roformer code accordingly to unwa's guidelines on his repo” - jarredou Tried, doesn’t work.
If you want to train with FNO1d or Conformer, you might check out this repo.
- Turns out Duality model is very good for pops and clicks of 45 RPM vinyl, moving them to instrumental stem (bratmix)
- Google broke installing dependencies in many Colabs.
For the inference Colab by jarredou, see here for troubleshooting (pushed the changes - not tested, might be useful for other Colabs; fixed).
In the case of AudioSR, use huggingface.
- NVIDIA released their own audio upscaler, with also an ability of inpainting (so it can fill short silences between damaged segments of audio).
https://github.com/NVIDIA/diffusion-audio-restoration
But maybe don’t try it out the upscaler just yet, as the code is currently so messy and difficult to deploy e.g. on MacOS, that it took 9 hours for two of our users, and they still didn't succeed even with help of AI chats. And to make it work on Colab, the code needs to be completely rewritten, jarredou says.
- mesk released a beta version of his metal Mel-Roformer fine-tune instrumental model called “Rifforge” focused more on bleedless.
“training is still in progress, that's why it's a beta test of the model; It should work fine for a lot of things, but it HAS quirks on some tracks + to me there's some vocal stuff still audible on some tracks, I'm mostly trying to get feedback on how I could improve it” known issues.
https://drive.proton.me/urls/5XM3PR1M7G#F3UhCU8RDGhX
- Custom model import Colab might have currently some issues with the model above. Probably, using that old version will work (at least locally).
"My old MSST repo I'm using, but I removed all the training stuff
https://drive.proton.me/urls/P530GFQR4W#VCAsF0E1TPje
pip install -r requirements.txt (u gotta have Python and PyTorch installed as well) for the script to work.
You just gotta put all the tracks you want to test on in the "tracks" folder then double-click on "inference.bat" to run the inference script
it's like if you were to type in the command in cmd, but it's simpler, and I'm lazy" - mesk
- Shared bias added during weight conversion was removed from the SW model, making it compatible with UVR and normal MSST repo code (it was just a leftover not doing anything, just zeroes). Also, delete the shared bias line from the yaml.
Also, it was possible to trim the model size to have only vocals (although it probably can be achievable quicker in the config). mask_estimators.0 is responsible for vocals (each mask estimator is responsible for the other stem).
- The new violin model on MVSEP sometimes does better than the strings model for strings (dynamic64)
- Aufr33’s Mel-Roformer Denoise average variant (link | yaml | Colab) can be also used for crowd removal (Gabox)
- (MVSEP) “I released a new MVSep Violin (violin, other). It’s based on BS-Roformer model with SDR: 7.29 for violin on my internal validation.
Link: https://mvsep.com/home?sep_type=65
Example: https://mvsep.com/result/20250809120109-f0bb276157-mixture.wav”- ZFTurbo
“I've only played around with it a little bit, but it can even separate violin quartets from cellos, so cool.” - smilewasfound
“Very neat model. (...) Sometimes the model does seem to pick up more than just violins imo, but yeah for separating high strings in particular it is really cool.” - Musicalman
- MVSEP now has also official YouTube channel:
https://www.youtube.com/@MVSEP
- Issues with https://huggingface.co/spaces/TheStinger/UVR5_UI have been fixed.
Mirror is still functional: https://huggingface.co/spaces/qtzmusic/UVR5_UI
- Unwa BS-Roformer Resurrection instrumental model added on MVSEP and on uvronline with these links for free/premium accounts.
- Gabox released experimental voc_fv6 model | yaml
“Sounds like b5e with vocal enhancer. Needs more training, some instruments are confused as vocals” - Gabox. “fv6 = fv4 but with better background vocal capture” - neoculture
bleedless: 26.61 | fullness: 24.93 | SDR: 10.64
For comparison:
SCNet XL very high fuillness on MVSEP has followin metrics:
Vocals bleedless: 25.30, fullness: 23.50, SDR: 10.40
- yt-dlp and their frontends like cobalt.tools are currently defunct. It might affect some Colabs YT downloading features, although JDownloader 2 still works.
- The below model added on x-minus/uvronline
https://uvronline.app/ai?discordtest - free accounts
https://uvronline.app/ai?hp&test - premium accounts
- Unwa released a new BS-Roformer Resurrection instrumental model | yaml | Colab
SDR: 17.25, bleedless: 40.14, fullness: 34.93
Compatible with UVR (model type v1). “Fast model to inference (204 MB only)”.
“One of my favorite fullness inst models ATM. Sounds like v1e to me, but cleaner. Especially with guitar/piano where v1e tended to add more phase distortion, I guess that's what you'd call it lol. This model preserves their purity better IMO” - Musicalman
“the way it sounds, is indeed the best fullness model, it's like between v1e and v1e+, so not so noisy and full enough, though it creates problems with instruments gone in the instrumental sadly, but apparently it seems Roformer inst models will always have problems with instruments it seems, seems like a rule. (...) Instrument preservation (...) is between v1e and v1e+ (...) Fixes crossbleeding of vocals in instrumental in a lot of songs, compared to previous models (...) No robotic voice bug at silent instrumental moments” - dca100fb8
“Some songs leaves vocal residue. It is heard little but felt” - Fabio
“Almost loses some sounds that v1e+ picks up just fine” - neoculture
Mushes some synths a bit in e.g. trap/drill tune compared to inst Mel-Roformers like INSTV7/Becruily/FVX/inst3, but the residues/vocal shells are a bit quieter, although the clarity is also decreased a bit. Kind of a trade.
So far, none models work for phase fixer/swapper besides 1296/1297 by viperx and unwa BS Large V1 to alleviate the remaining noise. ~ dca. SW model not tested.
Less crossbleeding than paid Dango 11.
- Gabox released a bunch of new models:
a) Gabox Inst_ExperimentalV1 model | yaml
b) Gabox Kar v2 Mel-Roformer | model | yaml
SDR is very similar with the v1 Gabox model: 9.7699 vs 9.7661.
Lead:
bleedless: 27.58 vs 28.18, fullness: 15.24 vs 14.79
Back-instrum:
bleedless: 50.67 vs 50.74, fullness: 32.46 vs 32.84
(but you’ll most likely get better results with Gabox denoise/debleed Mel-Roformer model instead ~Gabox, but it can’t remove vocal residues
c) Gabox Lead Vocal De-Reverb Mel-Roformer | DL | config | Colab
“just use it on the mixture” - Gabox, “sounds great” - Rage313
Sometimes removes back vocals, especially if they're panned to the sides.
“(...) also a vocal/inst separator. Dry vocals go in vocal stem, everything else goes to reverb. Don't think anvuew's models do that.
I might still preprocess with vocal isolation before dereverb. But only really worth it if you're after high fullness vocals.” - Musicalman
- Issues with https://huggingface.co/spaces/TheStinger/UVR5_UI occur.
“We have a problem with Zero GPU atm, waiting for a fix from HF staff
isn't related to the code or last commit” - Not Eddy
Meanwhile, you can use: https://huggingface.co/spaces/qtzmusic/UVR5_UI
- (MVSEP) Now 32-bit float for WAV will be used only if gain level falls outside 1.0 range to prevent clipping, otherwise 16 bit PCM will be used, when it won't occur. If you really need it anyway, 32-bit float output for all files unconditionally is available for paid users.
If you have troubles with nulling due to the new changes in free version, consider decreasing volume of your mixtures by e.g. 3-5dB, and you won’t be affected, although it might slightly affect separation results.
Also, FLAC now uses 16-bit instead of 24-bit.
- (MVSEP) “Sometimes we have complaints on speed from different parts of the world. The best way is to use VPN to solve them.” - ZFTurbo
- Gabox’ voc_Fv5, Inst_GaboxFv7z, Unwa’s voc Resurrection, voc_gabox2 and the new jarredou’s drumsep 5 stem added to the inference Colab
(Resurrection [now with the config] and Fv7z fixed). Also, previously released anvuew’s mono dereverb added.
- If you feel overwhelmed by this GDoc’s list, isling released his own, shorter version with recommended models for audio separation - click.
- Also, here you’ll find an excerpt of the current document with only models and their links, if you find it hard to navigate through the whole document (edit. 24.07.25)
- (MVSEP) BS-Roformer 2025.06 described previously below received two updates (11.81 -> 11.86 and 11.86>11.89) and it has been changed to:
BS-Roformer 2025.07. Full metrics.
“All Ensembles and models where this model is involved improved a little bit too.” - ZFTurbo
MVSEP Multichannel BS feature started using 11.81 model at some point, now sure if now uses 11.89.
If you tried achieving results any similar to BS-Roformer 2025.07, you could potentially try out splifft or its Colab. If you fail to use Spliff check the model before conversion on MSST repo (more)
- (MVSEP) Gabox INSTV7 instrumental model added
- (MVSEP) MelBand Karaoke (lead/back vocals) Gabox model added (SDR: 9.67)
- Fused model of Gabox and Aufr33/viperx weights 0.5 + 0.5 added (SDR: 9.85)
It gives maybe only slightly worse results than normal ensembling, but with separation time of just one model “it doesn't have the same quality and definition as Gabox Karaoke, fused doesn't separate well.” - Billie O’Connell.
You can perform fusion of models using ZFTurbo script(src) or by Sucial script (they’re similar if not the same). “I think the models need to have at least the same dim and depth but I'm not sure about that” - mesk.
Despite the higher SDR, the fusion model seems to confuse lead/back vocals more.
- The same goes to public Karaoke fusion models released by Gonzaluigi here
- Gabox released Mel-Roformer voc_gabox2 vocal model | yaml | Colab
Vocal bleedless: 33.13, fullness: 18.98, SDR: 10.98
- Unwa released a BS-Roformer vocal model called "Resurrection" | yaml which shares some similarities with the SW model (might be a retrain). The default chunk_size is pretty big, so if you run out of memory, decrease it to e.g. 523776.
Vocal bleedless: 39.99, fullness: 15.14, SDR: 11.34.
"Omg, this model is doing a really good job at capturing backing vocals (...)
Honestly, it sounds a bit muddy, and there's some instrumental bleeding into the vocal stems" neoculture
Not so good for speech denoising unlike some other models (Musicalman).
- mesk’s training guide updated and link changed
- (MVSEP) New All-in and 5 stem ensembles have been added for paid users
- AudioSR WebUI Colab by Sir Joseph got fixed
- MVSep Ensemble 11.93 (vocals, instrum) (2025.06.28) added.
Eventually surpassed sami-bytedance-v.1.1 on the multisong dataset SDR-wise.
Instrumental bleedless: 47.65, fullness: 28.76, SDR: 18.24
Vocal bleedless: 36.30, fullness: 17.73, SDR: 11.93
- corrected typos in the metrics (thx yakomotoo)
- (MVSEP) MultiChannel now uses 11.81 BS Roformer model.
- (MVSEP) New BS Roformer model is now available on site - it’s called 2025.06 (don’t confuse it with SW).
Vocals bleedless: 48.59, fullness: 27.85, SDR: 11.82
Instrumental bleedless: 37.83, fullness: 17.30, SDR: 18.12
“It has +0.5 SDR to the previous best [24.08] model. We reached ByteDance's best model quality [only 0.1 SDR difference). It is also TOP1 on the Synth dataset. It's balanced between both [instrumental and vocals]. I used metal dataset during training as well"
Compared to previous models, picks up backing vocals and vocal chops greatly where 6X struggles, and fixes crossbleeding and reverbs where in some songs previous models struggled before. Sometimes you might still get better results with Beta 6X or voc_fv4 (depending on a song). “Very similar to SCNet very high fullness without the crazy noise” - dynamic64, “handles speech very well. Most models get confused by stuff like birds churping (they put it in the vocal stem), but this model keeps them out of the vocal stem way more than most. I love it!”
“not a fan of the inst result. I feel like unwa and gabox sound better despite being less accurate” - dynamic64. Might be better than Fv7n “I think gabox tends to sound better but the new BS-Roformer is more accurate” dynamic64, “instrumentals are muddy” - santilli_,
“I think the Gabox [fv7n] model sounded more crispier than BS” - REYYY. “[voc_]fv4 sounds better” - neoculture, “instrumentals sound very good” - GameAgainPL.
“it did things i never thought it could before” “this model is insane wtf (...) never seen a model accurately do the ayahuasca experience before” - mesk.
“the first model to not produce vocal bleed in instrumental for "Supersonic" by Jamiroquai (not even Dango does it). It is also the case with "Samsam (Chanson du générique)" and "Porcelain" by Moby.” and "In the Air Tonight" by Phil Collins, also “removes very most of Daft Punk vocoder vocals" - dca. “my new favorite for vocals. It sounds fantastic” - dynamics64. “for the first time ever it managed to remove the reverb from one specific song. it is not perfect, but still much better than previous attempts” - santilli_
“It even seems to handle speech very well. Most models get confused by stuff like birds churping (they put it in the vocal stem), but this model keeps them out of the vocal stem way more than most. I love it!”. “sometimes 6x is better sometimes bs is better” - isling “for me it's picked up a lot that 6x hadn't for backing vocals
- Using this repo, you can convert Mel-Roformers, HTDemucs and Apollo models to OpenVINO (so to onnx)
- Lew, if you read it, some guy wants to add your Apollo uni model into a plugin for OpenVINO and Intel’s HF, but the model lacks an open source licence. If you could re-release it with the proper licence, it would be appreciated. More
- (x-minus/uvronline) “I added two new models to remove vocals and hid a few old ones.
So there are now only three main models in the menu for different purposes:
Mel-RoFormer by Gabox Fv7z - best bleedless, good fullness, almost noiseless
Mel-RoFormer by unwa v1e+ - best fullness, average bleedless
Mel-RoFormer unwa big beta6x - best vocals
Older models are still available at the link:
https://uvronline.app/ai?hp&test (premium)
https://uvronline.app/ai?test (free)” - Aufr33
“Oh! Lead vocal panning has been added for Mel Kar Old! (...)
Along with MDX Kar old and UVR Kar old to the test page!!” - dca
- Gabox released a new experimental Karaoke model. It’s one stem target so keep extract_instrumental enabled for the rest stem.
“really hard to tell the difference between this and becruily's karaoke model” minus the latter has more target stems.
- jarredou released his new MDX23C drumsep 5 stem model, which is public for everyone to download. All SDR metrics are better than the previous model (“on kick/snare/toms it's around +2 SDR better than previous version”):
SDR: kick: 16.66, snare: 11.54, toms 12.34, hihat: 4.04, cymbals: 6.36 (all metrics).
Metric fullness for snare: 25.0361, bleedless for hh: 12.3470, log_wmse for snare: 13.8959
“Quite cleaner than the previous one”, “it's more on the fullness side than bleedless”,
From all the metrics, only bleedless for snare is worse than in the previous model:
26.8420 vs 30.4149 and indeed “snare has a bit of bleed sometimes” - isling, “as well as cymbals bleed in hi hat track, but the stems sound clean” - dca.
“a lot noisier than other drumpsep models, but that's not necessarily a bad thing.”
“Surprisingly, it's the 2nd best model for hi hat and 2nd best model for cymbals on mvsep leaderboard. It's a bit biased because ZF's top mel model is 4 stem only.”
For comparison, metrics of the old 6 stem jarredou/Aufr33 MDX23C model
(which has cymbals divided into ride and crash which are not evaluated):
SDR: kick: 14.55, snare: 9.79, toms: 10.64, hihat: 3.20, cymbals: 6.08
Metric fullness for snare: 25.0361, bleedless for hh: 10.2765, log_wmse for snare: 12.4258
The model was trained with a lightweight config to train on a subpar T4 GPU on free Colabs and 10 accounts (“CRAAZY fast” for inferencing). The metrics do not surpass exclusive drumsep Mel-Roformer and SCNet models on MVSEP, but at least you can use this one locally.
“Most of the issues with my model are already known issues with mdx23c arch, it's bleedy and has band splitting artifacts. Like I said a few days ago, if I would have to redo it now, it would have probably gone with SCNet Masked. It's using 4x times lower n_fft resolution than InstVocHQ while using 2 times longer chunk_size (and with MDX23C, whatever number of stems, it's the same inference speed). A bit like the fruit’s model is doing”.
Trained on 511 tracks, MVSEP models were trained on almost the same dataset.
Maybe if we separate just snare with the old MDX23C model from an already separated drums stem, and mix/invert to get the rest, then pass it through the new model, the bleed would be gone.
Remember that you need already separated drums in one track to use this model effectively.
About used dataset: “It was around 2/3 acoustic drums and 1/3 electro drums dataset at start of training, I've added more electro drums at end of training to balance it a bit more.” - jarredou
- septcoco released macvsep which is “macOS client for the Mvsep music separation API”
- Added Clear Voice in speech separation
- Gabox released Inst_GaboxFv7z Mel Roformer | yaml
Inst. fullness: 29.38, bleedless: 44.95
“Focusing on the less amount of noise keeping fullness”
“The results were similar to INSTV7 but with less noise” - neoculture
Metrically better bleedless than Unwa v2 (although it’s even more muddy), for comparison:
Fullness: 31.85, bleedless: 41.73
- (MVSEP) “I added a new SCNet vocal model. It's called SCNet XL IHF. It has a better SDR than previous versions. Very close to Roformers now".
Vocal bleedless is the best among all SCNet variants on MVSEP. Metrics.
IHF stands for “Improved high frequencies”.
Vocal bleedless 28.31, fullness 17.98
“certainly sounds better than classic SCNet XL (...) less crossbleeding of vocals in instrumental so far, and handle complex vocals better (...) problems with instruments, compared to high fullness one. XL high fullness remain the one without too many instruments cut”, but some difficult songs used with previous models can yield better results - dca
- Great news! MVSEP now allows sorting scores on the Multisong Leaderboard by SDR, fullness, bleedless, aura_stft, aura_mrstft, log_wmse, l1_freq, si_sdr.
Be aware that Gabox (and probably sometimes becruily) used to give funny names to their evaluations, so finding proper model names on the leaderboard is sometimes impossible. But I’ve tracked down all possible models with their metrics and proper names in the instrumentals and vocal models section, so no worries.
Also, metrics beside SDR are not available for old evaluations where they weren’t listed in the model details yet. You can find more info about bleedless/fullness metrics here.
Log WMSE metric is good “at least for drums or anything rich in low frequency content” - jarredou
- Our server members send their warm regards to A5 whose account disappeared for the ~5th time :) And later reappeared weirdly mutated.
- “Dango launched their new instrumental model
https://tuanziai.com/en-US/blog/684841907c8c85686c1b3da6” It’s version 11.
“there is no opportunity to try at least 3 complete tracks for free.”
Some crossbleeding issues from v10 are still present, plus some songs are even getting worse results than in v10. You might want to use v1e (with phase fix) + Becruily vocal model (Max Spec) instead, although some people might still like Dango anyway.
“Some tracks are fuller than Gabox v8”. Conservative mode is less full than V1e.
“They have a tool called "edit & improve" [or “Advanced Repair tool”] that lets you use 'Conservative mode' for some of more complex parts of a song and 'Smart mode' for other parts. I find that way more convenient than processing the entire track in 'Conservative' mode.”
They plan to release a karaoke model in two months.
- Gabox released a “small” version of Mel instrumental model for faster inference | yaml
Be aware that it can have some audible faint constant residues.
- ZFTurbo: “I added BS Roformer SW to "MVSep Piano", "MVSep Guitar", "MVSep Bass", "MVSep Drums" algorithms. For Bass and Drums available new Ensembles.”
“drum SDR jumped by .6 on the ensemble! Atho fullness took a hit” - heuhew
“Same with bass, sdr leaped but fullness shot down 4 points” - dynamic64
- BS-Roformer SW 6 stem model replaced the old one.
Iirc, the model didn’t change, just the inference code. SW stands for “shared weights”
“I got better drums/bass separation with that model than with any others when input is some live/rehearsal recordings with shitty sound”
Also, there’s better SDR and fullness for instrumentals when you invert vocals against mixture instead of mixing down drums/bass/other stems.
- undef13 released these “bs-roformer weights stored in fp16 precision, half the size of frazer's initial version. quality is the exact same as the fp32 version”.
- First vocal retrain was published shortly after a day “May not perform very well” (at least for now)
If you were to fine-tune it “this model generalizes like crazy (...) hasn't failed yet to confuse instruments and just chews through whatever you put through it (i ignore the overall mudiness)” you can retrain it to just being inst/voc model “I’m currently training it to my 2 stem (...) dataset (...) I was pleasantly surprised)” iirc on even laptop RTX 3070...
- Added on MVSEP as BS-Roformer 6 stem (no clicking issues)
- The new Logic Pro model has been reversed/cracked and shared as a standalone model for inference. Full metrics for all stems added later below. It has the best SDR on multisong dataset for all stems besides vocals (but still not bad).
It uses a BS-Reformer arch. .MIL (CoreML) model file was converted to .PT.
“The only change they made was a global parameter for bias which I've never seen before so I guess it's Apple secret sauce”. No quantization was used “they had a shared bias across QKV and the out_proj”.
Since then, weight compatible with UVR with deleted shared bias was shared (there were actually only zeroes). Also, with the mask estimator method, just one stem file can be extracted out of the full weight. Vocals only were shared, but config for UVR will rather require some tweaks.
A bit scared to share it, but seek and ye shall find.
“It is wonderful to achieve such results with dim 256. It seems that what was still needed was depth.”
Usage for the old model-specific inference code before shared bias was deleted:
python inference.py --audio_path="./sample.flac"
For: ModuleNotFoundError: No module named 'hyper_connections'
Run: pip install hyper_connections
“looks like chunks aren't overlapping? Getting clicks in output.”
“A very small edit, line 13:
parser.add_argument('--chunk_size', type=int, default=588800)
- this produces 99% identical results with the DAW.
Previous 117760 chunk size was adding clicks and was lower quality in general.”
Still, the code doesn’t use overlap, and it will result in click, just less than before.
Also, you can run out of memory with 588800 with 5GB VRAM free.
882000 was tested to have the biggest SDR in that model (not lower or higher).
On a CPU without an Nvidia GPU it will probably be long.
The inference script and model probably still needs the validation to ensure the metrics are the same with the validation made from DAW lately, but it’s rather the same (at least other inference code got 0.03 SDR difference or same results based on the same converted weights).
To use the old model version with ZFTurbo MSST repo:
“You need to replace bs_roformer.py in the repo with file from the archive (...) and change line 8 to:
from models.bs_roformer.attend import Attend” and then use separately shared config for the MSST repo and the model. Using MSST repo for inferencing fixes the clicking issue.
For “unrecognized arguments” issue, “you must put your path inside quotation marks or apostrophes”.
“If you're using the script GUI, be aware that the browser popup window when choosing checkpoint has some predefined extension and .pt is not part of it”
- Becruily guitar model added to inference Colab, bleed suppressor by unwa/97chris model fixed, denoise-debleed by Gabox added, Revive 3e fixed, Revive 2 added
- Gabox released instv7plus bleedless model (experimental)
fullness: 29.83, bleedless: 39.36, SDR 16.51
- And Inst_FV8b
fullness: 35.05, bleedless: 36.90, SDR 16.59
“Very clean” although muddier than V1E+.
- wesleyr36/Dry Paint Dealer Undr HTDemucs Phantom Center model was added to the inference Colab
- (Unwa) “After a long time, I'm uploading a vocal model specialized in fullness.
Revive 3e is the opposite of version 2 — it pushes fullness to the extreme.
Also, the training dataset was provided by Aufr33. Many thanks for that.”
bs_roformer_revive3e | config | Colab (should be fixed now)
Voc. SDR: 10.98, fullness: 21.43, bleedless: 30.51
- Logic Pro updated their stem separation feature, which now incorporates guitar
Overall, it’s “surprisingly good” - dynamic64. And a piano separator was also added to it. More
“Guitar & Piano separation seems to be really on point. So far it separated super well, also didn’t confuse organs for guitars and certain piano sounds as well.” - Tobias51
“guitar model sounds better than demucs, mvsep, and moises” - Sausum
“it's not a fullness emphasis or anything, but it's shockingly good at understanding different types of instruments and keeping them consistent sounding” - becruily
You don’t need to process L and R for bleeding across channels like in other models, there isn’t any in this one - A5
Full evaluation on multisong dataset (besides instrumental):
SDR piano 7.79, bleedless 31.96, fullness 14.42
SDR other 19.90, bleedless 58.68, fullness 49.85
SDR guitar 9.00, bleedless 31.54, fullness 15.95
SDR other 15.94, bleedless 49.36, fullness 31.57
SDR drums 14.05 (although lower fullness than MVSep SCNet XL drums 14.26 vs 21.21),
SDR bass 14.57 (-||-), other 8.66, vocals 11.27 (only that is not SOTA)
MVSep Piano Ensemble (SCNet + Mel) has only other fullness higher: 56.96 (click)
- Since 23.05.25 jarredou (Discord: rigo2) and dca100fb8 (Discord) also have writing privileges to this document. You can find it mirrored to this date here in docx, pdf and html.
- Becruily released Melband guitar model | Colab
“Not SOTA, but much more efficient and comparable to existing guitar models, and for some songs it might work better because it picks up more guitars (though it can also pick some other instruments).
For better results you might try first removing vocals.”
- (MVSEP) “We added a new GUI example to work with the MVSep API. Now it allows to use multiple files and multiple algorithms at once.
It exists as standalone .exe file, so it doesn't require python installation
Repository: https://github.com/ZFTurbo/MVSep-API-Examples
Exe for Windows: https://github.com/ZFTurbo/MVSep-API-Examples/raw/refs/heads/main/python_example5_gui/mvsep_client_gui_win.exe” - ZFTurbo
TL;DR “You can process a song with multiple models, and process multiple songs”
- Sucial released v1/2 de-breath VR models:
https://huggingface.co/Sucial/De-Breathe-Models/tree/main
Alternatively, for this purpose you can also try out free/abandonware:
https://archive.org/details/accusonus-era-bundle-v-6.2.00
- (x-minus/uvronline) Aufr33 added new Lead and Backing vocal separator.
It uses big beta 5e model as preprocessor for becruily Mel Karaoke model “In fact, the big beta 5e model is run after becruily Mel Karaoke” Aufr33 (so you don’t need the additional step to use this separator), plus it also allows controlling option for lead vocal panning like for BVE v2 (it’s to “to "tell" the AI where the main vocals are located (how they are mixed).”. Becruily’s model “doesn't even need Lead vocal panning a lot of the time, [the] ability to recognize what is LV and what is BV [is] impressive” - dca).
The difference from using single becruily Kar model (without preprocessor) is that, here, “you get the third track, backing vocals.”.
“The new separator is available in the free version, however, due to its resource intensity, only the first minute of the song will be processed.” if you don’t have premium.
Becruily:
“Probably too resource-intensive, but you could try adding demudders to each step
1) karaoke model + demudding
2) separate vocals of bgv + demudidng
But not sure how much noise this will bring
(Or even a 50:50 ensemble with BVE OG)”
- Unwa released Revive 2 variant of his BS-Roformer fine-tune of viperx 1297 model
Voc. bleedless: 40.07, fullness: 15.13, SDR: 10.97
“has a Bleedless score that surpasses the FT2 Bleedless” and fullness lower by 0.64.
“can keep the string well” better than viperx 1297 (...) in my country they have some song with Ethnic instruments. Only 1297 and Revive2 can keep them in Instrumental while other model notice them as Vocal” ~daylight
“it does capture more than viperx's” - mesk
It’s depth 12 and dim 512, so the inference is much slower than with some newer Mel-Roformers like voc_fv4 (even two times), with the exception of Mel 1143 which is as slow as BS 1297 (thx dca, neoculture).
- BS-Roformer Revive unwa’s vocal model (viperx 1297 model fine-tuned) was released.
Voc. bleedless: 38.80, fullness: 15.48, SDR: 11.03
“Less instrument bleed in vocal track compared to BS 1296/1297” but it still has many issues, “has fewer problems with instruments bleeding it seems compared to Mel. (...) 1297 had very few instrument bleeding in vocal, and that Revive model is even better at this
(...). Works great as a phase fixer reference to remove Mel Roformer inst models noise” it doesn’t seem to remove instruments like FT3 Preview for phase fixing (thx dca100fb8)
Added to phase fixer Colab.
- Inst_GaboxFv8 model | yaml | Colab checkpoint has been updated, metrics could have changed, but most of the model qualities might remain similar
- (MVSEP) “I added new Drumsep MelBand Roformer (4 stems) model on MVSep (old one was removed). It gives the best metrics with big gap for kick, snare and cymbals.” - ZFTurbo
(metrics; only toms are worse SDR-wise vs previous SCNet Drumsep models)
- Gabox released voc_fv5 vocal model | yaml
voc bleedless: 29.50, fullness: 20.67, SDR: 10.56
“fv5 sounds a bit fuller than fv4, but the vocal chops end up in the vocal stem. In my opinion, fv4 is better for removing vocal chops from the vocal stem” - neoculture. Examples
“v5 is slightly fuller, v4 is less full but also slightly more careful about what it considers as vocals. I think b5e is the fullest overall, but it's a bit much sometimes. Pretty sure the gabox models are a little more accurate with vocal/instrument detection.” Musicalman
Passes the Gregory Brothers - Dudes a Beast test (before - trumpets in vocal stem at 0:51; unwa’s beta4 and inst v1e tested) - maxi74x1
- Some of our less active users have been accidentally kicked out of our Discord server during some administrative tasks. You’re free to rejoin using this invite link (unless you were banned before in some other unrelated event).
- Dry Paint Dealer Undr (a.k.a. wesley36) released new Phantom Centre Models:
HTDemucs Similarity/Phantom Centre Extraction model:
https://drive.google.com/drive/folders/10PRuNxAc_VOcdZLHxawAfEdPCO6bYli3?usp=sharing (it tends to be more “correct” in center extraction than the last MDX23C model)
The Demucs model won’t work with UVR giving bag_num error even with the yaml prepared in the same way as for Inagoy Drumsep and after renaming ckpt to th (it’s probably because it needs ZFTurbo inference code).
SCNet Similarity/Phantom Centre Extraction model:
https://drive.google.com/drive/folders/1CM0uKDf60vhYyYOCg2G1Ft4aAiK1sLwZ?usp=sharing
And also, difference/Side Extraction model based on SCNet arch was released:
https://drive.google.com/drive/folders/1ZSUw6ZuhJusv7HE5eMa-MORKA0XbSEht?usp=sharing
- Aufr33 released his UVR Backing Vocals Extractor v2 model, previously available only on x-minus/uvronline (VR arch).
“Note that this model should be used with a rebalanced mix.
The recommended music level is no more than 25% or -12 dB.
If you use this model in your project, please credit me.”
Should work in UVR. Just place the model file in Ultimate Vocal Remover\models\VR_Models and config file in lib_v5\vr_network\modelparams. Then pick “4band_v4_ms_fullband.json” when asked to recognize the model (it has the same checksum as in lib_v5\vr_network\modelparams folder if it’s there already). Also, I think it's not VR 5.1 model. And it was used with vocal model as preprocessor.
More about its usage in Karaoke section (scroll down a bit).
- squid.wtf doesn't work anymore “it just downloads 30 seconds of a song, just a random 30 second snippet” lucida works.
- USS-Bytedance Colab has been fixed (Python “No such file or directory” fix) - thx epiphery.
https://colab.research.google.com/drive/1rfl0YJt7cwxdT_pQlgobJNuX3fANyYmx?usp=sharing
- (MVSEP) “I added a new MVSep Saxophone (saxophone, other) model. It has 3 versions:
SCNet XL (SDR saxophone: 6.15, other: 18.87)
MelBand Roformer (SDR saxophone: 6.97, other 19.70)
Ensemble Mel + SCNet (SDR saxophone: 7.13, other 19.77)” ZFTurbo
“SCNet XL take[s] wurlitzer as sax tho. Mel Rofo one (...) didn't” - dca
- (x-minus) Server code updated [might fix the issue with bleeding at first seconds in e.g. Mel Decrowd; edit. it didn’t]
Added Lead vocal panning setting for Mel-RoFormer Kar by becruily model.
[It’s] “to "tell" the AI where the main vocals are located (how they are mixed).
Added Demudder for the Mel-RoFormer Kar by becruily model.” - Aufr33
“doesn't even need Lead vocal panning a lot of the time, [the] ability to recognize what is LV and what is BV [is] impressive” - dca
- Anjok released a new UVR Roformer patch #15 fixing CUDA for RTX 5000 Series GPUs and Windows users (it’s based on CUDA 12.6 and newer PyTorch). It might not be backward compatible with older GPUs, so be aware (src).
Download
- (MVSEP) “I added becruily Karaoke model. It's available as option in MelBand Karaoke (lead/back vocals) algorithm.” ZFTurbo
- (MVSEP) Since at least February there's a normalization for all input unless WAV is chosen as output format.
Sometimesi it can be "annoying when you have to combine the outputs later".
“No, if you turn off normalization, FLAC will cut all above 1.0
And if it was normalized, it means you had these values.”
FLAC doesn’t support 32-bit float, it’s 32 int, so normalization is still needed.”
So if your stems don’t invert correctly, just use WAV output format - it's 32-bit float.
- Audioshake now have strings model
- Fast Separation Colab by Sir Joseph has been updated with the following models:
MelBand Roformers: FT 3 by unwa, Karaoke by becruily, FVX by Gabox, INSTV8N by Gabox, INSTV8 by Gabox, INSTV7N by Gabox, Instrumental Bleedless V3 by Gabox, Inst V1 (E) Plus by Unwa, Inst V1 Plus by Unwa
- (stephanie/UVR) “Those of you on Linux running the current roformer_add+directml branch that cant get becruily's karaoke model working due to the same error:
it seems editing line 790 in separate.py setting the keyword argument strict to False when calling load_state_dict seems to make the karaoke model load and infer properly, so i think it will work
model.load_state_dict(checkpoint, strict=False)
I don't know if this is a robust workaround, but I haven't observed anything behaving differently than it should yet, so if you want to give it a shot I think it will work
TL;DR change line 790 in separate.py to the codeblock and then run again and karaoke model should work”
- Aname’s Mel-Roformer 4 stems Large added to inference Colab
- Apollo Lew Uni model can be also used as denoiser.
It tends to smooth out some noise in higher frequencies, making the spectrum more even there, smoothing out the sound in general (example). “tends to "Clean" audio noise and flatten the sound a bit.” - CC Karaoke
More about the model and its usage - click.
- Becruily’s Mel-Roformer Karaoke model added on x-minus/uvronline under “Keep backing vocals” option and in the inference Colab
Most likely, you’ll have “”’norm’”” AttributeError when trying out that model in UVR. Read here for troubleshooting. Use melband-roformer model type, not v2.
Make sure you use the latest UVR Roformer patch - older patches like #2 will show RuntimeError about layers.
- (becruily) “I'm releasing my first karaoke model.
It's a dual model trained for both vocals and instrumental. It sounds fuller + understands better what is lead and background vocal, and to me, it is better than any other karaoke model.”
“Compared to Aufr33’s Melband model, it can achieve e.g. cleaner pronunciation in some songs (examples) - neoculture “It is the best available, better than Mel Kar, UVR BVE v2, lalal.ai, Dango...” - dca “This sounds amazing” - Rege 313 “It performs very well with male/female duets, nice work” - Gabox
“Important note: This is not a duet or male/female model. If 2 singers are singing simultaneously + background vocals, it will count both singers as lead vocals. The model strictly keeps only actual background vocals. The same goes for "adlibs" such as high notes or other overlapping lead vocals.
The model is not foolproof. Some songs might not sound that much improved compared to others. It's very hard to find a dataset for this kind of task.
Tip: For even better results, first extract the vocals with a fullness model (like mine) and combine the results with a fullness instrumental model.” becruily
The model outputs 2 stems like duality models, so you might end up with three outputs if you check the option to invert stem - don’t use it, it will rather have worse quality than what the model outputs.
- (MVSEP) “I added 2 more models for DrumSep based on MelBand Roformer architecture.”
a) 4 stems (kick, snare, toms, cymbals) - average SDR of hihat ride, crash is 11,52 (but in one stem) and so far it’s the best SDR out of all models (even vs the previous ensemble consisting of three MDX23C and SCNet models).
b) 6 stems (kick, snare, toms, hihat, ride, crash) - average SDR of hihat ride, crash is 8.18 (but from separated stems), while
The snare in a) has the best SDR out of all available models.
Kick and toms are still the best SDR-wise in the previous 3x MDX23C and SCNet ensemble (new ensemble with these new Mel-Roformers so far)
- The new models “are very great for ride/crash/hh. And overall they have the best metrics almost for all stems.” - ZFTurbo
- Aname released two 4 stems Mel-Roformer models:
https://huggingface.co/Aname-Tommy/melbandroformer4stems/tree/main
a) Large (4GB) SDR drums: 9.72, bass: 9.40, other: 5.11, vocals 8.65 (multisong dataset)
b) XL (7GB) SDR drums: 9.83, bass: 9.37, other: 5.31, vocals 8.57 (multisong dataset)
The latter doesn’t work in the custom model import Colab with at least the default chunk_size, and works slow on e.g. 3060 (?12GB). Both models were trained with chunks set to 15 seconds (chunk_size = 661500).
“I tried a song on 4070 Super it took like 6 mins on XL 4 stems compared to 30 seconds on Large 4 stems” On 3060 XL is very slow.
Despite lower AVG SDR on musdb18 dataset vs demucs_ft (8.54 vs 9), it seems to outperform that model (SDR is only better in other stem), public SCNet, SCNetXL, BS-Roformer have better metrics (still musdb18 dataset, not multisong on MVSEP)
“Drums are sounding really good in particular, tested a couple songs with the large model after using unwa's v1e+ for instrumental” “drums are absolutely the standout“
“Large works in like 99% use case” “Large split sounds amazing so far tho”
XL “result would take so much longer, but the large results sounded better imo” 5B
“The Colab is forcing a different value than the one from the config. You can try to edit the inference cell code and add 661500 as possible value and see if it goes better.
The Colab only changes chunk_size (value from GUI), batch_size (forcing =1) and overlap (value from GUI), it doesn't touch other settings from config.” - jarredou
“It may change audio setting, chunk_size=485100, n_fft=2048 will work, but it will go lower SDR maybe” while the lowest reasonable value will be rather 112455 (2,5 s).
Large model uses 7GB VRAM on Nvidia GPU in UVR with default config settings.
- Sir Joseph released SESA Fast Separation Colab based on UVR. It’s faster than the regular SESA Colab (whiuch now has “added Apollo to Auto Ensemble and fixed a few technical glitches. It’s running smoother now!”)
More changes in the fast Colab:
V1e+ and Gabox inst fv8 are missing because the model list cannot be updated in the Fast Colab yet.
“auto-ensemble feature is included here too.
Background noise suppression is a bit more polished.
You can specify unwanted stems to filter out.”
- (x-minus/uvronline) “1. A new Mel-RoFormer by unwa v1e+ model has been added. It removes vocals very gently while preserving instruments. It is recommended to use it with correct_phase post-processing.
2. Mel-RoFormer by Kim & unwa ft3 and some other models are hidden. As before, you can find them here: https://uvronline.app/ai?hp&test” - Aufr33
“The only problem is the phase correction, it still uses FT2 as a reference [for phase fixer], and FT2 cuts instruments still, so I'm waiting for FT3 release by unwa so it can be added as phase fixer reference and preserve instruments well” dca
“Results are still better with phase fixer though, right”
Make sure you’re “clicking on "Ensemble"? It should "reveal" that option” since the last website layout changes.
“phase fixer [on the site] swaps the v1e+ vocals with the ft2 vocals”
Iirc phase fixer feature requires premium.
- SESA Colab by Sir Joseph is back! The Colab link has changed - click
Apollo Integration: Added Apollo audio enhancement feature. Supports Normal and Mid/Side methods.
UI Updates: Added new Apollo settings components under the Settings tab.
Bug Fixes:
Fixed Apollo output not showing in the terminal.
Corrected "Phase Remix" and "Overlap Info" display in the UI.
Translation Updates: Added new translation keys for Apollo, removed unused keys.
Colab Support: Added 10 new languages: EN_US (English), TR_TR (Turkish), AR_SA (Arabic), RU_RU (Russian), ES_ES (Spanish), DE_DE (German), ZN_CN (Chinese), HI_IN (Hindi), JA_JP (Japanese), IT_IT (Italian).
and new models added
Note: Enhanced UI and processing stability.
- Gabox released Inst_GaboxFv8 model (yaml) [weight has been replaced by v2]
Inst. bleedless: 38.06, fullness: 35.57, SDR: 16.51 [outdated]
Might have some “ugly vocal residues” at times (Phil Collins - In The Air Tonight) - 00:46, 02:56 - dca.
VS v1e ”it seems to pick up some instruments better” Gabox
“a bit cleaner-sounding and has less filtering/watery artifacts.
Both models are prone to very strange vocal leakage [“especially in the chorus.”].
And because Fv8 can be so clean at times, the leakage can be fairly obvious. For now, my vote is for Fv8, but I'll still probably be switching back and forth a lot” - Musicalman
“sometimes v1e+ have vocal residues which sound like you were speaking through a fan/low quality mp3” - dca
- Added Mesk Metal Model Preview, Unwa v1+ Preview, and Unwa v1e+ Mel instrumental models and Beta6X and FT3 Preview by Unwa vocal models, and Bandit v2 multilingual model to inference Colab
- Unwa released a new V1e+ Mel-Roformer instrumental model | yaml | Colab
Inst bleedless: 36.53, fullness: 37.89, SDR: 16.65
Less noise than v1e (esp. in the lower frequencies), but it’s also less full - “somewhere between v1 and v1e.”. It has fewer problems with quiet vocals in instrumentals than the V1+, “issues with harmonica, saxophone, elec guitar and synth seem to have been fixed. Theremin and kazoo are still problematic [like] for models from MDX-Net or SCNet [archs]). Only dango seems to correctly detect kazoo as an instrument it seems” - dca, “The loss function was changed to be more fullness-oriented, and trained a further 50k steps from the v1+ test.” Unwa
“v1e keeps better instruments like trumps than v1e+
With v1e+ there is less noise, but some instruments are hidden” koseidon72
“v1e+ has a strange problem of almost vocoding the vocals and keeping them in quietly” even with phase fixer
“has some problems with cymbals bleed in vocals (not the case with other instrumental roformer models)” dca
“trained with additional phase loss which helps remove some of that metallic fullness noise, and also has higher sdr I believe” - becruily
- Unwa released V1+ Mel-Roformer instrumental model | yaml | Colab
Inst. bleedless: 38.26, fullness: 35.31, SDR: 16.72
"It is based on v1e, but the Fullness is not as high as v1e, so it is positioned as an improved version of v1." Unwa
"very nice model, the multistft noise is gone"
It's probably due to:
"Unwrapped phase loss function added" Unwa
BTW. It was already proven before, that adding artificial noise to separations was increasing fullness metric.
"Seems to have significantly less sax and harmonica bleed in vocal, which is an awesome thing (...) It still struggles with other things like FX and Kazoo." dca
"It sounds clean. The only thing [is] that some instruments are deleted, and in some tracks leaves remnants of voice in the instrumental." Fabio
"Screams are not removed from the track" Halif
Training details
"I made a small improvement to the dataset and trained about 50k steps with a batch size of 2.
8192 was added to multi_stft_resolutions_window_sizes.
As it was, the memory usage increased too much, so it was rewritten to use hop_length = 147 when window_size is 4096 or less and 441 when window_size is greater than that." Unwa
- Mesk released a preview of his instrumental model retrained from Mel Kim on metal dataset consisting of a few thousands of songs.
https://huggingface.co/meskvlla33/metal_roformer_preview/tree/main | Colab
These are not multisong metrics, but made with private dataset!
Instr bleedless: 48.81, fullness: 42.85, SDR: 13.7621
"currently restarting from scratch because I think I know what all the problematic vocal tracks were, and I removed them, we'll see if it's gonna be better"
"vocals could follow if requested.
Should work fine for all genres of metal, but doesn't work on:
- hard compressed screams
- some background vocals
- weird tracks (think Meshuggah's "The Ayahuasca Experience")
P.S: Use the training repo (MSST) if you want to [separate] with it. UVR will be abysmally slow (because of chunk_size [introduced since UVR Roformer beta #3])”
- Yusuf fixed Apollo and AudioSR WebUI Colabs and mid/side method of upscaling was added to Apollo
- Unwa released Big Beta 6X vocal model (yaml)
Vocal bleedless: 35.16, fullness: 17.77, SDR: 11.12
“it is probably the highest SDR or log wmse score in my model to date.”
Some leaks into vocal might occur.
“dim 512, depth 12.
It is the largest Mel-Band Roformer model I have ever uploaded.”
“I've added dozens of samples and songs that use a lot of them to the dataset”
- (MVSEP) “I added new Apollo model with Aura MR STFT: 22.42
It's available under "Apollo Enhancers (by JusperLee and Lew)" with option:
"Universal Super Resolution (by MVSep Team)".
It requires a hard cutoff on frequency for best experience.” - ZFTurbo
Iirc, it was trained by his student.
“It's doing well on more transient stuff like snare hits, but it seems to really struggle to actually add harmonics. Has this really weird quality of sounding high quality and low quality at the same time”
“It doesn't seem to like 8 kHz cutoff, it has generated almost nothing”
“I tried with a 10 kHz cutoff and just got quitet-ish transients”
“Lew told me the same while training his, the model would learn transients/drums but struggle with harmonics. Maybe it’s an Apollo limitation. I don’t recall if the OG model by jusper lee has this issue too, since it rarely works”
Advice
You might want to process your song even 4 times to potentially get better results.
Also, you can split mids and sides, and upscale them separately to get better results, although it’s not always better solution (spectrograms | tutorial), thx AG89.
Using e.g. MDX23C Similarity/Phantom Centre extraction model instead with 2x slowdown (to reduce smearing artefacts) gives less high-end recovery, but less noise resulting in more proper cancelling of both channels (spectrograms by AG89).
Avg ensemble will be rather diminishing returns, so consider manual weighted ensemble in DAW.
Getting rid of noise or dithering above real frequencies by making cutoff can make a night and day difference for the result (example)
Sometimes cutting off some more existing frequencies might be beneficial too (the model was trained with hard cutoff)
For noise artefacts after upscaling you can use some denoisers
- Gabox released new INSTV8N instrumental model in experimental folder (yaml)
“noticed too many vocal residues. (...) there is no noise” although N stands for noise in its name.
- Some upscaling Colabs are also affected by the last runtime changes in Colab made by Google. Maybe downgrading !pip install torch==2.5 would help.
- We're aware of the issues in some Colabs like MDX by HV (numpy errors related to its wrong version). Any fixing will be announced. Stay tuned.
- Fixed, but initialization is slow till further notice, and you need to click initialization cell second time when you’re prompted to restart environment.
- Fixed, but now you need to click the initialization cell again after Numpy has been installed (happens briefly after launching the initialization cell).
- Unwa's FT3 test vocal model added on x-minus/uvronline
“make vocals sound a bit lower at chorus compared to other parts of songs”, doesn’t happen with big beta 5e - oak
- ZFTurbo: “I added 2 new super resolution algorithms on MVSep in Experimental section:
1) AudioSR. Metrics: https://mvsep.com/quality_checker/entry/8067
2) FlashSR. Metrics: https://mvsep.com/quality_checker/entry/8071”
Be aware that both can give some errors occasionally. Some problems with mono audio were fixed already.
- Unwa released FT3 preview vocal model | yaml
Vocal bleedless: 36.11, fullness: 16.80, SDR: 11.05
“primarily aimed at reducing leakage of wind instruments to vocals.
I will upload a further fine-tuned version as FT3 in the near future.”
For now, FT2 has less leakage for some songs (maybe till the next FT will be released).
- Gabox added some new experimental instrumental models in a separate repo folder.
They are called V8, V9, V10, don’t consider them as newer/better, but forgotten to upload in the meantime.
They’re less full than V7, but have less vocal residues. Also, the results from V8 and V10 are the same (“inverted polarity between 2 results, and it's just silence”), and also for V9.
“Both remove some instruments from the music, like V7.
As for noise, however, they are less noisy”
- Gabox released inst_gaboxBv3 instrumental model (B for bleedless)
Inst. bleedless: 41.69, fullness: 32.13
“can be muddy sometimes”
- mesk’s training model guide link has been changed (the previous one has been deleted)
- Apart from new drumsep models on MVSEP, also moises.ai has their own drumsep model (paid).
Probably their base drums model used for drumsep is not better than other solutions, so check this section of the doc to get better drums to separate first to test it out, although one user reported that moises’ drums model (free), probably vs Mel-Roformer on MVSEP or x-minus (not sure) can give “better results (...) if the input material is for example cassette-tape sourced or post-FM).
- Joseph made the SESA Colab private till some stuff will be fixed in the future.
Consider using this Colab with newer models added at this time.
- (x-minus) Inst V7 model by Gabox replaced v1e model by Unwa.
It can be still accessed by these links:
https://uvronline.app/ai?hp&test (premium)
https://uvronline.app/ai?test (free)
(v1e might be still fuller, and impair fewer instruments in cost of more noise, also be aware that separation on x-minus might differ from Colabs, MSST or UVR, possibly due to different inference parameters)
- Training (and inferencing) locally on Radeon using MSST, specifically RX 7900 XTX, was confirmed to work by Unwa on Ubuntu 24.04 LTS using Pytorch 2.6 for ROCm 6.3.3.
Currently, officially supported consumer GPUs with ROCm are:
RX 7900 XTX, RX 7900 XT, RX 7900 GRE and AMD Radeon VII. But in fact, there are more consumer Radeons confirmed to work already too.
“No special editing of the code was necessary. All we had to do was install a ROCm-compatible version of the OS, install the AMD driver, create a venv, and install ROCm-compatible PyTorch, Torchaudio, and other dependencies on it.” More
“So far I have not had any problems. Running the same thing appears to use a little more VRAM than when running on the NVIDIA GPU, but this is not a problem since my budget is not that large and if I choose NVIDIA I end up with 16GB of VRAM (4070 Ti S/4080 S).
Processing speeds are also noticeably faster, but I did not record the results on the previous GPU, so I can't compare them exactly.“ More
- Inst_GaboxFVX model was released (which is “instv7+3” - so probably fuller than instv3) and
- INSTV7N (so more noisy than INSTV7; “it's [even] closer to fv7 than inst3”) yaml
- Gabox Karaoke model got updated (links have been replaced, and the old deleted from the repo),
- and also final INSTV7 was released (“I hear less noise compared to v1e, but it has worse bleedless metric”)
- Gabox released instv7 beta 2
Inst. bleedless: 34.66, fullness: 38.96
and instv7 beta 3
“Both are noisy with small vocal residuals in places where music is low and deletions of some musical instruments.”
- New 4 stem drumsep SCNet model (kick, snare, toms, cymbals) has been added on MVSEP (best SDR for kick and similar for toms to previous 6s modelm -0.01 SDR difference), and also 8 stems ensemble of all other drumsep models (besides the older Demucs model by Inagoy) metrics
- Gabox released instv7beta model yaml
Inst. bleedless: 35.01, fullness: 38.39
“sound is good, but sometimes some instruments are lowered or deleted”
“while the annoying buzzing/noise is still present, it seems to be more contained.”
- Gabox released Mel KaraokeGabox model (uses Aufr’s config) | Colab
“the lead vocals are good and clean!
While the backing tracks are lossy for this model, [it still] provide[s] great convenient for those who need LdV”
“The model doesn't keep the backing vocals below the main vocals, sometimes the backing vocals will be lost even though there are backing vocals there.”
- New FullnessVocalModel (yaml) vocal model was released by Aname | Colab
Voc. bleedless: 32.98 (less than beta 4), fullness: 18.83 (less than big beta 5e/voc_fv4/becruily, more than beta 4)
“While it emphasizes fullness, the noise is well-balanced and does not interfere much. (...)
in sections without vocals, faint, rustling vocals can be heard.”
We have some report of very long separation of this model in UVR on Macs.
> Try to change chunk_size: 529200 to 112455 for that model/yaml (but it’s dim_t 256 equivalent, so something higher to test might be a better idea too)
- (SESA) No audio file found bug fixed
- “In my testing, I've found that SCNet very high fullness (on mvsep) put through Mel-Roformer denoise (average) and UVR denoise (minimum) has the best acapella result
would love to see people's thoughts” dynamic
- Gabox released voc_fv4 | yaml | Colab
Voc. bleedless 29.07, fullness 21.33
“Very clean, non-muddy vocals. Loving this model so far” (mrmason347)
“lost some of the trumpet sound while on Becruily model can keep it, but some also was lost”
- Joseph fixed some bugs and errors in SESA Colab, and also added new interface
There are still some issues with auto ensemble till further notice.
- Fixed
- unwa released Big Beta 6 vocal model | yaml | Colab
“Although it belongs to the Big series, the characteristics of the model are similar to those of the FT series. (...) this model is based on FT2 bleedless with the dim increased to 512”
Muddier than Big Beta 5, might be better than FT2 at times.
“If you liked the output of the Big Beta 5e model, you may not like 6 as much; it does not have the output noise problem of 5e, but instead sacrifices Fullness. (...) Simply put, it is a more conservative model” unwa
- To get rid of noise in INSTV6N, use denoisedebleed.ckpt (yaml) on mixture first, then use INSTV6N - “for some reason it gives cleaner results” (Gabox)
- New Gabox model released: INSTV6N (noisy) | yaml | Colab | SESA | metrics:
inst bleedless: 32.63, fullness: 41.68 (more than v1e)
Interestingly, some people find it having less noise vs v1e, and more fullness.
Also, it has more fullness vs INSTV6, and more noise.
“v1e sounds like an "overall" noise on the song, while v6n kind of mixes into it.
v6n also sounds like two layers, one of noise that's just there. And the other one mixes into the song somehow.
Using the phase swap barely makes it any better than phase swapping with v1e though” - vernight
Also Kim model for phase swap seems to give less noise than unwa ft2 bleedless
- Demudder in UVR using at least DirectML (Intel/AMD) works only if "Match freq cut-off" is enabled in MDX settings. Otherwise, you’ll get “Format not recognised” error.
- SESA Colab might undergo some issues with hyper_connections at the moment.
It might be fixed tomorrow.
- Done
- SESA Colab update:
Voc_Fv3 (by Gabox)
dereverb_mel_band_roformer_mono (by anvuew)
MelBandRoformer4StemFTLarge
INSTV5N (by Gabox)
denoisedebleed (by Gabox)
- Gabox released denoisedebleed.ckpt | yaml | Colab for noise from fullness models (tested on v5n) - it can't remove vocal residues
- Aname released small inst/voc 200MB Mel-Roformer with null target stem (link)
- v5_noise inst model released | yaml | metrics | Colab
- New Gabox vocal model released: voc_Fv3.ckpt | yaml | Colab
Enthusiastic opinions so far
- INSTV6 by Gabox and De-reverb (Mono) by anvuew models added on x-minus | Colab
V6 “is slightly better than v5” (although not for everyone), but “v1e still gives better fullness, but noise [in v1e] is a problem” old viperx 12xx models have less problems with sax.
- (added on MVSEP as SDR 13.72) ZFTurbo trained new SCNet XL model for drums.
“I have 2 versions: one is slightly higher SDR and avg Bleedless.
Second is better for fullness and L1Freq.
Previous best SDR model had 13.01 (it's SCNet Large).” Metrics
15.7180 (13.72) one has much better fullness metric.
“It's far superior to the other one, but I still hear some weird parts.
It still messes up on some percussion.
The drums stem sounds really weird.
The no drums is alright except for some bleeding but yeah the drums is quite muddy” - insling
- Gabox released new fine-tunes of his inst Mel-Roformer models (click):
inst_gabox2.ckpt, inst_gabox3.ckpt. INSTV5.ckpt. INSTV6.ckpt
with one opinion that the last one is his best inst model so far.
“seems like a mix between brecuily and unwa's models”
“confuses way less instruments for vocals than v1e, but it's still not as full as v1e (...) But it's a very good model”
Rarely it can give “Run out of input error” in UVR when installing using the new Model install option (moved model has 0 bytes), while V5 worked correctly, then move the ckpt to Ultimate Vocal Remover\models\MDX_Net_Models manually.
- We’re aware that the x86-64 version of the latest UVR patch for Mac went offline.
Anjok was pinged about it.
- anvuew released new dereverb_mel_band_roformer_mono_anvuew_sdr_20.4029 model.
“supports mono, but ability to remove bleed and BV is decreased
should not matter whether it's singing or speech, because my dataset contains speech.”
- MedleyVox Colab is currently broken (you can use MVSEP instead)
> fixed:
https://colab.research.google.com/drive/10x8mkZmpqiu-oKAd8oBv_GSnZNKfa8r2?usp=sharing (although initialization now takes 7 minutes, GDrive integration added)
- Phase remix functionality was added to SESA model inference Colab
https://colab.research.google.com/drive/1U28JyleuFEW6cNxQO_CRe0B2FbNoiEet
- (MVSEP) ZFTurbo added new SCNet XL “high fullness” and “very high fullness” models on the site (metrics).
They’re good for both vocals and instrumentals, and sometimes are fuller than v1e, although with more noise, which can be too strong for some people, but not all.
“very high fullness” variant have both vocals and instrumental fullness and bleedless metric better than the “high fullness”.
“they also correctly detect "complex" (for the AI) instruments as part of the instrumental track rather than in vocals (like flute or sax for example), which isn't the case for v1e and Fv5. Example: sax solo of Shine On You Crazy Diamond detected the sax solo as part of acapella using v1e or Fv5 or becruily inst.” dca100fb8
The noise "gonna go nuts with distortion, compression and other vct plugins" John.
Both variants "have the same amount of buzzing noise"
"Instrumentals are very good. It's holy shit level. Unwa v1e/Gabox Fv5 are still amazing, it's just nice to have such a decent model like these new ones on a different arch" dca
From the songs I’ve tested, SCNet is incredible. Very full sounding" mrmason347
"Regular high fullness though has a less full instrumental but quite good acapella" theamogusguy
VHF leaves some vocal residues in metal, but seems to do well for e.g. alt-pop.
"scnet doesn't pick up the drone backing vocals, but 10.2024 has mad violin bleed in the vocals" dynamic64
For "mainly orchestral tracks with choir" "it gave me noticeably fuller results than v1e" Shintaro5034.
"For noisy/dense mixes though, Roformers are probably better, especially for inst.
scnet seems better at preserving treble in some vocals. These high fullness models especially so. So maybe teaming SCNet up with Roformer might give a nice middle ground"
“Rofos are really bad for some kinds of EDM that are very aggressive (Dubstep, Trance, Breakcore, etc...), also it has a very hard time with Experimental (IDM)”
VHF seems to have more crossbleed in some songs, along with also basic XL model. Some songs which sound full enough even with basic SCNet XL. While others sound muddy (dca)
- (X-Minus) Mel Kim model has been replaced for phase correction by Unwa’s Kim FT2 model for premium users
- New sites and rippers added:
https://yams.tf/ (Qobuz, Tidal, Spotify, Apple Music [currently 320kbps], Deezer) - for URLs
https://us.deezer.squid.wtf/ (Deezer only) - for queries
https://github.com/ImAiiR/QobuzDownloaderX (local ripper for premium accounts or provided ARLs)
- FlashSR has been released (Colab with chunking and overlap by jarredou).
It’s a diffusion distillation of AudioSR, and has lower Aura MR STFT metric, and usually lower quality as well, but it might give better results for music for some people
- (MVSEP) “We trained new DrumSep models (5 stem and 6 stem) based on SCNet XL.
* 5 stems: cymbals, hh, kick, snare, toms
* 6 stems: ride, crash, hh, kick, snare, toms”
Both have better SDR than the previous MDX23C model by jarredou and Aufr33.
The 5 stems variant has e.g. better snare SDR than the 6 stems variant. Full metrics.
It doesn't work correctly on the site yet, it will be announced in the link above by ZFTurbo when it will be fixed.
- They work already
- Unwa released ft2 bleedless vocal model | Colab
https://huggingface.co/pcunwa/Kim-Mel-Band-Roformer-FT/tree/main
voc bleedless 39.30 | fullness 15.77 | SDR 11.05
- instv5 model released by Gabox (39.40 inst fullness | inst. bleedless 33.49) link | yaml | Colab | x-minus
“it seems that most vocal leakage is gone, and the noise did significantly decrease, although there's still a bit more noise presence than v1e.
In terms of fullness though, for some reason it sounds as if it's actually less full than v1e, despite the higher instrumental fullness SDR.
Despite v4's significant amount of noise, it seems to be the only model that gave me a fuller sounding result compared to v1e that's actually perceivable by my ears.” Shintaro
- New inst/voc SYH99999 models released
https://huggingface.co/SYH99999/MelBandRoformerSYHFTB1/tree/main
- (x-minus) Phase fixer added for Gabox fv3 and becruily models for premium users
- For vocals, you can alleviate some of the noise/residues in unwa’s 5e model by using phase fixer/swapper and using becruily vocals model as a reference (imogen).
- For instrumentals, you can try unwa's v1e with phase swap at 500/500 with original mel band of kim. It consistently gives less noise - midol
"500 / 500 means you use original phase below 500 Hz and hard cutoff/swap to transferred phase above 500hz. (this can potentially create phase artifacts at 500Hz because of hard swap)
500 / 20000 means you use original phase below 500hz and progressively crossfade to transferred phase until 20000hz and transferred phase is used above 20000hz. So it's softer phase swap below 20kHz" - jarredou
“using 500 on both parameters really does make me have the illusion that I have produced the official instrumental. Even tho it's unofficial haha” - midol
- Deezer on Lucida doesn’t work. Doubledouble.top came back (probably temporarily), but returns mp3 128kbps from Deezer now. Also, it supported Apple Music unlike Lucida, but now it doesn’t work, (check current services status). Besides, occasionally it can happen that rips from Amazon only on doubledouble have quality higher than 44/16. Plus, downloading full albums frequently fails, swhile single songs downloading works.
- Lots of new Gabox models added since then, including:
a) BS-Roformer instrumental variant, which doesn’t struggle so much with choirs like most Mel-Roformers, although may not help in all cases (link)
b) inst_gaboxFv3.ckpt - like v1e when it comes to fullness (added on x-minus)
Inst SDR 16.43 | inst. fullness 38.71 | inst. bleedless 35.62
It might pick up entire sax in vocal stem.
- Gabox models have been added to SESA Colab (you’ll find more info about them later below).
- Along with UVR Roformer beta patch #14, Anjok released the long anticipated demudder.
It's in Settings > Advanced MDX options (so works only for Roformers and MDX models).
It consists of three methods to choose from (each separates your twice):
- Phase Rotate
- Phase Remix (Similar to X-Minus) - “the fullest sounding, but can leave a lot of artifacts with certain models. I only recommend that method for the muddiest models. Otherwise, Combined Methods is the best” “I don't recommend using phase remix on the Instrumental v1e model. I recommend combined methods or phase rotate for models produce fuller instrumentals.” Anjok
It might leave some choruses when using V1E (Fabio)
- Combine Methods (weighted mix of the final instrumentals generated by the above). More in the full changelog. You cannot use demudder on 4GB AMD GPUs with 800MB Roformers with even 2 seconds chunk size set (memory allocation error).
“It's meant to solely target instrumentals. The vocals should stay exactly as before.
For Roformer models, it must detect a stem called "Instrumental” so for some models like Mel-Kim, you need to open model’s corresponding yaml, and change “other” to “instrumental”.
“I've noticed with the few amounts of tracks I've tried, demudding can sometimes accentuate instances of bleeding or otherwise entirely missed vocal-like sounds”
In case of file not found error on attempt of using demudder, reinstall UVR.
“I put the demudded instrumental in the bleed suppressor, and it sounds really good, almost noise free. I either do a bleed suppressor or a V1/bleed suppressor ensemble” gilliaan
“With the new config editor feature you could probably edit the configs of models to have the vocal stem labelled as the Instrumental stem so the demudder demuds the vocal stem, it definitely still makes a difference.
I accidentally did this when installing another model, but it seems to actually have an effect on vocal stems too.
You just change the target instrument from vocals to instrumental I think (don't move the stems around)
You can verify it works if the stems are the other way around when processing (vocals are in the file labelled as Instrumental). Then you can use the demudder on the vocals that way
I think. If you want to use the demudder with other models that aren't labeled with instrumental, you'll have to select the stem you want to demud and replace it with Instrumental.
Though demudding the vocal stem will definitely make it quite noisy depending on what model you use, though there appears to be instances where demudding the vocal stem can mildly help with certain effects but i did not test this enough” stephanie
Anjok: “Just a few quick notes on the Demudder:
It works best on tracks that are spectrally dense (ex. Metal, Rock, Alternative, EDM, etc.)
I don't recommend it for acoustic or light tracks.
I don't recommend using it with models that emphasize fuller instrumentals (like Unwa's v1e model).
I do plan on adding options to tweak the phase rotation.
I also plan on adding another combination method that may work better on certain tracks.”
UVR_Patch_1_21_25_2_28_BETA:
Small patch (you must have a Roformer Patch [e.g. #13] previously installed for this to work): Link
Also, minor bugs fixed, calculate compensation for MDX-Net v1 models added.
The MacOS version will be released later (observe).
Be aware that at least Phase Rotate doesn’t work on AMD and 4GB VRAM GPUs on even 88200 chunk size (prev. dim_t 201 - 2 seconds) and 800MB Roformers like Becruily’s, while 112455 (2,55s, prev. dim_t = 256) works just fine for normal separation.
- BS-RoFormer 4 stems model by yukunelatyh / SYH99999 added on x-minus
Since then, a new version was added (later epoch, but it has lower SDR for all stems).
https://uvronline.app/ai?discordtest
Some people liked the v1 more than Demucs, but “it's like demucs v4 but worse i think
the vocals have a ton of bleed, the bass is disappointing tbh
the other stem has a ton of bgv and adlib bleed in it” Isling
It has SDR metrics for all stems worse than 4 stem BS-Roformer by ZFTurbo and demuics_ft.
- Aname also released 4 stem BS-Roformer model | yaml
It has better SDR than the above (as in the SDR metrics link above), but worse than the other two mentioned
- Gabox released Mel-Roformer instrumental model (Kim/Unwa/Becruily FT): https://huggingface.co/GaboxR67/MelBandRoformers/tree/main/melbandroformers
inst bleedless: 37.40 (better than v1e by 1.8), fullness 37.07 (better than unwa inst v1 and v2)
“It’s like the v1 model with phase fixer, but it gets more instruments,
like, it prevents some instruments getting into the vocals”, “sometimes both models don't get choirs”.
- instrumental variant called fullness v1 (“noisier but fuller”)
inst bleedless: 37.19, fullness 37.26
(thanks for evaluation to Bas Curtiz and his GSheet with all models.)
- fullness v2 released
- fullness v3 released
- B (bleedless) v1/v2 variants released
- voc_gabox.ckpt:
voc bleedless: 34.66 (better than 5e), fullness 18.10 (on pair with beta 4)
- Vocal model F v1
- Vocal model F v2
voc bleedless: 33.4013, fullness: 19.3064
- Issues with dataset 4 in MSST repo were fixed
“I think that issue could also explain why training de-reverb models with pregenerated reverb audio files was not working that well, as reverb was not aligned with clean dry audio as it should have been.” jarredou (more)
- Aufr33 BS-Roformer Male/Female beta (model | config | config for UVR | tensor match error fix) added on Colab (based on BS-RoFormer Chorus Male Female by Sucial) along with Unwa’s Kim FT2
- Anjok released the MacOS versions of UVR Roformer beta patch #13.1 applying hotfix to address a few graphics issues:
- Mac M1 (arm64) users - Link
- Mac Intel (x86_64) users - Link
- Anjok released UVR beta Roformer patch #13 for Mac (Windows further below):
UVR_Patch_1_15_25_22_30_BETA:
- Mac M1 (arm64) users - Link
- Mac Intel (x86_64) users - Link
- mesk wrote a good comprehensive training guide for beginning model trainers. Later you can proceed to read further section of this doc for more details and arch explanations
- Apple Music bot link added in this section (thx mesk)
- mrmason347 and Havoc shared an interesting method to get cleaner vocals. The last point of tips to enhance separation here:
Separate with becruily Mel Vocal model and its instrumental model variant, then get vocals from the vocal model, and instrumental from instrumental model, import both stems for the DAW of your choice (can be Audacity) so you’ll get a file sounding like original file, then export - perform a mixdown of both stems, then separate it with vocal model
- If you somehow still struggle with “norn” issues in UVR, see at the bottom of the section here
- Dango released “Reverb Remover” - click
“it's very similar to RX11 Dialogue Isolate, good/real-time set to 5
it's like listening to the same inference files” John; probably also works in mono, you can get 30 seconds for free)
- filegarden added to the list of cloud services (seems to be unlimited, registration required, link shortener with custom name available)
- your-good-results and your-bad-results channels have been reopened on the server, but you need to paste links to uploads instead of uploading audio files directly on Discord due to copyright issues the server was undergoing
- If you want to use Phase fixer Colab with cut-offs suggested by CC Karaoke, check here
- Unwa’s Kim FT2 model added to the inference Colab (both inst and voc becruily models are added too)
- jarredou released Custom Model Import Version of the inference Colab. You can use it if we don’t add any new model to the main Colab on time, or you test your own models.
Just make sure that pasted link haven’t “downloaded the webpage presenting the model instead of the model itself.”
So, e.g. for yamls pasted from GH, use:
https://raw.githubusercontent.com/ZFTurbo/Music-Source-Separation-Training/main/configs/config_vocals_mdx23c.yaml'
Instead of:
https://github.com/ZFTurbo/Music-Source-Separation-Training/main/configs/config_vocals_mdx23c.yaml'
And for HF, follow the pattern presented in the Colab example (so with the resolve in the file address)
- model_fusion.py by Sucial
This script seems to save the weighted ensemble of three models into a checkpoint called "fused". The result is not bigger than a single model.
Probably you could basically create one checkpoint getting the same or similar results of manually weighted models, and not inference every of them one by one.
- Becruily models added on MVSEP and instrumental on variant on x-minus
- ZFTurbo “added new organ model: MVSep Organ (organ, other).
Demo: https://mvsep.com/result/20250116160630-f0bb276157-mixture.wav”
- Anjok released a patch #13 fixing following issue with no sound on some Roformer models (like avvuew’s and sucial’s de-reverb) on GTX 10XX or older (Windows):
UVR_Patch_1_15_25_22_30_BETA:
“- Full Install: Link
- Patch Install (use if you still have non-beta UVR installed): Link
- Small Patch Install (have a Roformer patch previously installed for this to work): Link
The issue was some older GPU's are not compatible with Torches "Inference Mode," (which is apparently faster) so it's now using "No Grad" mode instead. Users can switch back to using "Inference Mode" via the advanced multi-network options.
The MacOS version will be released in a few days. I just need to finish testing out all the models and networks and ensure all the kinks are worked out.” More
- Users undergo some issues (no sound) with Mel-Roformer de-reverb by anvuew (a.k.a. v2/19.1729 SDR) since the latest UVR beta #11/12 updates (the issue seems to occur only on GTX 10XX series, and maybe older). Anjok’s working on the issue.
You should be able to use more than one UVR installation at the same time when one’s been copied before updating (patch #10 still works) or use MSST repo and/or its GUIs.
- Anjok released patch #12 which is a hotfix for the 4 stem BS-Roformer model by ZFTurbo (trained on MUSDB)
UVR_Patch_1_13_0_23_46_BETA_rofo_fixed.exe (Windows only)
- Anjok released a new UVR beta Roformer patch #11 (Windows only for now):
UVR_Patch_1_13_0_23_46_BETA_rofo
It fixes 4 bugs: with VR post-processing threshold, Segment default in multi-arch menu, CMD will no longer pop-in during operations, and error in phase swapper.
More details/potential updates.
Standalone (for non-existent UVR installation)
UVR_1_13_0_23_46_BETA_full.exe
For 5.6 stable (so for non-beta Roformer installation)
UVR_Patch_1_13_0_23_46_BETA_rofo.exe
Small (for already existing Roformer beta patch installation)
UVR_Patch_1_13_25_0_23_46_rofo_small_patch.exe
- New beta UVR Roformer patch #10 released by Anjok (for now, only small patch for already existing beta Roformer installation is available, and only for Windows, check here for Mac later)
UVR_Patch_1_9_25_23_46_BETA_rofo_small_patch - Link
Added SCNet and Bandit archs with models in Download Center (SCNet models using ZFTurbo’s unofficial code update will not work since they appear to require a library "mamba_ssm" that is only available in Linux), fixed compatibility with some newer Roformer models (wesley’s MDX23C and Roformer Phantom center models, and 400MB inst small by Unwa), new Model Installer option added, model configuration menu enhanced, allowing aliases to selected models, added compatibility for Roformer/MDX23C Karaoke models with the vocal splitter, VIP code issue is gone, issues with secondary models options and minor bugs and interface annoyances are addressed, “improved the "Change Model Settings" menu. Now, any existing settings associated with a selected model are automatically populated, making it easier for users to review and adjust settings (previously, these settings were not visible even if applied).”.
If you have Python DLL error on startup, reinstall the last beta update using the full package instead, then the small installer from the newer patch.
“If you see a different usage of VRAM than with previous Roformer beta version, it could also be because the new beta version doesn't rely on 'inference.dim_t' value anymore (if you were using edited "dim_t" value)
You have to edit audio.chunk_size now (see here for conversion between dim_t and chunk_size” it’s “In model yaml config file, at top of it, chunk_size is first parameter (...) you can edit model config files directly inside UVR now.”
“Unfortunately, SCnet is not compatible with DirectML, so AMD GPU users will have to use the CPU for those models.
Bandit models are not compatible with MPS or DirectML. For those with AMD GPU's and Apple Silicon, those will be CPU only.
The good news is those models aren't all that slow on CPU.” - Anjok
Annoying CMD window will randomly pop up again when ffmpeg and Rubber Band are used. Regression will be fixed.
- Newer Mel-Roformer Male/Female model was added by ZFTurbo on MVSEP (SDR: 13.03 vs 11.83 - the previous SCNet one, and much better bleedless metric 41.9392 vs 26.0247 with only 0.2 fullness decrease)
“ I find it acts differently from Rofo or UVR2. Sometimes it's the one of the three that gets it right., and not strictly for male/female.” CC Karaoke
- Aufr33 released his own BS-Roformer Male/Female (currently beta) model based on BS-RoFormer Chorus Male Female by Sucial.
“this model only works with vocals. You need to pre-isolate the vocals.”
Added on MVSEP and x-minus for premium (in the new Other menu).
Weights: https://mega.nz/file/XZwV2QwB#5nvWpmvtoBMTJkpor-lMUZCbBZWDH-3i52ELJS_JmcU
- Unwa released a new version of his Mel-Kim fine-tune (ft2)
https://huggingface.co/pcunwa/Kim-Mel-Band-Roformer-FT/tree/main
It tends to muddy instrumental outputs at times, similarly like the OG Kim’s model was doing, which didn’t happen in the previous ft model by Unwa.
Metrics. PS. All unwa models were trained on 3060 Ti!
- Unwa released 400MB experimental BS-Roformer inst model
https://huggingface.co/pcunwa/BS-Roformer-Inst-EXP-Value-Residual
It’s using a new Value Residual Learning added to Roformer arch by Lucidrains in the OG Roformer. If it wasn’t made compatible with MSST repo already, replace bs_roformer.py from this repo and
from bs_roformer.attend import attend
⇩
from models.bs_roformer.attend import attend
in bs_roformer.py file
“I think it sounds better than large rn but still not good, needs some [more] epoch[s]!”
[later the VRL was added as Mel-Roformer v2 model type in UVR so it’s compatible with the model]
- New dereverb model(s) released by Sucial - “fused”: model | yaml
“trained two new models specifically targeting large reverb removal. After training, I combined these two models with my v2 model through a blending process, to better handle all scenarios. At this stage, I am still unsure whether my new models outperform the anvuew's v2 model overall, but I can confidently say that they are more effective in removing large reverb.” More
- Becruily Mel inst and voc models added on MVSEP and inst variant on x-minus
- ZFTurbo released new models on MVSEP:
a) a new Male/Female separation model based on SCNet XL
SDR on the same dataset: 11.8346 vs 6.5259 (Sucial)
Model only works on vocals. If the track contains music, use the option to "extract vocals" first. Sometimes the old Sucial model might still do a better job at times, so feel free to experiment.
b) SCNet XL (vocals, instum)
Inst SDR: 17.2785
Vocals have similar SDR to viperx 1297 model,
and instrumental has a tiny bit worse score vs Mel-Kim model.
- “All Ensembles on MVSep were updated with latest release [SCNet XL] increasing vocals SDR to 11.50 -> 11.61 and instrum SDR: 17.81 -> 17.92”.
- (MSST) You can now inference mono files without any issue
- You can now use “batch_size=1 without clicks issues (with overlap >= 2 of course)” - jarredou
- Becruily’s released instrumental and vocal Mel-Roformer models | Colab | UVR beta |
Instrumental model files | Inst SDR 16.4719 | inst fullness 33.9763 | bleedless 40.4849
Vocal model file | Vocals SDR 10.5547 | voc fullness 20.7284 | bleedless 31.2549 |
config with ensemble fix in UVR.
Instrumental model is as clean as unwa’s v1, but has less noise and, and it can be got rid well by Mel denoise and/or Roformer bleed suppressor. Inst variant “removed some of the faint vocals that even the bleed suppressor didn't manage to filter out” before”. Doesn’t require phase fix from Mel-Kim like unwa models below.
“it handles the busy instrumentals in a way that makes VR finally an arch of the past”
Correctly removes SFX voice. More instruments correctly recognized as instruments and not vocals, although not as much as Mel 2024.10 & BS 2024.08 on MVSEP, but still more than unwa’s inst v1e/v1/v2. (dca100fb8).
Trumpet or sax sound which on unwa model was lost, can be recovered
on becruily's model (hendry.setiadi)
The instrumental model pulled out more adlibs than the released vocal model variant - it pulled out nothing (isling).
“Vocal model pulling almost studio quality metal screams effortlessly. Wow, I've NEVER heard that scream so cleanly” (mesk)
The model was trained on dataset type 2 and single RTX 3090 for two days (although with months of experimentation beforehand). SDR metrics are lower than Mel-Kim model.
If you use lower dim_t like 256 at the bottom of config for slower GPU, these are the first models to have very bad results with that setting.
You can experiment with phase fixer with santilli_ suggestion “Using becruily's vocals as source and inst [model] as target, and changing high frequency weight from 0.8 to 2 makes for impressive results”.
- Phase fixer Colab (update 2) by santilli_ released - it can use e.g. Mel-Kim model phase for unwa’s v1e/v1/v2 models to automatically get rid of some noise during separation (it might no longer work due to the last changes in MSST repo), it includes also becruily models
- A small UVR Roformer beta patch #9 fixing mainly Apollo arch released also for Mac (UVR_Patch_12_8_24_23_30_BETA):
Mac M1 (arm64) users - Link
Mac Intel (x86_64) users - Link
- New “MVSep Bass (bass, other)" SCNet model available on MVSEP
“It achieved SDR: 13.81. In Ensemble it gives 14.07 - which is a new record on the Leaderboard.” ZFTurbo
“It passes Food Mart - Tomodachi Life test. That's the first model to.”
“All bass models have problems with fretless bass”
There’s already an option to combine all SCNet+BS Roformer+HTDemucs bass models for 14.07 SDR.
Ensembles have been updated with this model too.
- Reverb removal by Sucial v2 (Mel-Roformer) model added on MVSEP (update of the previous model)
- Lew universal upscaling model has been added on x-minus/uvronline too (premium users).
Just a reminder - it’s not for badly mixed music, it’s for lossy files (also on Colab/MVSEP/UVR beta [at least support for a model file])
- ZFTurbo released a new 4 stem XL model trained on SCNet.
“I have great results comparing with SCNet Large model (by starrytong).”
SCNet Large MUSDB test avg: 9.70 (bass: 9.38, drums: 11.15 vocals: 10.94 other: 7.31)
SCNet XL MUSDB test avg: 9.80 (bass: 9.23, drums: 11.51 vocals: 11.05 other: 7.41)
SCNet Large Multisong avg: 9.28 (bass: 11.27, drums: 11.23 vocals: 9.05 other: 5.57)
SCNet XL Multisong avg: 9.72 (bass: 11.87, drums: 11.49 vocals: 9.32 other: 6.19)
A new SCNet bass model is incoming and already surpassed metrics of ZFTurbo’s HTDemucs and BSRoformer bass models.
- Anjok released a small UVR Roformer beta patch #9 fixing mainly Apollo arch:
UVR_Patch_12_8_24_23_30_BETA
Windows only for now: Full | Patch (Use if you still have non-beta UVR installed) |
Small Patch (You must have a Roformer patch previously installed for this to work)
Changelog:
Apollo fixes: “Chunk sizes can now be set to lower values (between 1-6)
Overlap can be turned off (set to 0)”
Fix both for Apollo and Roformers: now 5 seconds or shorter input files no longer cause errors.
OpenCL was wrongly referenced in the UVR. It was actually DirectML all the way, and Anjok changed all the OpenCL names in the app into DirectML.
- Unwa released a new Kim-Mel-Band-Roformer-FT vocal model | Colab
It enhances both our new bleedless (36.95 vs 36.75) and fullness (16.40 vs 16.26) metric for vocals vs the original Mel Kim model. SDR-wise it’s also a tad lower (10.97 vs 11.02)
(thx Bas Curtiz)
- Male/female BS-Roformer separation model has been released by Sucial
https://github.com/ZFTurbo/Music-Source-Separation-Training/issues/1#issuecomment-2525052333
If they sing at intervals (one by one), they cannot be separated.
Works pretty good, bleed might occur occasionally. Also, it seems to pick up various people dialogues.
If you want to use the model in UVR, use this config (thx Essid)
If you have "The size of tensor a (352768) must match the size of tensor b (352800) at non-singleton dimension 1" e.g. in python-audio-separator, use this config (thx Eddycrack864)
- Anjok released UVR Roformer beta patch #8 for Win: full | patch | Mac: M1 | x86-64
- UVR_Patch_12_3_24_1_18_BETA
Apollo arch was made compatible with MacOS MPS (metal) and OpenCL, but with it, it might be unstable and very RAM intensive - use chunk size over 7 to prevent errors (currently it’s not certain that all models will work with less than 12GB of VRAM).
Apollo is now compatible with all Lew models (fixed incompatibility with any other than previously available in Download Center). Fixed (presumably regression with) Matchering.
How would I assign a yaml config to an Apollo model on the new UVR [patch]?
“1. Open the Apollo models folder
2. Drop the model into the folder
3. From the Apollo models folder, drop the yaml into the model_configs directory
4. From the GUI, choose the model you just added and if the model is not recognized, a pop-up window will appear, and you'll have the option to choose the yaml to associate with the model.” - Anjok
“I found some overlapping issues in the UVR [using Apollo vs Colab]. like some short parts sounding duplicated overlaid” The issue is caused by different chunking, which on Colab is preconfigured to use 15GB of VRAM. “chunk_size has influence on results, colab uses 25sec (or 19 for latest lew model)”
- Anjok released UVR Roformer beta patch #7 for Windows: full | patch
- UVR_Patch_12_2_24_2_20_BETA (version for Mac probably tonight, observe or on GH)
It introduces support for Apollo arch. The OG mp3 enhancer and Lew v1 vocal enhancer were added to Download Center. Probably, now you’ll be able to add newer Lew uni enhancer and v2 vocal enhancer manually. The arch is located in Audio Tools. Sadly, this arch cannot be GPU-accelerated with OpenCL, so using AMD and Intel cards (you’re forced to use CPU, which might be long).
Also, “Phase Swapper” a.k.a. Phase fixer for Unwa inst models was added to Audio Tools.
- New SCNet and MelBand DnR v3 (SFX) models were added on MVSEP (along with optional ensemble). “The metrics turned out to be better than those of the similar model Bandit v2” (25.11)
- We fixed some issues (IndexError) with jarredou’s inference Colab due to the recent updates in the ZFTurbo code (thx for the heads-up, MrG).
- Anjok released UVR Roformer beta patch #6 for MacOS as well: M1 | x86-64
- Lew released a new Apollo universal model for upscaling various lossy audio files (added on MVSEP and x-minus premium).
Unlike the previous mp3 model, it’s able to enhance any formats and not only mp3, including files with hard cutoff like in AAC 128 kbps (see), it struggles with 48 kbps files.
“If anyone wants to run the new model in Colab [already added], set chunk_size to 19. Then the model uses 14.7GB VRAM” (Essid). Sometimes a lower setting is necessary (e.g. for a 3 minute song, otherwise memory error will appear).
As for 27.11.24 it doesn’t work with MSST yet, later added support in UVR patch.
“Actually much better than the original Apollo model. It handles artifacts really well
and also noise, it understands noise while [the] OG model doesn't for some reason” John UVR/simplcup
Specifically for any muddy Roformer vocals, still use Lew vocal enhancer v1/2 as they're better for this task, though they can be noisy (available in the Colab).
“I also included checkpoint you can continue training from” (mirror)
“Q: segments: 5.4 - can I assume that chunck_size if 5.4 * 44100?
A: Yes, and dims 384
The smaller one for inference and bigger one for training
Q: Does that model has a dataset of a wide variety of compression noise and artifacts?
A: mp2, mp3, ogg, wma, aac, opus, low band width wavs. Random speed change augmentations were used too.” “It was mostly trained on music” Lew
- Two weeks ago, unwa and 97chris released a bleed suppressor Mel Roformer model dedicated for instrumentals (made with unwa v1 in mind). It can work with e.g. v1 and v1e. Sometimes it can remove some bleed also after using phase fixer (by becruily) dedicated for v1 model, or used also on x-minus for premium users
- Anjok released a new UVR Roformer beta patch #6
UVR_Patch_11_25_24_1_48_BETA (Windows: standalone | patch | Mac: M1 | x86-64)
addressing “All stem” error issue with viperx’ models.
And with it, a long anticipated MDX-Net HQ_5 model has been released | Colab | MVSEP
(it’s also added to Download Center in previous UVR patch versions).New version of the HQ_5 model is announced to be released in two weeks already.
Instrumentals are slightly muddier than in HQ_4, but vocal residues are also a bit quieter (although rather still present where they were before, maybe with some exceptions).
E.g. some hi hats might get a tad quieter in the mix.
The new model variant in two weeks was said to have fuller instrumentals.
vs unwa’s v1e “HQ5 has less bleed but is prone to dips in certain situations. (...). Unwa has more stability, but the faint bleed is more audible. So I'd say it's situational. Use both. (...) Splice the two into one track depending on which part works better in whichever part of the song is what I'd do.” CC Karaoke
Model | config: "compensate": 1.010, "mdx_dim_f_set": 2560, "mdx_dim_t_set": 8, "mdx_n_fft_scale_set": 5120
- We have some reports about user custom ensemble presets from older versions no longer working (since 11/17/24 patch and in newer ones). Sadly, you need to get rid of them (don’t restore their files manually) or the ensemble will not work and model choice will be greyed out. You need to start from scratch.
- Sucial released a new Mel-Roformer dereverb/echo model (model | MVSEP).
It’s good but doesn’t seem to be better than the more aggresive variant of Mel anvuew's model (models list).
Still, might depend on a use case.
- People experience All stems error with viperx’12xx models in newer versions of UVR Beta Roformer patch (patch #2 was the last confirmed to work with these older models)
- Lucida.to is undergoing some issues with Qobuz links. Tidal and Deezer work, but poorly, occasionally giving errors too, just retry. Doubledouble redirects to Lucida now. In case of problems with accessing the domain in your country, check out lucida.su or VPN.
If you have any problems during downloading files, try out in incognito mode without any browser extension, also download accelerators might cause issues too (FAQ).
- Unwa released a new beta 5 model dedicated for vocals | Colab | MSST-GUI | UVR instr
https://huggingface.co/pcunwa/Mel-Band-Roformer-big/tree/main | yaml: big_beta5e.yaml
It seems to fix some issues with trumpets in vocal stem (maxi74x1).
It handles reverb tails much better (jarredou/Rage123).
“It's noisy and, IDK, grainy? When the accompaniment gets too loud. (...) Definitely not muddy though, which is a welcome change IMHO. I think I prefer beta 4 overall” - Musicalman
“to me, the noise sounds similar to how VR arch models sounded, except it's not poor quality”
“Perhaps a phase problem is occurring (...) The noise is terrible when that model is used for very intense songs” - unwa
Phase fixer for v1 inst model doesn’t help with the noise here (becruily).
“it's a miracle LMAO, slow instrumentation like violin, piano, not too many drums...
it's perfect... but unfortunately it can't process Pop or Rock correctly” gilliaan
“feel so full AF, but it has noticeable noise similar to [Apollo] lew's vocal enhancer”
“the vocal stem of beta5e may have fullness and noise level like duality v1, but it may also suffer kind of robotic phase distortion, yet may also remove some kind of bleed present in other melrofo's.” Alisa/makidanyee
“bigbeta5e is particularly helpful when you invert an instrumental and then process the track with it. It really keeps the quality. Even if the instrumental was a lossy mp3 inverted to a lossless flac file, it cleans it up without making a mess. (...) some songs gets their instrumentals leaked online. And a lot of the time it's a lossy 160kbps mp3 file or even worse, you invert that instrumental file to the real song and process the result using bigbeta5e [to clean the invert]” gilliaan/heauxdontlast
“Ensemble AVG Big Beta 4 + Big Beta 5e is really good to reduce the noise while keeping the fullness” - heauxdontlast
- Unwa released a new Inst v1e model (“The model [yaml] configuration is the same as v1”)
https://huggingface.co/pcunwa/Mel-Band-Roformer-Inst/tree/main | Colab | MSST-GUI | UVR instructions (added in Download Center) | x-minus (link for premium users) | MVSEP
“The "e" stands for emphasis, indicating that this is a model that emphasizes fullness.”
“However, compared to v1, while the fullness score has increased, there is a possibility that noise has also increased.” “lighter compared to v2.”
While SDR-wise it’s worse than previous unwa’s models, it has the best full fullness factor (you can read more about this new method of evaluation later in this section).
The phase fixer doesn’t really fix the noise in this model like in v1.
Like other unwa models, this can also confuse flute, trumpets and saxophone with vocals.
- You might want to use this max ensemble by dca100fb8 (e.g. the BS model here is capable of detecting flute correctly and the Mel - sax and trumpet):
unwa’s v1e + Mel 2024.10 + BS 2024.08 (Max FFT; the latter models on MVSEP, also sometimes unwa's big beta5e can also retrieve missing instruments from v1e when those two fails)
- You might want to check max ensemble of instv1, instv2 and inst v1e - erdzo125
(for even better fullness but more noise - you can consider the phase fix for instv1)
- Anjok released a new beta Roformer patch #5 for UVR (Windows only): UVR_Patch_UVR_11_17_24_21_4_BETA_patch_roformer
“- Fixed OpenCL compatibility issue with Roformer & MDX23C models.
- Fixed stem swap issue with Roformer instrumental models in Ensemble Mode.”
The patch is rather not standalone like patch #3, so have a previous UVR installation.
- Anjok released a new beta Roformer patch #4 for UVR: UVR_Patch_11_17_24_21_4_BETA (Windows: full | patch | Mac: M1 | x86-64)
Minor bug fixes. Most importantly, MacOS version fix:
“Roformer checkbox now visible for unrecognized Roformer models” so now you can use custom Roformer models on MacOS Roformer patch without copying/modifying configuration files from Windows version or other users in order to circumvent the lack of option from Windows version to set that the recognized model is Roformer, so separation will work on that model. Plus it includes all the previous fixes in the previously released patch (so overlap code fixed, so no stem misalignment should occur on certain overlap settings - probably higher overlap now means longer separation)
- Anjok released a new beta Roformer patch #3 for UVR (Windows version for now) UVR_Patch_11_14_24_20_21_BETA_patch_roformer - “this is a patch and requires an existing UVR installation” (so either previous beta Roformer patch or stable 5.6 version).
The new patch fixes the issue with stem misalignment when using incorrect overlap setting for Roformers. Now it uses ZFTurbo code (also for MDX23C), meaning that probably now increasing overlap for Roformers will result in increasing separation times and potentially better SDR (the opposite of what it used to be in the previous beta Roformer patch). Potentially, it might allow using faster settings without stem misalignments or segment popping (when overlap and dim_t was set to 201 and overlap 2) for 4GB VRAM cards and some heavier models.
Among other minor fixes: “Roformer stem naming issues resolved. Fixed manual download link issues in the Download Center. Roformer models can now be downloaded without issue.”. Implementation of SCNet and Bandit archs is still in works.
Full changelog.
- Becruily made a Python script fixing phase with unwa v1 model, so it removes its noise.
You need to run: pip install librosa
in case of “no module named librosa found” error.
“The results are almost, if not the same as x-minus' phase correction.
To use, you need to have the song separated with Kim's melband model and unwa's v1 model.” 32 bit output switch added
“the output length is few ms shorter than the input
the output has little popping in the end”
- SYH99999/yukunelatyh released a MelBandRoformerSYHFTV3Epsilon model.
VS previous SYH’s models “this version is more consistent with separation. It's not what I'd call a clean model; It sometimes lets background noise bleed into the vocal stem. But only somewhat, and depending on how you look at it, it can be a good thing since it makes the vocals sound less muddy.” Musicalman
Since then, there was also a newer MelBandRoformerBigSYHFTV1Fast model released.
- Lew released a v2 of the vocal enhancer model for Apollo trained on Roformer vocal outputs
Added for paid users on x-minus in the Ensemble menu or in the Restoration menu (formerly De-noise) and on Colab. Model files | config.
Works the best potentially on BS and Mel Roformer ensemble, but it might add some noise as well.
The model stopped progressing during training, so probably there won’t be any newer epoch of this model.
- Unwa released v2 version of the inst Mel-Roformer model.
“Sounds very similar to v1 but has less noise, pretty good”
“the aforementioned noise from the V1 is less noticeable to none at all, depending on the track”.
“V2 is more muddy than V1 (on some songs), but less muddy than the Kim model.
(...) [As for V1,] sometimes it's better at high frequencies” Aufr33
Also, SDR got a bit bigger (16.845 vs 16.595)
https://huggingface.co/pcunwa/Mel-Band-Roformer-Inst/tree/main | Colab | MSST-GUI
“It's the same size as the big model with depth 12 and mask_estimator_depth 3.
The improvement was stagnant with the same model size as v1.” - unwa
- The model has been added to UVR Beta Roformer Download Center and x-minus.
- MSST-GUI is now included in ZFTurbo's repo, it's the "gui-wx.py" file” just don’t run it by double-clicking, but run it from CMD.
GPU acceleration working only Nvidia GPUs will give out of memory errors on 4GB VRAM GPUs for Roformers (you can use CPU instead).
"UnicodeEncodeError" means there is disallowed character in your input file name, e.g. “doesn't work with [ and ] in the foldername - known bug”.
- Both duality models and inst v1/2 are now added to UVR Beta Roformer Download Center (problems with duality models in UVR have been fixed)
- Unwa released v2 version of the duality model, slightly a bit better SDR and fewer residues (available in the link below)
"other" is output from model
"Instrumental" is inverted vocals against input audio.
The latter has lower SDR and more holes in the spectrum.
So using MSST-GUI, leave the checkbox “extract instrumental” disabled for duality models.
- Unwa released a new inst-voc Mel-Roformer called “duality”, focused on both instrumental and vocal stem.
https://huggingface.co/pcunwa/Mel-Band-Roformer-InstVoc-Duality/tree/main | Colab
Vocals sound similar to beta 4 model, instrumentals are deprived of the noise present in inst v1 model, but as a downside, they don't sound similarly muddy to previous Roformers.
You can use it in the MSST-GUI for ZFTurbo script (already added) or with the OG ZF repo code. The model will now work in UVR (added in Download Center, but the problem was also fixed by Anjok and added in the OG repo’s yaml)
- New Ensemble button added on x-minus for premium users for the new inst unwa’s model. It corrects the phase and almost removes the noise existing in this model
“This post-processing uses Kim's model. After post-processing, the vocals will be replaced with those of this model.” Examples
Using Mel-Roformer de-noise might be better alternative:
“removes more noise from the song, keeping overall instrument quality more than the new button” koseidon72. But the more aggressive variant of the model sometimes deletes parts of the mix, like snares.
- New Bas Curtiz fine-tuned on MVSEP and unwa’s inst Mel-Band added on MVSEP and x-minus.
Although there were only 5 submission sent to ZFTurbo for fine-tuning, and 30+ is needed, so there is not so much of a difference in the new FT.
“I suggest to all of you, if there is any voice left [in inst v1], use the Mel-Roformer de-noise with minimal aggression. “not only for little voices left, but also for some background noise.
Unfortunately, this new [unwa’s] model doesn't eliminate vocoder voices well from an instrumental”
The model is much faster than beta 4.
- unwa released a new Mel-Roformer model focused on instrumental stem this time (a.k.a. v1):
https://huggingface.co/pcunwa/Mel-Band-Roformer-Inst/tree/main | Colab | UVR instructions | MVSEP
"much less muddy (...) but carries the exact same UVR noise from the [MDX-Net v2] models"
But it's a different type of noise, so aufr33 denoiser won't work on it.
“you can "remove" [the] noise with UVR-Denoise, aggr. -10 or 0” although at least with -10 it will make it sound more muddy like Kim model and synths and bass are sometimes removed with the denoiser (~becruily). UVR-Denoise-Lite doesn’t seem to damage instruments that badly, but still more than Mel denoise (recommended aggr. - 4, with 272 vs 512 windows size it’s less muddy, TTA can stress the noise more, somewhere above 10 aggr. it gets too muddy). UVR-Denoise on x-minus is even less aggressive (it’s medium aggression model for free users without aggression pick), but it might catch ends of some instruments like bass occasionally. Premium minimum aggression model is somehow more muddy, but doesn’t damage instruments. Minus the noise, this is a groundbreaking instrumental model among public models or existing Roformers.
(more training details)
“Flipping the target seems to definitely have effect on the instrumental part!” Bas Curtiz
“I got an error when I set num_stems to 2.” unwa
You can use “target_instrument: null” instead, which is also required for multistem training like on this example ~jarredou
“It's because of the PHASE. I found a way to fix it. Today I will add a new ensemble button.”
- Similarity / Phantom Center Extractor model by wesleyr36 added on MVSEP (Experimental section) and x-minus.pro (Extract backing vocals).
“This model is similar to the Center Channel Extractor effect in Adobe Audition or Center Extract in iZotope RX [and Audacity/Bertom], but works better.
Although it does not isolate vocals, it can be useful.” Aufr33
You can find more on the topic in Similarity Extractor section.
- ZFTurbo released MVSEP Wind model on the site (MelBand/SCNet/ensemble)
Some songs might be separated better vs the model on x-minus, not all.
- A GUI for ZFTurbo's Music Source Separation script for inference called MSST-GUI was released by Bas Curtiz (link with instruction in the description):
https://www.youtube.com/watch?v=M8JKFeN7HfU (reupload)
It has screen reader compatibility, although people can't navigate with the arrow keys in the web view for now, but at least you have HTML source of the page so you can just download models from there.
- Multiple updates were made since that excerpt was written and new models were constantly added
- If you have “ERROR: Could not build wheels for diffq, pesq, which is required to install pyproject.toml-based projects” then
“Edit the requirements.txt file and remove or comment out that line with asteroid” click
Then rerun pip install -r requirements.txt
- If you have decent Nvidia GPU, and no GPU acceleration maybe “Check these commands to install torch version that handle cuda”:
pip install torch torchvision torchaudio --index-url https://download.pytorch.org/whl/cu118
or
pip install torch==2.3.0+cu118 torchvision torchaudio —-extra-index-url https://download.pytorch.org/whl/cu118
or
pip install torch==2.3.0+cu118 --extra-index-url https://download.pytorch.org/whl/cu118
or
pip install torchaudio==2.3.0+cu118 --extra-index-url https://download.pytorch.org/whl/cu118
- New beta 4 of unwa’s Mel-Roformer fine tune of Kim’s voc/inst model released:
https://huggingface.co/pcunwa/Mel-Band-Roformer-big/tree/main | Colab
Be aware that the yaml config has changed, and you need to download the new beta4 yaml.
“Metrics on my test dataset have improved over beta3, but are probably not accurate due to the small test dataset. (...) The high frequencies of vocals are now extracted more aggressively. However, leakage may have increased.” - unwa
“one of the best at isolating most vocals with very little vocal bleed and still doesn't sound muddy” “gives fuller vocals”. Can be a better choice on its own than some ensembles.
- ZFTurbo, the owner of MVSEP, seeks help on improving Bas Curtiz’ ft Mel-Roformer model on MVSEP. How can you help?
1) Find a badly separated song with this model (e.g. bleeding)
2) Find other model which separates your song correctly
3) Send the good results (instrumental stem + vocal stem in stereo/44kHz with the same length) to ZFTurbo on Discord (or email)
The result will be used as input for training a new, fine-tuned model.
- “1) All [MVSEP] ensembles now use Bas Curtiz MelRoformer model. SDR for Multi dataset stayed almost the same, but greatly increased for Synth
2) Drums model were updated for all Ensembles too.
https://mvsep.com/quality_checker/entry/7197” ZFTurbo
- New SCNet Large Drums model added on MVSEP
- https://studio.gaudiolab.io introduced new Noise Reduction feature
- For those having problems with too slow functioning of lucida.to, you can use https://mp3-daddy.com/. Sometimes FLAC option might not work, then download mp3 first, and then FLAC will work (the files have full 22kHz spectrum). Although, sometimes it may fail anyway (not always in incognito mode and with third party cookies allowed, and after long wait after error appeared). Downloading might be possible by manual download with Inspect option in your browser (it starts downloading and interrupts like on GSEP in the old days). Don’t even bother reading their site description - it’s full of AI-written sh&t. Contrary to what they say, it doesn’t support YT or YT Music/Tidal/Deezer links, so you need to use their search engine. So probably the max output quality is 44kHz/16 bit. It doesn’t seem to use Tidal (maybe Deezer).
https://doubledouble.top/ is now also back online and supports Apple Music unlike Lucida, but it might be slower and go offline eventually as before.
- Strings model based on MDX23C arch added on MVSEP. It has low SDR yet (3.84), so it’s hit or miss whether it will work for your song, but some people had even good results at times. ZFTurbo plans to work on it further.
- Finally, HQ_4 released in March has been added also on MVSEP (it was also added on x-minus/uvronline.app not long ago via this link at least)
- Beta 3 of the unwa’s Mel-Roformer fine-tuned Kim’s model released. Fine-tuning was started from scratch on enhanced dataset made with help of Bas Curtiz. As the result, the model is free from the high frequency ringing present in the previous beta models.
“I've added hundreds of GB worth of data to my dataset”.
"definitely better than Kim's now" although vocal residues might occur yet, and then use unwa’s BS-Roformer fine-tune instead. SDR is slightly lower than Kim Mel-Roformer. It’s good for RVC.
- The Lew’s model was added to jarredou’s Colab
- Lew released a model for Apollo, serving to enhance vocal results of Roformers https://ufile.io/09560o34
“You can use it in Music Source Separation Training repo, and it should be compatible with jarredou Apollo Colab” Links a bit below (not compatible with UVR).
- Beta 2 of the unwa’s Mel-Roformer fine-tuned model released (Colab).
Be aware that both models have some ringing issues in higher frequencies. Hard to say if it will be fixed in the further training, Unwa explaining said it was mainly made with vocals in mind so it’s not sure.
- Unwa released beta version of still-in-training Mel-Roformer fine-tuned model of Kim’s. Not tested SDR-wise, but might give better results than the old Unwa’s BS-Roformer model already. Download:
https://huggingface.co/pcunwa/Mel-Band-Roformer-big/tree/main
In UVR consider using dim_t = 1501 at the bottom of the yaml (can be slow), but 1333 or 1301 can be better for e.g. 40 second snippets, while the biggest SDR is for 1101 for all Roformers, but it still depends on a song what gives the best results (in reality, even SDR for each song is different, and bigger SDR not always means better quality, the quality using specific parameters might even differ in certain fragments).
- (uvronline.app/x-minus) New electric and acoustic guitar models by viperx' added on the site for premium users.
Acoustic seems to be good, while electric might be more problematic at times.
- Now lalal.ai have some voc/inst models sounding like some ensemble of public Roformers, but still not as good, although close. Some of their specific models are worth trying out, e.g. lead guitars - the model got better by the time, or also piano model
- https://github.com/JusperLee/Apollo | jarredou Colab
jarredou: “New tool for heavily compressed mp3 restoration, using bandsplitting and roformers. It does work really great if the audio was compressed at 44.1khz sample rate, whatever bitrate [<=128kbps]. BUT if there was some resampling leading to hard cutoff, it will wrongly behave.
The current model of Apollo was only trained on mp3 compressed audio. If you use ogg/opus/m4a/whatever else compressed audio as input, it's not guaranteed that it will work as expected.”
It was also added on MVSEP as "Apollo MP3 Enhancer":
Demo: https://mvsep.com/result/20240919224117-f0bb276157-mixture.wav
Advice
Input has no hard cutoff (quality slowly degrades toward high freq).
Generated output is as expected. It can fill holes, and it can remove artifacts (and probably bleeding too) and is working great with highly degraded audio here. If trained on clean source vs separated stem, which is not as much degraded content than 32kbps mp3 like previous example, I think it could work really great” [bad use case]
“so far the overlap magic is needed, cause u hear the transition”
“It seems to alter the tempo. It's not a constant alteration, it just shifts stuff, and you can't invert” becruily
> “I've seen this too in my tests, but it seems to happen only at the end of the chunk.
In the updated version (in which the end of chunks is ditched), I haven't seen that issue again.
>Overlap feature added [to the Colab].
New inference.py created for easier local CLI use.
I have set chunk_size at 3 seconds as default in the Colab because it was the chunk_size used to train the model, but it seems that the highest is the best.” jarredou
It was also added into ZFTurbo’s training dataset (read more).
Also, for non-mp3 input files, you might want to experiment with compressing them to 64kbps first.
- Also, TS-BSmamba2 was added to the repo. So it's available for training now too. But currently it works only on Linux.
- Aufr33 added MDX HQ4 to x-minus/uvronline via this link: https://uvronline.app/ai?hp&test-mdx
- (x-minus/uvronline.app) “viperx has updated the piano model!
I just replaced it on the website.” Aufr33
“The new piano model is incredible, I have even been able to separate a harpsichord by passing it over and over again through the model until the other instruments are left alone and it doesn't sound bad at all.”
There was some update to MVSEP piano models lately too, and there are SCNet and viperx models and ensemble with metrics added on the website (at least on separate page beside multi song dataset chart).
“both similar but mvsep has a teeny bit more bleed during the choruses and whatnot”
- (MVSEP) “I added possibility to use Bas Curtiz’ MelRoformer model with large score on Synth dataset. You must choose it from MelRoformer options. By default, my model is used.
The problem with Bas's model is that it's very heavy and slow, with almost the same score on Multi dataset.” Aufr33
“I've tried some songs and have great result! Music sounds fuller than original Kim's one & the finetuned version from ZFTurbo. Even [though] the SDR is smaller than BS Roformer finetuned last version, but almost song has the best result in instrumental.
1 song I found is bad result is from Wham - Where did your hearts go. The trumpet or sax whatever sound was lost, the model detects it as vocal, and the 1st beginning of vocal still heard. On other mel roformer, that trumpet or sax sound can still separate it as well.” Henry
- (MVSEP) “Guitar model was updated. I added BSRoformer model by viperx with SDR: 7.16. And
- I replaced [guitar] Ensemble. Earlier it was MDX23C + Mel. Now its BS + Mel. SDR increased from 7.18 to 7.51.
Demo: https://mvsep.com/result/20240914110542-7ab0356600-song-000-mixture.wav
All these models are available for all users.” ZFTurbo
- The new MDX HQ5 beta model is now online!
Use this link to access it:
https://uvronline.app/ai?hp&test-mdx - link for premium users
Go to "music and vocals" and there you will see it (scroll down).
It's not a final model yet, and the model is in training from April and is still in progress.
It seems to be muddier than HQ_4 (and more than Kim’s and MVSEP’s Mel-Roformer), it has less vocal bleeding than before, but more than Kim Mel-Roformer. Sometimes struggles with reverb.
"Almost perfectly placed all the guitar in the vocal stem" it might get potentially fixed in the final version of the model, which is planned for release in the mid-November as Anjok said at 04.11.24.
The model is not available in UVR yet (only on uvronline.app)
- Using the UVR Roformer beta patch for Mac doesn’t allow you to choose the Roformer parameter to check for manually copied Roformer models to UVR like: Kim Mel-Roformer or unwa’s Roformer, and only config name can be chosen, but no confirm button is available to make the model work. Place corresponding hash-named file to models\MDX_Net_Models\model_data after placing model file to MDX_Net_Models and non-hased model’s yaml to mdx_c_configs and start the UVR.
- Aufr33 released files for the new UVR de-reverb model made with jarredou
(based on VR 5.1 arch).
“1. Download this and unzip into your Ultimate Vocal Remover folder
2. Select VR architecture and DeReverb model from the menu
3. Set the parameters as shown here”
(PS: Dry, Bal: 0, VR 5.1, Out:32/128/Param: 4band_v4_ms_fulband -
An already existing json config file in modelparams folder has the same checksum)
Bas Curtiz’ “Conclusion so far:
- MDX[23C] De-Reverb seems to be cleaner, takes the reverb away, also between the words,
whereas VR leaves a little reverb
- [The new] VR De-Reverb seems to sound more natural, maybe therefore actually.
Also, MDX tends to 'pinch' some stuff away to the background, which sounds unnatural.
This is just based on my experience with 3 songs/comparisons, but both points are a pattern.
Overall, they're both great when u compare them against the original reverbed/untouched vocals.” Video
- SCNet Large vocal model on MVSep published.
Multisong dataset:
SDR vocals: 10.74
SDR other: 17.05
“just like the new bs roformer ft model, but with more bleed. [BS] catches vocals with more harmonies/bgv” isling
- Cyrus repaired pip issues with Medley Vox Colab
- Aufr33 released MDX23C de-reverb model files
https://a19p.uvronline.app/public/dereverb_mdx23c_sdr_6.9096.ckpt | config
“If you will use this model in your project, please credit us (me and jarredou)”
Also added on MVSEP.
UVR instruction:
“1. Just copy model to Ultimate Vocal Remover\models\MDX_Net_Models
2. Copy .yaml config to Ultimate Vocal Remover\models\MDX_Net_Models\model_data\mdx_c_configs
3. When opening UVR, selecting dereverb_mdx23c_sdr_6.9096 from the MDX-Net process method, don't click 'RoFormer model' cause it's not.
4. Select config_dereverb_mdx23c from the dropdown. Done.” ~Bas Curtiz
5*. In case of “no key” error in UVR, changed line 30 in the config to:
No dry
But it doesn’t happen to everyone.
- New UVR Dereverb model added on uvronline.app for premium users.
It seems to handle room reverb better than the previous MDX23C model, and the Foxy’s model sometimes cut “way too much” than this new model.
- People cannot separate using Ripple since longer than August 12th. There's an error "couldn't complete processing please try again"
- (x-minus.pro/uvronline.app) “Hipolink was a temporary solution. Now I can accept payment via Patreon as well.” Aufr33
- MDX23C De-reverb model by Aufr33 released for premium users of uvronline.app.
“Thanks to jarredou for helping me create the dataset”
- Jarredou released v. 2.5 of MDX23 Colab adding the new Kim Mel-Roformer model. Final SDR is higher (17.64 vs 17.41 for instrumentals, with 2024.08.15 MVSEP Ensemble being 17.81).
“Baseline ensemble is made with Kim Melband rofo, InstVocHQ and selected 1296 or 1297 BS Rofo” switching from 1296 to 1297 produces more muddy/worse instrumentals in this Colab (more sudden jumps of dynamics from residues).” VitLarge is no longer used by default.
- unwa’s fine-tuned BS-Roformer model released (12.59 for instr) - worse SDR than other fine-tuned models on MVSEP by ZFTurbo, but better SDR than Kim’s MelRoformer and viperx base model https://drive.google.com/file/d/1Q_M9rlEjYlBZbG2qHScvp4Sa0zfdP9TL/view
- Mel-RoFormer Karaoke / Lead vocal isolation model files released by Aufr33 and viperx
“If you will use this model in your project, please credit us” (download)
UVR instructions. Be aware that online version on uvronline/x-minus seems to work better.
- doubledouble.top will be soon replaced by https://lucida.to/
- Kimberley Jensen released her Mel-Band Roformer vocal model publicly (download)
(simple Colab/CML inference/x-minus/MVSEP/jarredou Colab too now)
Works in UVR beta Roformer (model | config - place the model file to models\MDX_Net_Models and config to model_data\mdx_c_configs subfolder and “when it will ask you for the unrecognised model when you run it for the first time, you'll get some box that you'll need to tick "roformer model" and choose it's yaml”.
Use overlap 2 for best SDR, or 8 for faster inference in UVR)
- SCNet model published on MVSEP. Similar metrics to MDX23C model, but seems to leave lot of vocal residues.
“ it is based on SCNet-small config from the paper, the SCNet-large config is almost 1 SDR above in the reported eval, so hopefully, next SCNet model trained by ZFTurbo with that large config will be better too.” So far he had some problems training on large config, sadly.
- Slightly better Roformer 2024.08 model (0.1 SDR+) was added on MVSEP
vs 2024.04 model “it seems to be much better at taking the vocals when there are a lot of vocal harmonies.”
- (x-minus.pro/uvronline.app) “In the new interface, the BS-RoFormer model now also has De-mudder
Select the Music and vocals, BS-RoFormer and after processing you will see the De-mudder button appear.” Aufr33
It works for premium.
- If you got an error while using jarredou’s Drumsep Colab (object is not subscriptable):
change to this on line 144 in inference.py:
if type(args.device_ids) != int:
model = nn.DataParallel(model, device_ids = args.device_ids)
(thx DJ NUO)
- (GSEP) I received email about deletion of my files on one of my accounts which is inactive, if I don’t buy premium (haven’t received it on my main account with premium), so it's probably due to inactivity and no premium. It’s probably for accounts not using the service since the release of the new paid site and/or maybe didn’t have premium since then (email is from July 9th, so 3 months after the release of the new site, so possibly your files can get deleted after 3 months after premium was disabled on your account). Normally, new separations for free users are deleted after 3 days now, but older files were preserved at least for accounts using beta till now. The account wasn’t used since the end of October 2022.
Check your mailbox to ensure, I didn’t find that mail in spam on the main account with premium, so hopefully it’s not for everyone (at least not for those with premium or who used the site since the last 3 months):
“All files from Gaudio Studio will be deleted on August 7, 2024 [Wednesday]. (...) If you purchase a Studio Plan, your files will be preserved.”
Be aware that they function in Japan, which is GMT+9, so it’s 6 hours sooner than CEST (Warsaw, Skopje, Zagreb).
If you currently have premium, you can download all your previous separations in WAV without any charge (at least that’s how it used to be), without premium it says (misleadingly, I assume) “The song processed in the beta service do not support WAV file downloads.” but probably you’ll be able to do that if you buy premium if nothing has changed. It’s no loner possible, and there are no references to WAV in dev tools as before.
- Aufr33 released files of his Mel-Roformer de-noise models publicly:
Less aggressive & More aggressive | yaml file
“If you will use this model in your project, please credit me”
Added in jarredou Colab too (and on x-minus.pro/uvronline.app for premium users and MVSEP).
Both models work in UVR too (don’t forget setting overlap to 2 to avoid stem misalignment issues like for other Roformers in UVR Roformer beta, overlap 3 or above will break separation)
- Jarredou released manual ensemble Colab with drop-down menus (based on ZFTurbo code)
- To fix issues with BS variant of anvuew’s de-reverb model in UVR “change stft_hop_length: 512 to stft_hop_length: 441 so it matches the hop_length above” in the yaml file. It doesn’t happen on (thx lew).
If that line is not present in your model config go to the settings, then choose MDX In the advanced menu, then click the "clear auto-set cache" button.
Then go back to the main settings, click "reset all settings to default" and restart the app (thx santilli_).
- If you still have error on every attempt of using GPU Conversion in UVR on AMD GPU (you might potentially use outdated drivers and/or Windows), go to Ultimate Vocal Remover\torch_directml and replace DirectML.dll from C:\Windows\System32\AMD\ANR (make backup before). Experimentally, you can use this older 1.9.1.0 version of the library. Restart UVR after replacing the file!
Be aware that results achieved without GPU Conversion that way, at least on certain configurations, might have noisy static instead of bleeding in less noisy parts of stems vs when using only CPU (basically, MDX noise can be somehow different on GPU and denoise standard only alleviates the issue to some extent, and you need to use Denoise Model option to get rid of this noise, or better solution - min spec manual ensemble of denoise disabled result and denoise model to get rid of more noise. Aufr’s Mel-Roformer minimum denoise works worse for it.
- GSEP introduced a new model called “Vocal Remover” dedicated for vocal extraction and is only used for vocals, instrumental stem still uses the old model. Might be good at extracting SFX as well. (becruily/wancite)
- (uvronline.app) Mel-Roformer De-noise released for premium users.
“This model is optimized for music and vocals. You can choose between two aggressiveness settings:
minimum - removes fewer effects such as thunder rolls
average - usually removes more noise”
“The new model works as good as my UVR De-noise model, or even better.”
- drumsep model by aufr33 and jarredou added on MVSEP and uvronline.app too
- Not Eddy’s multi-arch Colab released in form of UI (like in e.g. KaraFan)
https://colab.research.google.com/github/Eddycrack864/UVR5-UI/blob/main/UVR_UI.ipynb
In case of “FileNotFoundError: [Errno 2]” try other location than “input”, or other Google account in case of ERROR - mdxc_separator (helps for both).
- New Mel-Roformer de-reverb model by anvuew was released
https://github.com/ZFTurbo/Music-Source-Separation-Training/issues/1#issuecomment-2226805511
(to make it work with UVR, delete “linear_transformer_depth: 0” from the YAML file, copy the model to MDX_Net_Models and YAML config to model_data\mdx_c_configs)
Also added on MVSEP.
“I'm definitely hanging onto it. It reminds me of the equivalent dereverb mdx model, which I've always liked (when it works). The roformer model is cleaner in some ways, though slightly more filtered and aggressive.
Neither the roformer or mdx models respond to mono reverb. However, adding a stereo reverb on top solves that, especially with roformer.” (Musicalman)
“anvuew's models can remove reverb effect only from vocals. Old FoxJoy's model works with full track.”
- BS-Roformer -||- - a bit better SDR
https://github.com/ZFTurbo/Music-Source-Separation-Training/issues/1#issuecomment-2229279531
(To fix “The size of tensor a”... error with BS variant of anvuew’s de-reverb model “change stft_hop_length: 512 to stft_hop_length: 441 so it matches the hop_length above” in the yaml file.) thx lew
Added in the Colab:
- The below model has been added. Ensembles updated as well.
Some users report “bleed from some synths and bass guitar” “Some drums instruments are low volume on drums only. While mel roformer makes a good clean one” “On some parts it's almost like it doesn't separate anything for a few seconds and on some other parts, it's working just really great. The demucs one is way more stable when listening to individual model separations on the same song.” (or simply older ensemble)
- (MVSEP) I finished my drums models. Results:
MelRoformer SDR: 12.76
Demucs4 (finetuned) SDR: 12.04
Ensemble Mel + Demucs4 SDR: 13.05
for comparison:
Old Best Demucs4 SDR: 11.41
Old Best Ensemble SDR: 11.99
New models will be added on site soon.” ZFTurbo
For comparison, the Mel-Roformer available on x-minus trained by viperx has 12.5375 SDR.
- (for models trainers) “Official SCNet repo has been updated by the author with training code: https://github.com/starrytong/SCNet”
“ZF's script already can train SCNet, but currently it doesn't give good results”
https://github.com/ZFTurbo/Music-Source-Separation-Training/releases/
The author’s checkpoint:
https://drive.google.com/file/d/1CdEIIqsoRfHn1SJ7rccPfyYioW3BlXcW/view
“One diff I see between author config and ZF's one, is that dev has used learning rate of 5e-04 while it's 4e-05 in ZF config. And main issue ZF was facing was slow progress (while author said it worked as expected using ZF training script https://github.com/starrytong/SCNet/issues/1#issuecomment-2063025663)”
The author:
“All our experiments are conducted on 8 Nvidia V100 GPUs.
When training solely on the MUSDB18-HQ dataset, the model is
trained for 130 epochs with the Adam [22] optimizer with an initial
learning rate of 5e-4 and batch size of 4 for each GPU. Nevertheless,
we adjust the learning rate to 3e-4 when introducing additional data
to mitigate potential gradient explosion.”
“Q: So that mean that you have to modulate the learning rate depending on the size of the dataset?
I think it's the first time I read something in that way.
A: Yea, I suppose because the dataset is larger you need to ensure the model sees the whole distribution instead of just learning the first couple of batches”
- jarredou/frazer
SCNet paper: https://arxiv.org/abs/2401.13276
On the same dataset (MUSDB18-HQ), it performs a lot better than Demucs 4 (Demucs HT).
“Melband is still SOTA cause if you increase the feature dimensions and blocks it gets better
you can't scale up scnet cause it isn't a transformer. It's a good cheap alt version tho”
Still, it might potentially give interesting results when training will be mastered to the point when e.g. SDR will be in pair with at least MDX-Net models as they can still be better than Roformers for instrumentals in many cases (e.g. MDX-Net tend to have less muddy instrumentals - every arch can have its own unique sound characteristics and might be potentially useful for ensembling).
- (jarredou) “I've released the Drums Separation model trained by aufr33
(on my not-that-clean drums dataset).
Stems: kick, snare, toms, hihat, ride, crash
It can already be used, but training is not fully finished yet.
The config allows training on not so big GPUs [n_fft 2048 instead of 8096], it's open to anyone to resume/fine-tune it.
For now, it's struggling a bit to differentiate ride/hh/crash correctly, kick/snare/toms are more clean.
Download
[attached config includes also necessary training parameters for training further using ZFTurbo repo]: https://github.com/jarredou/models/releases/tag/aufr33-jarredou_MDX23C_DrumSep_model_v0.1 (dead)
Use on Colab: https://colab.research.google.com/github/jarredou/Music-Source-Separation-Training-Colab-Inference/blob/main/Music_Source_Separation_Training_(Colab_Inference).ipynb” (dead)
https://colab.research.google.com/drive/1IC6Q1hLF55_tK6mhky0SWYKGVF9T5WsY?usp=drive_link
It works in UVR too. All models should be located in the following folder:
Ultimate Vocal Remover\models\MDX_Net_Models
Don't forget about copying the config file to: model_data\mdx_c_configs
The model achieved much better SDR on small private jarredou's evaluation dataset compared to the previous drumsep model by Inagoy which was based on a worse dataset and older Demucs 3 arch.
The dataset for further training is available in the drums section of Repository of stems/multitracks - you can potentially clean it further and/or expand the dataset so the results might be better after resuming the training from checkpoint. Using the current dataset, the SDR might stall for quite some amount of epochs or even decrease, but it usually increases later, so potentially training it further to 300-500-1000 epochs might be beneficial.
“I’ve had models where SDR changes by 0.01 but fullness/bleedless change with 10-15 points, I wouldn’t trust it that much” - becruily
Current model metrics:
“Instr SDR kick: 18.4312
Instr SDR snare: 13.6083
Instr SDR toms: 13.2693
Instr SDR hh: 6.6887
Instr SDR ride: 5.3227
Instr SDR crash: 7.5152
SDR Avg: 10.8059” Aufr33
And if evaluation dataset hasn't changed since then, the old Drumsep SDR:
“kick : 13.9216
snare : 8.2344
toms : 5.4471
(I can't compare cymbals score as it's different stem types)” - jarredou
After initial jarredou’s training in Colab, Aufr33 decided to train the model for additional 7 days, to at least above epoch 113 (perhaps around 150, it wasn't said precisely), while using the same config, but on a faster GPU (2x 4090).
Even epoch 5 trained on jarredou's dataset casually in free Colab (which uses Tesla T4 15GB with performance of RTX 3050, but with more VRAM) with multiple Colab accounts and very light and fast training settings, already achieved better SDR than Drumsep:
“epoch 5:
Instr SDR kick: 13.9763
Instr SDR snare: 8.4376
Instr SDR toms: 6.7399
Instr SDR hh: 0.7277
Instr SDR ride: 0.8014
Instr SDR crash: 4.4053
SDR Avg: 5.8480
epoch 15:
Instr SDR kick: 15.3523
Instr SDR snare: 10.8604
Instr SDR toms: 10.3834
Instr SDR hh: 4.0184
Instr SDR ride: 2.7248
Instr SDR crash: 6.1663
SDR Avg: 8.2509”
Don't forget to use already well separated drums (e.g. from Mel-Roformer for premium users on x-minus) from well separated instrumental as input for that model, or Jarredou MDX23 Colab fork v. 2.4 or MVSEP 4/+ ensemble (premium).
Purely for drums separation from even instrumentals, the model might not give good results. It was trained just on percussion sounds and not vocals or anything else.
Also, e.g. the kick and toms might have a bit of weird looking spectrograms. It’s due to:
“mdx23c subbands splitting + unfinished training, these artifacts are [normally] reduced/removed along [further] training.” Examples
- BTW, just for inference (separation), “ONNX and Demucs models don't work with multi-GPU”
- In the “experimental” section of MVSEP, there’s been added a new multispeaker model at the bottom.
E.g. it can work well splitting rapping and singing overlapped in the same, previously well separated vocal stem, but:
“It works more or less ok on my validation [5 quite different "songs"], but it's a disaster on real data. I opened it for everyone, but don't expect really good results” ZFTurbo
- Also, there has been added a new multichannel section in “experimental” it’s just for songs with 3 or more audio channels like e.g. Dolby Atmos (FLAC/WAV input supported). It’s just BS-Roformer and there’s “no reason to process stereo tracks with it”. Also, the original sample rate of the input file is preserved here.
- One of MVSEP’s GPU died recently, so the separations will be probably slower than usual.
- jarredou updated his AudioSR Colab. Now “each processed chunk is normalised at same LUFS level (fixes the volume drop issue)” plus “input audio is resampled accordingly to 'input_cutoff' (instead of lowpass filtering)”
Now also some errors associated with mono files are fixed.
- New drums model available on x-minus.pro
“SDR is: 12.4066.
Thanks to @viperx for model training! The model is trained on 995 songs. A small number of my pairs were included in the dataset.” Aufr33
Very positive reviews so far.
- New guitar model added on MVSEP
“Previous old model mdx23c: 4.87
New mdx23c model: 6.34
New MelRoformer model: 6.91
Ensemble MDX23c + MelRoformer: 7.10
Extract vocals and after apply ensemble MDX23C + MelRoformer: 7.28”
- If x-minus.pro site doesn’t work for you, use the clone instead:
- Demudder on x-minus was updated on 13.06 (cosmetic differences)
- “Some interesting updates to SL 11
https://www.youtube.com/watch?v=2BoEgBGiafM”
Seemingly, separation features got better. Coming on 19th June.
Their new algo was evaluated, and SDR is a bit worse than htdemucs 4 stem non ft model.
Every stem has some bleed, vocals are decent, and actually have better SDR than Demucs_ft. GPU processing in options has low utilization and is slow, they say it’s planned to be fixed in patch. 16GB VRAM recommended at least while using brass and saxophone. Around 18 models can be used in total.
Unmix Mulitple Voices is for speech case, not for singing case.
Unmix Drums option can serve for further separation of drums
“the residual kick/snare problem is much better, but the cymbal split does still contain bleed from the rest of the song sadly” [vs drumsep] - jasper waffles
- Multi-arch Colab by Not Eddy
incorporates: MDX-Net, MDX23C, Roformers (incl. 1053), Demucs, and all VR models, YouTube support and batch separation. It uses broken overlap from OG beta UVR code. Use the one below for just Roformers and now also 1053 instead:
- New Mel-Roformer De-Crowd model released on MVSEP. It slightly surpassed SDR of the previous MDX23C model.
It's also available publicly in the repository below:
https://github.com/ZFTurbo/Music-Source-Separation-Training
To use it in UVR, Go to UVR\models folder, and paste that there.
Then change "dim_t" value to 801 at the very bottom of: “model_mel_band_roformer_crowd.yaml” in mdx_c_configs subfolder. Don’t use overlap above 4.
- Drums Roformer model shared publicly by Yolkis
https://github.com/ZFTurbo/Music-Source-Separation-Training/issues/1#issuecomment-2156069553
Not totally bad results as for 7.68 SDR, but it was trained on subpar GPU for Roformers for only 5 days. To use it in UVR, delete linear_transformer_depth line in the config.
- (x-minus.pro) “The new Strings model by viperx has been added!” based on Mel-Roformer arch.
Good results reported so far.
Sometimes it can pick up brass.
- (x-minus.pro) “Demudder has been added!
This only works with the mel-roformer inst/vocal model. You need premium to use it.”
It works only for instrumentals. Vocals are unaffected. It fills holes in the spectrum, basing on both vocals and instrumental stems (e.g. it won’t serve to just recover lossy mp3).
The option shows after you uploaded/processed a track (at least with mel-roformer model).
It’s capable of providing better results than max_mag a.k.a. BS and Mel Roformer ensemble (premium), depends on a song.
SDR-wise, it’s not much worse than original model results (16.48 vs 17.32).
- “New VST for real time source separation (probably same models [like in] MPC stems)
https://www.youtube.com/watch?v=0Js5bWQWY7M
https://products.zplane.de/products/peelstems/”
- Mel-RoFormer de-crowd by aufr33 and viperx model files have been released publicly.
DL: https://buzzheavier.com/f/GV6ppLupAAA
Conf: https://buzzheavier.com/f/GV6psmJpAAA
“You can use ZFTurbo's code [check his GitHub] to run this model. If you use it in your project, please credit us”
To use it in UVR 5, change “the name of the model itself to the name of the YAML file.
This model only works the best when at 2 overlap, since anything higher than that it'll stop isolating parts of the song entirely.
Or else, you can also check out setting “inference.dim_t” parameter at the bottom of the yaml file to 801. “Leaving dim_t at 256 (2.5seconds) makes the model only usable with overlap=2 (2 seconds) with current beta code. Higher value will result is missing/non-processed parts.” - jarredou
“The Roformer model does a better job at retaining the instruments and vocals as well as some sound effects and synths better than the MDX-NET decrowd, but at the cost of crowd bleed. While the MDX-NET decrowd model does a better job at removing most of the crowd at the cost of instrumental bleed into the crowd stem.
Sometimes [the old model] mistakes the fuzzy sounds of guitar as crowd noise
Also isolates some kicks in songs” - Kashi
“For really difficult live songs (where the crowd is overwhelmingly loud to the point where you can't hear the band properly) sometimes filtering vocals with mel roformer on xminus THEN running the vocals stem through the mdx decrowd model sometimes helps” - isling
- (x-minus) “The new Wind / Saxophone model has been added! It completely replaces the old UVR model [on the site]. Thanks to viperx for model training.”
“Really great model! Big step since the last VR winds model.” It works better for brass instruments than wind.
- (x-minus) “BS-RoFormer Bass model added! This is a model by viperx.” Aufr33
”much better at treble-heavy bass tones than demucs”
It’s different from the latest MVSEP bass model. Viperx’ model might be cleaner, but pick up less bass at times. Both are improvement over demucs_ft ~drypaintdealerundr
It’s best in no piano stem.
- (x-minus) “Piano beta model added! Thanks to viperx for model training.”
- (x-minus) “Now all models except BVE are available for free, even without registration!
The only restrictions:
Only mp3 downloads are available
No ensemble
10 minutes of audio per day (past 24 hours). This is enough for testing 2-3 songs.
Max song duration is 8 minutes
It's not available through Tor, and it's not available in some countries.” Aufr33
Q: Why BVE models are excluded? [from free option]
A: “Because in the free version, wav files are deleted immediately after processing is complete. This makes it impossible to download some stems. In addition, this model necessarily uses MDX for preprocessing, which is very compute-intensive.”
- (mvsep) Bass model is online. The metrics:
Single models:
HTDemucs4 bass SDR: 12.5295
BSRoformer bass SDR: 12.4964
MelRoformer bass SDR: 11.86
MDX23C bass SDR: 11.20
Models on site:
HTDemucs4 + BSRoformer Ensemble (It's available on site as MVSep Bass (bass, other)): 13.25
Ensemble 4 stems and All-In (from site): 13.34
For comparison:
Ripple lossless (bass): 13.49
Sami-ByteDance v1.0: 13.82
- GSEP announced works on a new model
- Mel-RoFormer Karaoke model added on x-minus.pro
“one of the cleanest lead vocals result[s]”
“I noticed that the new karaoke model considers vocals as lead vocals, even if they are quite wide. In other words, it has a much larger tolerance for vocal width than other karaoke models. This means that backing vocals that sound almost centered can be removed along with the lead vocal. If I apply a stereo expander, the model produces more adequate results. So when I add the Lead vocal panning setting, the "center" will actually work as "stereo -20%" (for example).”
“Q: wouldn’t this mean that there will be more backing vocal bleed in the lead vocal stem too?
A: The model behaves differently. In some cases, it completely isolates vocals, in other cases it gets confused and vocals appear in both stems at once, in other cases it doesn't isolate at all.”
Q: What are the differences between mel-roformer karaoke and the last model?
A: “If the vocals don't contain harmonies, this model (Mel) is better. In other cases, it is better to use the MDX+UVR Chain ensemble for now.”
Although your mileage will still vary on a song, e.g. “For most of the songs I tried it worked very well. Example: "From Souvenirs to Souvenirs" by Demis Roussos. It's the only model as of now which can seperate the lead from the back vocals correctly” dca100fb8
- Izotope RX11 officially released. Seeing by the unencrypted onnx model file names, it uses demucs_ft for stem separation now, but maybe it’s not the same model as the public one, as all stems “null back to the input stereo which is something standard demucsht doesn't appear to do (...) mdx seems to always null luckily.” The feature still doesn’t use CUDA/Nvidia GPU for processing (and there’s no such option anywhere). It’s still an improvement over RX8-10 as they used in-house Spleeter 22kHz models before.
https://www.youtube.com/watch?v=MhUEmvneerc
New features (e.g. clean up dialogue in real time).
- Logic Pro 11 now incorporates a Stem Splitter. Results vary from good to bad (and worse than known solutions) depending on a song.
- Mel-Roformer De-crowd model added on x-minus.pro.
Results are more accurate than in the old MDX model.
- GSEP has been updated.
Free option for all stems has been removed. There's only a 20 minutes free trial. WAV is only for paid users.
Vocals and all other stems (including instrumentals/others) are paid, and length for each stem is taken from your account separately for each model.
No credit is not required for the trial.
For free, only mp3 output and 10 minutes input limit.
For paid users there's a 20 minutes limit, and mp3/wav output, plus paid users have faster queue, shareable links, and long term results storage.
7$/60 minutes
16$/240 minutes
50$/1200 minutes
Seems like there weren't many changes in the model (if there weren’t even more vocal residues introduced since then). People still have similar complaints to it. Comparison video.
There was an average of 0.13 SDR increase for mp3 output and first 19 songs from multisong dataset evaluation, but judging by no audible difference for most people, they could simply change some parameters for inference.
The old files from previous separations on your account didn't get deleted so far.
- (x-minus) max_mag of (?-)Roformer and Demucs (drums only) added
“now the synths and everything else feels muddy
noticed the drums in some places (mainly louder-ish bits) sound a bit weird
mostly lower end like bass drum instead of hi hats
great improvement overall” isling
- Doubledouble might have some occasional hiccups on downloading. If you encounter very slow download, don’t attempt retrying the same download, but generate a new download query. Do it even three times in a row if necessary or wait half an hour and retry. Also, you can check the option to upload your result on external hosting.
- (x-minus) “Added max_mag ensemble for Mel-RoFormer model! It combines Mel and BS results, making the instrumentals even less muddy, while better preserving saxophone and other instruments.”
- New Mel-Roformer model trained by Kimberley Jensen on Aufr33 dataset dropped exclusively on x-minus.
“This model will now be used by default and in ensemble with MDX23C (avg).”
It’s less muddy than viperx model, but can have more vocal residues e.g. in silent parts of instrumentals, and can be more problematic with wind instruments putting them in vocals, plus it might leave more instrumental residues in vocals.
“godsend for voice modulated in synth/electronic songs”
SDR is higher than viperx model (UVR/MVSEP) but lower than fine-tuned 04.24 model on MVSEP.
- New UVR patch has been released. It fixes using OpenCL on AMD and Intel GPUs (just make sure you have GPU processing turned on in the main window and (perhaps only in some cases) OpenCL turned on in the settings).
Plus, it fixes errors when the notification chimey in options is turned on.
https://github.com/TRvlvr/model_repo/releases/download/uvr_update_patches/UVR_Patch_4_14_24_18_7_BETA_full_Roformer.exe (be aware that you can lose your current UVR settings after the update)
To use BS-Roformer models, go to download center and download them in MDX-Net menu (probably temp solution).
For 4GB VRAM and at least AMD/Intel GPUs, you can try out segments 32, overlap 2
and dim_t 201 with num_o 2 (dim_t is at the bottom of e.g. model_bs_roformer_ep_368_sdr_12.9628.yaml) to avoid crashes.
You might want to check a new recommended ensemble:
1296+1297+MDX23C HQ
Instead of 1297 and for faster processing and similar result, make a manual ensemble with a copy of 1296 result instead. It might work in similar fashion like weighting in 2.4 Colab and model ensemble on MVSEP (source).
- VIP code allowing access to extra models in UVR currently doesn’t work using Roformer beta patch older than #10, and MDX23C Inst Voc HQ 2 models disappeared from download center and GH. You can try to download VIP model files manually from this link and place them in Ultimate Vocal Remover\models\MDX_Net_Models directory:
Of course, it’s not all. E.g. 390 340 models and old beta MDX-Net v2 fullband inst models epochs are not reuploaded. This situation might cause errors on an attempt of using Inst Voc HQ 2 in AI Hub fork of Karafan.
Decrypted VIP repo leads to links which are offline, and also it doesn’t contain all models. Possibly the only way to access all the VIP models in beta UVR, is to roll back to stable 5.6 version from UVR official repo, and after downloading all desired VIP models, update to the latest patch.
- According to their forum leak, iZotope RX11 might be released between May and July, and contain some “pretty big changes”, among others, a novel arch for separation is rumored, and a lot of options reworked. (cali_tay98)
Official announcement is out:
https://www.izotope.com/en/learn/rx-11-coming-soon.html
(overhauled repair assistant, real time dialogue isolation for better separation of noise and reverb from voice recording)
- GSEP announced an update on May 9th with a WAV download option and redesigned UI.
The site will be unavailable on 8th May.
Noraebang (karaoke) service “due to low usage” will be shutdown, and your separated files deleted (you can make a backup of your files before).
Paid plan will be offered with faster processing times and “additional features”.
No model changes are announced so far. The update schedule might change.
- MDX23-Colab Fork v2.4 is out. Changes:
“BS-Roformer models from viperx added, MDX-InstHQ4 model added as optional, FLAC output, control input volume gain, filter vocals below 50Hz option, better chunking algo (no clicks), some code cleaning” - jarredou
- (x-minus) “Added mixing of MDX23C and BS-RoFormer results (avg/bs-roformer option). So far, it works only for MDX23C.” Aufr33
- “Output has released a free AI based generator that create multitrack stem packs
https://coproducer.output.com/pack-generator”
12-seconds long audio, fullband, 8 stems (drums in one stem, electric and rhythm guitar, hammond organ, trumpet, vocals) with 8 variations
“this looks more like it's mixing different real instruments, rather than actually making up songs (like a diffusion based generator)” ~jarredou/becruily
- Ensemble on MVSEP updated
- The site is up and running after some outage
- ZFTurbo released fine-tuned viperx model (“ver. 2024.04”) on MVSEP (further trained from checkpoint on a different dataset). Ensembles will be updated tomorrow. Clicking issue has been fixed.
SDR vocals: 11.24, instrumental: 17.55 (from 17.17 in the base model)
Depends on a song if it’s better. Some vocals can be worse vs the previous model.
- Test out ensemble 1296 + 1143 (BS-Roformer in beta UVR) + Inst HQ4 (dopfunk)
Ensembles with BS-Roformer models might not work for everyone, use manual ensemble if needed.
- Viperx model added also to beta Colab by jarredou. It gives only vocals, so perform inversion on your own to get instrumental
https://colab.research.google.com/drive/1pd5Eonbre-khKK_gn5kQPFtB1T1a-27p?usp=sharing
Update: now BS-Roformer is also added in the newer v.2.4 Colab
- Viperx’ BS-RoFormer models have been implemented by Anjok to UVR
- BS/Mel-Roformer UVR beta patch
For GPU acceleration, UVR currently supports:
a) CUDA (NVIDIA GPUs)
b) DirectML (AMD and Intel GPUs; previously misnamed as OpenCL)
c) MPS (Mac M1 [ARM]/x86-64)
And even old CPUs are supported (AMD A6-9225 Dual-Core or Intel Core 2 Quad [models with SSE4.1 tested]), but for at least MDX-Net HQ (v2) models speeds will be good on such CPUs without GPU acceleration.
DirectML acceleration is not supported for Apollo, Bandit (also incompatible with MPS), SCNet and probably Demucs 2 archs - CPU will be used automatically.
Minimum reasonably good enough NVIDIA GPU for Roformers might be:
desktop RTX 3050 6GB, 2304 CUDA cores (at best 8GB variant with more 2560 CUDA).
E.g. Colab’s Tesla T4 11GB (older RTX 2000 gen) has 2560 (which without TTA [not implemented for Roformers in UVR] is just alright - see separation times).
The old 980 Ti 6GB with 2816 CUDA will be rather slower than these due to older architecture.
I'd refrain from getting a mobile RTX 3050 - the Ti variant has 2560 CUDA cores, but both Ti and regular have 4GB, and using shared memory is slower - Roformers generally use more than that.
In FP16 inference, if we “convert” CUDA cores performance to match the Blackwell:
RTX 5060 Ti 4608 | 4608 | 100%
RTX 4070 Ti 7680 | ~4800-5120 | ~105-110%
RTX 3080 8704 | ~3200-3500 | ~70-75%
Plus, you can always convert current Roformers to FP16 if some model uses FP32 (more below). Performance comparisons: 1 | 2 | 3 | 4
Using CUDA, even with currently the biggest 2GB size inst model, Rifforge, doesn’t reach 6GB VRAM usage on RTX 2000 (but its separation time on T4 using TTA in Colab takes a long 36 minutes for a 5 minute file), so consider min. 6GB VRAM as bare minimum for reasonable separation times with around 2560 CUDA cores on at least RTX 2000 (newer gens should have faster cores in inferencing). Unless you want to use the 53 stem model.
For AMD/Intel GPUs, if you want to use default high chunk_size with certain models, 16GB VRAM is recommended, so you won’t have to decrease it manually (DirectML is way less memory-efficient than CUDA).
Minimum AMD GPU capable for Roformers using DirectML might be RX 6600 XT or RX 7600 (might be even faster than 3050, DML will probably use its new RDNA3 AI cores).
DirectML is slow nevertheless, so if you want some serious GPU acceleration boost on AMD GPUs, use MSST/pymms with ROCm instead. DirectML gives tremendous overhead over CUDA. For older GPUs than AMD Vega (which doesn't support ROCm PyTorch or only potentially with ZLUDA with modded libs) or Intel GPUs with Vulkan support, instead of MSST use: https://github.com/chenmozhijin/BSRoformer.cpp (CLI, C++ ready build, model conversion required besides Deux and FV6, allows quantization for faster separation, Vulkan backend faster than ROCm 5+ZLUDA on Windows) or also: https://github.com/pymss-project/pymss-mnn (CLI, requires building and also model conversion, Vulkan backend too, potentially a bit slower)
Full list of model architectures supported by UVR:
MDX-Net (a.k.a. v2), MDX23C (archs by kuielab), VR (voice-remover by tsurumeso, versions: 4, 5 [UVR fork], and 5.1), Demucs (by Meta; v. 1-4, only models trained on OG code, not MSST ones), BS-Roformer, Mel-Roformer (arch by Bytedance & implementation from paper by lucidrains; issues on Linux explained later), SCNet, Apollo (in Tools; for upscaling, no DirectML acceleration), BandIt (SFX, no DirectML support).
It has also a feature of ensembling models of various archs. If some models don't appear on the ensemble list reach the appropriate section below (or rich the document outline in options).
Demudder functionality added in newer UVR patches (currently not on Linux and MacOS).
Models for Roformers/SCNet/Bandit arch are located altogether in the MDX-Net menu.
Apollo audio upscaler is located in Tools.
Don’t forget to enable GPU Conversion - if it works, it speeds up separation hugely. Even Tiger Lake iGPUs are capable of working with at least MDX-Net HQ (v2) models.
- Min. 4GB VRAM AMD GPUs tested/recommended for Roformers (with chunk_size between 112455 and 200655 depending on the model in the Roformer yaml; table of chunks).
- Min. 2GB VRAM NVIDIA GPU (even 2GB Rifforge worked with low enough chunks), but generally it can give memory allocation errors randomly and sometimes working or not with Roformers. Setting your NVIDIA Control Panel UVR profile to “CUDA Sysmem Fallback Policy” to “Prefer Sysmem Fallback” makes it crash less - mohammedmehditber
- Bare minimum for GPU acceleration on NVIDIA is Maxwell/900 series GPUs/compute compatibility 5 (unsupported for CUDA in UVR are at least NVIDIA GTX 600 and GT 700 series and older, returning: “AssertionError: "" Traceback Error:” or "CUDNN_STATUS_NOT_INITIALIZED), although DirectML should be theoretically supported by all DX12 GPUs (I’m not sure if in newer versions than 5.6.0 it’s still possible to switch to DirectML setting on NVIDIA GPUs).
- For AMD, at least RX 4GB models tested (not sure about R9 200 4GB GPUs - either if on newer modded Radeon-ID drivers and/or with downgraded DirectML.dll attached with drivers, copied to UVR\torch_directml folder, but seems like someone had occasional memory issues on HD 7870 2GB, but GPU Conversion still worked). “AssertionError: "" Traceback Error:” also exists when your AMD driver/Windows is outdated (then use e.g. 1.9.1.0 library).
- Intel was confirmed to work with ARC GPUs, and Xe integrated graphics (e.g. Tiger Lake 2021) with at least MDX-Net HQ (v2) models.
- If your separation on DirectML is stuck (AMD/Intel), and takes an enormously long time, decrease chunk_size.
- RTX 5000 series support on Windows was added in a separate UVR patch (or possibly you can use OpenCL (DirectML) in options instead [slower]). The patch is not compatible with Intel/AMD GPUs, and potentially also older NVIDIA GPUs, giving the following error:
“AttributeError: module 'torch_directml' has no attribute 'is_available’
Be aware that some newer models are incompatible with this version giving:
“torch._dynamo.polyfills.fx” error
- In patch #12 a new “Inference mode” option in Advanced MDX-Net>Multi Network was implemented (disabled by default). In the current state, it fixes silence in separation on GTX 10XX and maybe older, but might make separations longer for other compatible GPUs. So if you have slower separations after updating UVR, check if GPU Conversion is still enabled (it’s rather disabled by default on new installations) and you can try to turn on Inference Mode if you have RTX GPU, or potentially GTX 16XX.
- To potentially increase separation time (from 45 to 32 minutes) using DirectML (AMD/Intel) on Windows, you can try using vkd3d-proton -
Extract it with PeaZip (if 7-Zip returns an error), copy the two .dll libraries from the X64 folder near UVR.exe.
UVR download
(if your download speed gets slow, use e.g. “Free Download Manager” on Windows or any other increasing connection count).
- for Linux
Sadly, official building instructions on GitHub are outdated, and the current branch working on Linux with Roformers is old and doesn’t support all those models (at least without some workarounds below). Also, normally you can’t use WINE and use at least DirectML GPU acceleration (error), and I doubt VKD3D would help.
It turns out someone found a way to use GPU Conversion with Wine by using Bottles and used newer UVR Windows codebase in the packages linked later below (although it might be slower than MSST due to additional translation layer, so consider it instead, it worked well on 5060, but slow on 4090):
"My way to easily run UVR on Linux:
I just downloaded Bottles
https://usebottles.com/ (which uses WINE) and used the provided .exe file from the repository's releases. I created a new bottle with a gaming profile (to utilize GPU) and moved the exe file into "drive_c" (otherwise it won't work), then just ran through the installer and it worked like a charm!" (src)
If you still want to use native Linux codebase:
- Some installation directions (from our Discord) from Roformer patch #1/2 period (or also here), plus comment out:
segmentation-models-pytorch 0.3.3 in requirements.txt here (line 136) (for Nvidia/CPU).
Current code repository for Roformers with DirectML support is located here (although at certain periods it might lack current patches):
(might work for non-Nvidia GPUs):
https://github.com/Anjok07/ultimatevocalremovergui/tree/v5.6.0_roformer_add%2Bdirectml
Judging by the date of files from 9 December 2024 in the repo at the moment, they seem to derive from outdated 12_8_24_23_30_BETA beta #9 patch (some newer models will fail with that codebase, plus it’s before chunk_size implementation, so it rather still uses dim_t).
- Some other potentially useful information:
https://github.com/Anjok07/ultimatevocalremovergui/issues/1674
- Or “you just need to use export PYGLET_SHADOW_WINDOW=0 and it'll work”
- See here if you get following errors:
“Getting requirements to build wheel did not run successfully. ... ModuleNotFoundError: No module named 'imp'”
edit requirements.txt:
pyrubberband==0.3.0
PyYAML==6.0
scipy==1.9.3
playsound
numpy==1.23.5
- Workarounding issues with Python 3.12:
https://github.com/Anjok07/ultimatevocalremovergui/issues/1789
- Becruily karaoke model error fix:
“Those of you on linux running the current roformer_add+directml branch that cant get becruily's karaoke model working due to the same error: it seems editing line 790 in separate.py setting the keyword argument strict to False when calling load_state_dict seems to make the karaoke model load and infer properly, so I think it will work
model.load_state_dict(checkpoint, strict=False)
I don't know if this is a robust workaround, but I haven't observed anything behaving differently than it should yet, so if you want to give it a shot I think it will work
TL;DR change line 790 in separate.py to the codeblock and then run again and karaoke model should work” stephanie
ROCm instructions for AMD (also for Windows, but currently only using WSL)
- For better separation speed than DirectML:
https://github.com/Anjok07/ultimatevocalremovergui/issues/1822#issuecomment-2824363747
- Fix for” “ModuleNotFoundError: No module named 'audioread’”
https://github.com/Anjok07/ultimatevocalremovergui/issues/1797#issuecomment-3315191055
- Workarounding issues with playsound & sklearn dependencies:
https://github.com/Anjok07/ultimatevocalremovergui/issues/2107
- Fixing issues in Matchering on Linux:
“Succeeded to run after modifying UVR.py :
match.process(
target=target,
reference=reference,
results=[match.save_audiofile(save_path, wav_set=self.wav_type_set),],
)
to
match.process(
target=target,
reference=reference,
results=[match.pcm16(save_path)]
)
“
- Fixing matching errors
https://github.com/Anjok07/ultimatevocalremovergui/issues/2018#issue-3564850746
- Python 3.11 dependencies fix
https://github.com/Anjok07/ultimatevocalremovergui/issues/2108#issuecomment-3808617707
- natsort issues
https://github.com/Anjok07/ultimatevocalremovergui/issues/2214
__________
Below you’ll find download links for ready packages for other platforms
(they contain isolated Python environment - so you don’t have to mess with your local Python installation - you don’t even have to have Python installed on your computer)
- for macOS
(it doesn’t incorporate demudder from Windows patch #14)
- UVR Roformer beta patch #13.1
(beta_0115_MacOS_arm64_hf)
which applies a hotfix to address a few graphics issues.
- Mac M1 (arm64) users - Link
- Mac Intel (x86_64) users - Link (it went offline for some reason; older version below)
Note: Some people get error about failed malicious software check. Then check out patch #9 below or read “MacOS Users: Having Trouble Opening UVR?” section in the GH repo. Plus:
“Functionality for systems running macOS Catalina or lower is not guaranteed”.
Also: “What ended up working for me was to make sure that UVR lived in my root level Applications folder. I normally have an 'Installs' folder inside of Applications to help me keep track of things I have downloaded and installed. Nothing would load from the Installs folder, but worked fine from the root Applications folder.”
Older Mac packages:
UVR beta Roformer patch #13
VR_Patch_1_15_25_22_30_BETA:
- Mac M1 (arm64) users - Link
- Mac Intel (x86_64) users - Link
UVR Roformer beta patch #9:
UVR_Patch_12_8_24_23_30_BETA:
Mac M1 (arm64) users - Link
Mac Intel (x86_64) users - Link
(Patch #9 fixes mainly Apollo arch issues)
Apollo arch was made compatible with MacOS MPS (metal) but it might be unstable and very RAM intensive - use chunk size over 7 to prevent errors.
Apollo is now compatible with all Lew models (fixed incompatibility with any other than previously available in Download Center). Fixed Matchering (presumably regression).
Changelog for Mac since patch #2 (older patches later below):
“Roformer checkbox now visible for unrecognized Roformer models” so now you can use custom Roformer models on MacOS Roformer patch without copying/modifying configuration files from Windows version or other users. Plus it includes all the previous fixes in the previously released patch (overlap code fixed, so no stem misalignment should occur on certain overlap settings - higher overlap now means longer separation - so it's the opposite now)
- for Windows
- (optional) Standalone UVR Roformer beta patch #15 for RTX 5000 (or more precisely 14b) fixing issues with not working or slower than CPU CUDA on those GPUs.
It’s not compatible with older NVIDIA GPUs or AMD/Intel.
- Full Install: Download
(standalone, although still called patch just for simplification to refer codebase functionality as to digits hashtag number)
Note: Some newer models are exclusively not compatible with this version giving torch dynamo error (then you need to use the patches below with CPU or maybe slower DirectML; you can have both installations after manual backup).
- UVR Roformer beta patch #13 (for most modern GPUs, but RTX 5000 will be slow)
UVR_Patch_1_15_25_22_30_BETA
It fixes the issue with no sound on some Roformer models (like avvuew’s de-reverb) on GTX 10XX or older:
- Full Install: Link
(means it's standalone, and doesn't need 5.6.0 on top, or any other existing UVR installation to work)
- Patch Install (for non-beta UVR installed, e.g. 5.6, not Roformer 5.6.x patch): Link
- Small Patch Install (for any Roformer patch previously installed): Link
The issue was some older GPU's are not compatible with Torches "Inference Mode," (which is apparently faster) so it's now using "No Grad" mode instead. Users can switch back to using "Inference Mode" via the advanced multi-network options. More
UVR demudder
- UVR Roformer beta small patch #14 - the long anticipated demudder added:
UVR_Patch_1_21_25_2_28_BETA: Link
It’s a small patch (you must have a Roformer Installation [e.g. full #13 above] previously installed for this to work).
- Included in the standalone RTX 5000 patch above (download)
Furthermore, minor bugs were fixed, calculate compensation for MDX-Net v1 models added.
The MacOS version of the patch is not released yet.
To enable demudder, go to Settings>Choose Advanced Menu>Advanced MDX-Net Options>Enable Demudder (in Demudder options pick one out of the three described below)
Troubleshooting:
- At least Phase Rotate doesn’t work on AMD and 4GB VRAM GPUs on even 88200 chunk size (prev. dim_t 201 - 2 seconds) and 800MB Roformers like Becruily’s, while 112455 (2,55s, prev. dim_t = 256) works fine for normal separation.
- In case of file not found error on attempt of using demudder, reinstall UVR.
- In case of Format not recognised error for demudder, keep Match freq cut-off enabled in MDX settings.
- “For Roformer models, it must detect a stem called "Instrumental” so for some models like Mel-Kim, you need to open the model's corresponding yaml, and change “other” to “instrumental”.”
“With the new config editor feature you could probably edit the configs of models to have the vocal stem labelled as the Instrumental stem so the demudder demuds the vocal stem, it definitely still makes a difference
I accidentally did this when installing another model, but it seems to actually have an effect on vocal stems too” stephanie
Demudder consists of three methods to choose from:
- Phase Rotate
- Phase Remix (Similar to X-Minus) - “the fullest sounding, but can leave a lot of artifacts with certain models. I only recommend that method for the muddiest models. Otherwise, Combined Methods is the best” “I don't recommend using phase remix on the Instrumental v1e model. I recommend combined methods or phase rotate for models produce fuller instrumentals.” Anjok
- Combine Methods (weighted mix of the final instrumentals generated by the above). More in the full changelog.
PS. You can also use phase remix in SESA Colab
Demudder is “meant to solely target instrumentals. The vocals should stay exactly as before.”
“It works best on tracks that are spectrally dense (ex. Metal, Rock, Alternative, EDM, etc.)
I don't recommend it for acoustic or light tracks.
I don't recommend using it with models that emphasize fuller instrumentals (like Unwa's v1e model).” Anjok
“I've noticed with the few amounts of tracks I've tried, demudding can sometimes accentuate instances of bleeding or otherwise entirely missed vocal-like sounds”. More in the full changelog.
“I put the demudded instrumental in the bleed suppressor, and it sounds really good, almost noise free. I either do a bleed suppressor or a V1/bleed suppressor ensemble” gilliaan
“I found that Phase Remix also works well on pop and other genres of music, but it only works using the vocal models (Phase Remix and VOICE-MelBand-Roformer Kim FT (from Unwa) or
VOICE-MelBand-Roformer Kim FT 2 (by Unwa).” Fabio
“I do plan on adding options to tweak the phase rotation.
I also plan on adding another combination method that may work better on certain tracks.” - Anjok
If you set 64-bit float output in Options>Additional settings, the results might be slightly less muddy, but also in very big size.
Demudder can be also used in a Colab and x-minus.pro/uvronline.app (more)
OG Discord channel to follow for updates
Old
____________
(old) Potential fixing of RTX 5000 series CUDA acceleration for #14 patch -
now unnecessary since dedicated patch was released above
(although there’s no full success with the below, so also refer here)
Some people reported that following the steps below might still result in slow separations:
Probably you’ll be able to install required CUDA 12.8 and nightly PyTorch to fix the compatibility issue when following steps for manual installation (unfold it), so UVR won't use its own Python environment. In addition to the above link, use this repo with newer code, although for now, not the newest code is attached from the patches below... and all if you fix "Getting requirements to build wheel ... error" afterwards. Then you’ll have a bug causing no GPU Conversion option functional - then you need to use:
python.exe -m pip install --upgrade torch --extra-index-url https://download.pytorch.org/whl/cu118
instead of cu117 as in the instruction above (although for RTX 5090 it probably won’t work, and you can use this https://download.pytorch.org/whl/nightly/cu128 instead.
- Similar issue might occur when you don’t install "onnxruntime-gpu"
when the current is "onnxruntime" library (which does not support GPU)
- For Demucs UnpicklingError issue using manual installation:
modify the line 46 in demucs/states.py:
package = torch.load(path, 'cpu', weights_only=False)
- Sometimes the same happens for e.g. Becruily inst model (different arch). It’s also the indicator that the model file is corrupted and has wrong CRC (most likely wrongly downloaded) - redownload the model.
At least on the old 5.6.0 version, there's OpenCL in options instead of DirectML in newer patches (although it's the latter).
Alternatively, you can use MSST-GUI.
__________
Old Roformer patches
- Anjok released a new UVR beta Roformer patch #11 (Windows only for now)
It fixes 4 bugs: with VR post-processing threshold, Segment default in multi-arch menu, CMD will no longer pop-in during operations, and error in phase swapper.
More details/potential updates.
Standalone (for non-existent UVR installation)
UVR_1_13_0_23_46_BETA_full.exe
For 5.6 stable (so for non-beta Roformer installation)
UVR_Patch_1_13_0_23_46_BETA_rofo.exe
Small (for already existing Roformer beta patch installation)
UVR_Patch_1_13_25_0_23_46_rofo_small_patch.exe
- Patch #12 which is a hotfix for the 4 stem BS-Roformer model by ZFTurbo (trained on MUSDB)
UVR_Patch_1_13_0_23_46_BETA_rofo_fixed.exe (Windows only)
Users undergo some issues (no sound) with Mel-Roformer de-reverb by anvuew (a.k.a. v2/19.1729 SDR) since the latest UVR beta #11 or #12 update. Patch #10 works.
The issue seems to occur only on GTX 10XX series, and maybe older.
You should be able to use more than one UVR installation at the same time when one’s been copied before updating (potentially patch #10 will still work) or use MSST repo and/or its GUIs.
UVR Roformer beta patch #9:
UVR_Patch_12_8_24_23_30_BETA
Windows: Full | Patch (Use if you still have non-beta UVR installed) |
Small Patch (for this you must have a Roformer patch previously installed for this to work)
UVR Roformer beta patch #10
UVR_Patch_1_9_25_23_46_BETA_rofo_small_patch - Link
For now, only small patch for already existing beta Roformer installation above is available, and only for Windows.
If you have Python DLL error on startup, reinstall the last beta update using the full package instead, then the small installer from the newer patch.
Since beta version #10, UVR doesn't rely on 'inference.dim_t' value for Roformers anymore (if you were using edited "dim_t" value in yaml configuration files).
You have to edit audio.chunk_size instead if need it (e.g. for 4-12GB VRAM on AMD/Intel).
It’s located “In model yaml config file, at top of it, chunk_size is first parameter (...) you can edit model config files directly inside UVR now.” or in the new config editor in newer versions.
“Conversion between dim_t and chunk_size
dim_t = 801 is chunk_size = 352800 (8.00s)
dim_t = 1101 is chunk_size = 485100 (11.00s)
dim_t = 256 is chunk_size = 112455 (2,55s)
dim_t = 1333 is chunk_size = 587412 (13,32s)
[more values later below]
The formula is: chunk_size = (dim_t - 1) * hop_length)” - jarredou
Generally, to have the best SDR, use chunks not lower than 11s in the yaml for inference, which is usually training chunks value (rarely higher). Although, at times people get better results with 2,55s chunks, although some models behave worse than others with such small values.
“Most of the time using higher chunk_size than the one used during training gives a bit better SDR score, until a peak value, and then quality degrades.
For Roformers trained with 8sec chunk_size, 11 sec is giving best SDR (then it degrades with higher chunk size)
For MDX23C, when trained with ~6sec chunks, iirc, peak SDR value was around 24 sec chunks (I think it was same for vit_large, you could make chunks 4 times longer)
How much chunk_size can be extended during inference seems to be arch dependant.” - jarredou
Changelog #10:
Added SCNet and Bandit archs with models in Download Center, fixed compatibility with some newer Roformer models (prob. the Phantom center and 400MB small Unwa models, not sure yet), new Model Installer option added, model configuration menu enhanced, allowing aliases to selected models, added compatibility for Roformer/MDX23C Karaoke models with the vocal splitter, VIP code issue is gone, issues with secondary models options and minor bugs and interface annoyances are addressed, “improved the "Change Model Settings" menu. Now, any existing settings associated with a selected model are automatically populated, making it easier for users to review and adjust settings (previously, these settings were not visible even if applied).”.
“Unfortunately, SCnet is not compatible with DirectML, so AMD GPU users will have to use the CPU for those models.
Bandit models are not compatible with MPS or DirectML. For those with AMD GPU's and Apple Silicon, those will be CPU only.
The good news is those models aren't all that slow on CPU.” - Anjok
Changelog #9:
Apollo fixes: “Chunk sizes can now be set to lower values (between 1-6)
Overlap can be turned off (set to 0)”
Fix both for Apollo and Roformers: now 5 seconds or shorter input files no longer cause errors.
OpenCL was wrongly referenced in the UVR. It was actually DirectML all the way, and Anjok changed all the OpenCL names in the app into DirectML.
Changelog for all platforms:
Patch #3 fixed the issue with stem misalignment when using incorrect overlap setting for Roformers. Now it uses ZFTurbo code (also for MDX23C), meaning that now increasing overlap for Roformers will result in increasing separation times and potentially better SDR [the opposite of what it used to be in the previous beta Roformer patches #1 and #2]. Also, it “Fixed manual download link issues in the Download Center. Roformer models can now be downloaded without issue.”). Also, new Roformer models were added to Download Center, so you don't have to download them manually.
- UVR Roformer beta patch #8 for Win: full | patch | Mac: M1 | x86-64:
UVR_Patch_12_3_24_1_18_BETA
Apollo arch was made compatible with OpenCL too, but it might be unstable and very RAM intensive - use chunk size over 7 to prevent errors (currently it’s not certain that all models will work with less than 12GB of VRAM). At least in newer patches, it can straight up say that Apollo is not compatible with DirectML, and fall back to CPU mode.
Apollo is now compatible with all Lew models (fixed incompatibility with any other than previously available in Download Center). Fixed (presumably regression with) Matchering.
UVR Roformer beta patch #7 (full | patch) Win
UVR_Patch_12_2_24_2_20_BETA.
It introduces support for Apollo arch. The OG mp3 enhancer and Lew v1 vocal enhancer were added to Download Center. The arch is located in Audio Tools. Sadly, this arch cannot be GPU accelerated with OpenCL so AMD and Intel cards (you’re forced to use CPU which might be long).
Also, “Phase Swapper” a.k.a. Phase fixer for Unwa inst models was added to Audio Tools.
Roformer beta patch #6: M1 | x86-64
UVR Roformer beta patch #6: Win
UVR_Patch_11_25_24_1_48_BETA (standalone or patch - you can install it on stable 5.6 version already installed)
Fixes issues with viperx’ models.
And with it, a long anticipated MDX-Net HQ_5 model was released (available for older versions in Download Center too. Changelog
UVR_Patch_UVR_11_17_24_21_4_BETA_patch_roformer (Beta patch #5 for UVR, Windows only):
“- Fixed OpenCL compatibility issue with Roformer & MDX23C models.
- Fixed stem swap issue with Roformer instrumental models in Ensemble Mode.”
That patch is probably not standalone like patch #3, so have a previous UVR installation.
Roformer patch #4 for MacOS: M1 | x86-64
UVR_Patch_11_17_24_21_4_BETA (Windows: full | patch | changelog) beta patch #4
UVR_Patch_11_14_24_20_21_BETA_patch_roformer (beta patch #3 “requires an existing UVR installation” so either the previous beta Roformer patch above or stable 5.6 patch.
Full changelog.
UVR_Patch_4_14_24_18_7_BETA_full_Roformer | mirror (standalone Roformer beta patch #2, fixed OpenCL separation for AMD/Intel GPUs)
UVR_Patch_3_29_24_5_11_BETA_full_roformer.exe (older Roformer patch #1)
With the following issue fixed in the newer patch #2 above -
if you have playsound.py errors, disable notification chimneys in settings>additional settings, using OpenCL GPU acceleration (AMD) for BS-Roformer doesn’t work (or at least not for everyone)
Older Roformer patch #1/2 for MacOS (ARM only) got deleted from the Discord server
____________
(Roformer models are also added on MVSEP and x-minus.pro/uvronline.app and MSST-GUI, and inference Colab and MDX23 v.2.4/2.5 Colab)
Instructions for UVR Roformer patches and installing custom models
- Your current settings might be lost after patching your current UVR installation
- Applying newer UVR Roformer versions over some older UVR versions might cause errors on startup (e.g. Python39.dll) - then perform clean installation to fix the issue. Just make sure that after uninstalling UVR, nothing is inside the old UVR folder
- To perform clean installation of the latest UVR version, for now you need:
Windows:
Roformer full patch #13
(it's actually standalone, and not just a patch - it doesn't need 5.6.0 on top):
UVR_Patch_1_15_25_22_30_BETA: Link
And then install this small patch #14 (demudder and other fixes; it needs the patch #13 in order to work):
UVR_Patch_1_21_25_2_28_BETA: Link
(it’s not standalone)
(Only for) RTX 5000 full patch #14:
UVR_Patch_4_24_25_20_11_BETA: Link
(iirc uses newer PyTorch and CUDA; standalone)
MacOS:
- Roformer full patch #13.1 (standalone)
(beta_0115_MacOS_arm64_hf)
which applies a hotfix to address a few graphics issues:
- Mac M1 (arm64) users - Link
- Mac Intel (x86_64) users - Link (it went offline for some reason; older version #13 below:
Link)
Linux:
- Old patch #9 only for now (no demudder, some fixable issues with certain models):
https://github.com/Anjok07/ultimatevocalremovergui/tree/v5.6.0_roformer_add%2Bdirectml
- Or use bottles with gaming profile using WINE on Windows newer codebase (still utilizes GPU)
“used the provided .exe file from the repository's releases. I created a new bottle with a gaming profile (to utilize GPU) and moved the exe file into "drive_c" (otherwise it won't work), then just ran through the installer and it worked like a charm” (ref)
Tested on at least RTX 5060. But using MSST instead might be faster.
(for older patches and more info and troubleshooting refer here)
- Some newer Mel/BS models really require the newest Roformer patches (otherwise you’ll get MLP error).
- In the Download Center you’ll find some BS/Mel models, but not all. Refer to the full list of models here.
Installing custom Roformer models in UVR
(those unavailable in Download More Models a.k.a. Download Center)
- To install models from external sources (those unavailable in Download Center) since patch #10, a new “Install Model” option was added (e.g. in MDX-Net menu).
Click RBM on the models’ list to access the option and follow the instructions on your screen.
- If you use older patch or native Linux, or have “Model not installed” error, alternatively you can also copy the model file to models\MDX_Net_Models and the .yaml config to models\model_data\mdx_c_configs, then after choosing the model in the app, press yes to recognize the model, and wait a while.
Models for Mel/BS Roformer, SCNet and Bandit Plus/v2 are located in the MDX-Net menu.
- In “Set Model Type”, most Roformers will use Roformer non-v2 (and most are Mel).
For now, you should pick v2 Roformer type probably only for unwa 400MB experimental model (if you have lots of layers errors using Roformers, it means you picked v2 config unnecessarily).
In the older beta you need to check “roformer model” option when asked for configuration file and confirm (you cannot press confirm or check the option on the oldest Mac Roformer patch, the issue is explained below, and fixed in newer versions, or the fix for older Mac patch step by step)
- For most model archs mentioned above, you need both ckpt and yaml file for the model to work (if it’s not the same with any config you already downloaded before - some models share the same config file - e.g. one folder with models in the repository might have only one yaml)
- If you opened a yaml file to download it, but its content opened without page layout instead of starting the download, press CTRL+S or go to options of your browser and find the option called “Save As”.
Now you’ll have a txt extension, but we need it to be yaml instead of txt (otherwise UVR won’t detect it), so choose All files in Extension, and edit it there manually to yaml (or manually after downloading).
If you see the config, but also page layout (e.g. on HF) just press “RAW” or “Copy download link”, and paste it to your browser and follow the above (so then save it with CTRL+A and change extension to yaml)
Or else, you can just see the list of files in the HF repo - then just click the download icon near the desired yaml file and that's it.
- If you don’t see “ckpt” on the extensions list to choose the model during using Install model option, consider performing clean UVR installation from the patches above
- Manual model installation for e.g. Demucs models, which doesn’t have “Install model” option (or if you use patch before #10) - on an example of MDX/Roformers/SCNet/Bandit located in MDX-Net menu:
Misc
- You might want to decrease default chunk_size in Edit Model Param (or yaml) for AMD/Intel GPUs with VRAM lower than 16GB if you have memory errors with GPU Conversion enabled or your separation is stuck on e.g. 5% (read more in Common issues later below)
- How would I assign a yaml config to an Apollo model on the new UVR [patch]?
“1. Open the Apollo models folder
2. Drop the model into the folder
3. From the Apollo models folder, drop the yaml into the model_configs directory
4. From the GUI, choose the model you just added and if the model is not recognized, a pop-up window will appear, and you'll have the option to choose the yaml to associate with the model.” - Anjok
(iirc not used in UVR) - “batch_size = 2 (not less or more)” - “1 can lead to some clicks in output, while with batch_size>=2, there are no clicks. Clicks are obvious in low freq of log spectrogram”
edit. iirc clicks with batch_size=1 could have been fixed (at least were with MSST from which the inference code was implemented in newer UVR patches, but iirc it wasn’t used, and later the clicks were fixed in MSST).
- Segment size in the UVR UI does nothing for these Roformer models due to Advanced arch options being set by default to Segment Default which makes it being read from the yaml file. While "Segment_Default" in Advanced MDX-NET23 settings is checked, it will use the dim_t value from the bottom of the config. Simply dim_t is the segments. Although now chunk_size is used in newer patches instead.
- “The overlap value in yaml files is never used by UVR, only the value in GUI is used.”
“UVR uses inference.dim_t from config as segment_size, but the inference.num_overlap is not used by UVR, it's always using the value in the GUI
(while ZFTurbo original script is using audio.chunk_size and inference.num_overlap but not inference.dim_t . That's a mess) ” jarredou
Overlap comparisons for Roformers
4 is a balanced value in terms of speed/SDR according to measurements (since the beta patch #3 or later used above, overlap 16 and 32 are now the slowest in UVR (not overlap 2 is the slowest anymore when it was set the opposite) and overlap 4 has a bigger SDR than overlap 2 now.
Some people still prefer using overlap 8, while for others it’s already an overkill.
There’s very little SDR improvement for overlap 32, and for 50 there’s even a decrease to the level of overlap 4, and 999 was giving inferior results to overlap 16.
Compared to overlap 2, for 8 “I noticed a bit more consistency on 8 compared to 2 (less cut parts in the spectrogram).” Instrumentals with overlap higher than 2 can get gradually muddier.
So the highest SDR gives overlap 32, but for drastic separation time increase vs 16.
You might try to experiment with overlap 23 as a value in between rarely still used.
The SDR info is based on evaluations conducted on multisong dataset on MVSEP. Search for e.g. overlap 32 and overlap 16 below, and you will see the results to compare:
https://mvsep.com/quality_checker/multisong_leaderboard?algo_name_filter=kim
“overlap=1 means that the chunk will not overlap at all, so no crossfades are possible between them to alleviate the click at edges.” - it got fixed in newer versions of MSST.
The setting in GUI overrides the one in yaml’s setting.
Technical explanation on overlap by ZFTurbo - click.
chunk_size
“Most of the time using higher chunk_size than the one used during training gives a bit better SDR score, until a peak value, and then quality degrades.
For Roformers trained with 8 sec chunk_size [can be found in the yaml], 11 sec is giving best SDR (then it degrades with higher chunk size)” - jarredou
All the notable chunk_sizes are described later below.
Sometimes chunk_size can influence the ability of picking up e.g. some screams in the model (“kim's can do it if you mod the chunk size”).
Lower chunks might sound a bit less muddy.
Unless they’re set too high and VRAM is exceeded, while separation still doesn’t fail on AMD/Intel, various values should provide similar separation times, esp. for NVIDIA users.
batch_size (not used in UVR, but in MSST)
Leave it default, but for faster inference, e.g. Gabox used 6. Using above 2 might increase VRAM usage. 1 is forced in the inference Colab (it has the clicking issue with that setting fixed in newer MSST code).
“i.e. instead of running a single song at batch_size = 1 you can run 2 at the same time (batch_size = 2)” so “you can use more than one song to process”
Most common issues
- Some models might occasionally disappear from your list (most likely sideloaded outside Download Center) and change their name (probably once they got added to Download Center - e.g. Melband Roformer karaoke ckpt)
- Since patch #3, Roformers’ separation times with even smaller overlap are longer than before. Also, remember about inference mode added in one of the later patches. It’s disabled by default as it causes silent separation issues on GTX 10XX GPUs and probably older. Enabling it on newer GPUs might be beneficial for performance
- If UVR freezes your PC occasionally during separation, you can change priority of UVR to Idle in Task manager, and the issue is gone sooner or later.
You can force it to be remembered every time you run UVR with Process Lasso (it autostarts with the OS, so you don't have to see the splash screen for a free user every time you want to use it).
- Read for RTX 5000 series issue (if you refrain from using dedicated patch), or use OpenCL (DirectML) in options (slower)
Ensembling impossible - model not visible in vocal splitter in UVR
- If user-imported Roformers aren't recognized in "instrumental/vocals" in ensemble or in vocal splitter, but are in "multi-stem ensemble":
“The .yaml associated with the model usually needs to be updated to match UVR's stem naming conventions. For example, if your config shows the instruments as "other" and "vocals", it will need to be updated to "Instrumental" and "Vocals" (case-sensitive)” - Anjok
E.g. for Karaoke models, you need to change “Karaoke” to “Vocals”.
But you can also use Tools>Manual Ensemble instead for ready separations.
Ensemble vs multi-stem ensemble explained (by stephanie)
“The multi-stem ensemble mode is designed to ensemble every stem in all the models that were selected. I'll explain how it's related:
Let’s say you have two vocal/instrumental models selected, and two 4-stem models selected (vocals, drums, bass, other), then it will process the song through all of the models. As a result, the final ensembled output will be 5 stems: vocals, instrumental, drums, bass, and other.
The drums, bass, and other stems will just be the ensemble of the 2 models in that selection that shared those outputs. However, all models in that particular selection shared a vocal output so it will be the ensemble of each model's vocal output
I'm not sure if that's very clear, but that's how it works and why the way it handles each stem is different from the other stem pair modes”
- Roformers might be sometimes slow/stuck/give memory allocation error during separation on AMD/Intel GPUs with VRAM lower than 16GB, if you don’t lower default “chunk_size” during importing the model, or in the corresponding yaml in: models\MDX_Net_Models\model_data\mdx_c_configs
or in Choose Model>Edit Models Config>Change Parameters>Edit Model Param.
So even if separation works, it can be unnecessarily slow with too big chunks esp. on AMD/Intel GPUs on DirectML.
Start with min. chunk_size = 112455 (dim_t 256 equivalent) and increase it depending on GPU or model size till you start getting errors to get the best possible SDR.
For older beta 1-9 and 4GB VRAM GPUs, lower dim_t at the bottom of the yaml (not at the top) to e.g. 256 or 201, sometimes 301 - some Roformers will require lower dim_t/chunk_size - the higher, the better till 1101, or training chunks value in the config.
So since patch #10 “new beta version[s] doesn't rely on 'inference.dim_t' value anymore (if you were using edited "dim_t" value).
Now you have to edit audio.chunk_size now “In the model yaml config file, at top of it, chunk_size is the first parameter (...) you can edit model config files directly inside UVR now.
Memory issues - chunk_size table for combatting “RuntimeError”
dim_t (old patches) to chunk_size conversion (new patches) for Roformers with hop_length = 441 in the model’s yaml (patch #1-2 and Linux uses dim_t instead of chunk_size - more)
Useful if you see insufficient memory error, so you can decrease the default by the following value steps.
The formula is: chunk_size = (dim_t - 1) * hop_length
Go to Choose “MDX-Net model” and scroll down the the bottom to reach “Edit Model Config”, or edit the model yaml directly in Ultimate Vocal Remover\models\MDX_Net_Models\model_data\mdx_c_configs e.g. in Notepad++ and find “chunk_size” value and edit as in the following.
For this to work, you need to ensure that in:
Settings>Choose Advanced Menu >Advanced MDX Options>>Multi-Network Only Options you have the “Segment Default” option checked.
Otherwise the value from the main UVR window in “Segment Size” will be used instead (at least if it’s not set to “Default”).
dim_t = 1700 is chunk_size = 749259 (17s) - used by inst resurrection
dim_t = 1333 is chunk_size = 587412 (13,32s)
dim_t = 1201 is chunk_size = 529200 (12s) - used by some newer models
dim_t = 1101 is chunk_size = 485100 (11.00s) - that dim_t value was giving the highest SDR for models trained with 8s chunks, at least in times of models released in beta Roformer beta patch #2 period, it’s default for e.g. duality models
dim_t = 801 is chunk_size = 352800 (8.00s) - default for most models, max working on Intel/AMD 8GB GPUs on 900MB models, 6-stem SW model on 4GB NVIDIA GPUs works faster with this or lower than 485100 setting
dim_t = 556 is chunk_size = 244755 (5,5s)
dim_t = 501 is chunk_size = 220500 (5s) - also, as below, but separation time slightly increased with at least a few browser tabs opened, higher crashes no matter what
dim_t = 456 is chunk_size = 200655 (4,5s) - works with Resurrection inst on 4GB AMD GPU
dim_t = 401 is chunk_size = 176400 (4s)
dim_t = 356 is chunk_size = 156555 (3,55s) - max supported for becruily inst and AMD 4GB VRAM (when in previous beta, dim_t needed to correspond with overlap to avoid stem misalignment, so probably halves caused misalignment before),
it can be more muddy in certain parts of songs vs 256
dim_t = 301 is chunk_size = 132300 (3s) - max working on NVIDIA 2GB VRAM GPU with deux model (memory management in CUDA is much better than in DirectML)
dim_t = 256 is chunk_size = 112455 (2,55s) - max working with e.g. becruily and smaller unwa’s 400MB exp. models on 4GB AMD GPU
dim_t = 201 is chunk_size = 88200 (2s) - required value for some more resource-hungry/bigger Roformers on AMD 4GB VRAM (some models only worked with that dim_t with earlier beta patches and here’s it’s probably the same or with 256 equivalent), but is still not low enough for most Roformers when demudder is used, and give “Could not allocate tensor” while single model separation previously worked with even bigger chunk setting. You shouldn’t go lower with that parameter, as even a 2 seconds chunk might sometimes give audio skips every two seconds, at least on UVR Roformer patch #2. 2 or 2,5 seconds will be rather bare minimum.” - jarredou (DTN edit)
All inst/voc Roformers seem to use hop_length: 441 (but ensure in yaml), so you always multiply that hop value by the desired dim_t - 1 to get correct chunk_size (e.g. corresponding with old dim_t values you were using in older beta patches)
Setting your NVIDIA Control Panel UVR profile to “CUDA Sysmem Fallback Policy” to “Prefer Sysmem Fallback” makes it crash less on e.g. 2GB GPU - mohammedmehditber
Performance issues
If the chunk_size is too big for your VRAM, it can slow down separation (esp. on Intel/AMD using DirectML).
If it’s slow, also make sure your overlap is not higher than 2 for the fastest separation.
[MSST only] “If not set already, batch_size=1 can also reduce memory footprint
And if the model is not converted to FP16 already [like Deux], converting it to FP16 should also reduce memory usage.
It can be done with that script https://github.com/ZFTurbo/Music-Source-Separation-Training/blob/main/scripts/prepare_weights_for_inference.py” - jarredou
But, sometimes FP16 might cause performance drops on GPUs not having FP16 acceleration (it starts from RTX 2000/RX 6000/ARC [rather only discrete Alchemist]), plus DirectML might have its own peculiarities in that regard.
So ideally, it would be best to have both model variants for various GPUs.
Usage of the script (install Python, and on fresh installation execute “pip install torch” then execute):
python C:\path_of_the_script\prepare_weights_for_inference.py --checkpoint "path_to_the_input_model/model_to_clean.ckpt" --output_file "path_to_the_output_model/cleaned_model.ckpt” --float16
- UVR won’t run without min. 3GB of disk space on C:\, but sometimes it’s still not enough for e.g. GPU Conversion unchecked to not trigger memory issues (e.g. “cannot allocate”), esp. on 2-3h files with e.g. Wind model (issues occuring on even 32GB RAM).
- If you don’t have much space on C:\ you can set the pagefile to even min. 500MB on C: drive and use other partition for pagefile if you run out of space on C:\ during the process and error occured
https://mcci.com/support/guides/how-to-change-the-windows-pagefile-size
- We have some reports about user custom ensemble presets from older versions no longer working (since 11/17/24 patch).
Sadly, you need to get rid of them (don’t restore their files manually) or the ensemble will not work and model choice will be greyed out. You need to start from scratch
- If your computer shuts down during separation, it’s usually the PSU's fault. Consider replacement to a higher-grade/tier one, and/or with more power.
- Black screens during separation might be an indicator of outdated drivers and/or lack of KB5077181 update occuring on NVIDIA (more).
Errors troubleshooting
“got an unexpected keyword argument 'linear_transformer_depth’ ”
In case of the error with any external Roformer model:
- delete “linear_transformer_depth: 0” line from the YAML file
- To fix issues with BS variant of e.g. anvuew’s de-reverb model in UVR, additionally to the above, also change the following in the yaml file:
“stft_hop_length: 512 to stft_hop_length: 441 so it matches the hop_length above” (thx lew).
If that line is not present in your model’s yaml config file, go to the settings, then choose MDX In the Advanced menu, and click the "Clear auto-set cache" button.
Then go back to the main settings, click "Reset all settings to default" and restart the app (thx santilli_).
These issues don’t happen in the ZFTurbo’s CML inference code of:
https://github.com/ZFTurbo/Music-Source-Separation-Training/
'use_amp' “Key error”
using the GH repo above, and referencing (separating) some models:
- “add:
use_amp: true
in the training part of [models’ yaml] config file (it's missing)”
“”’norm’”” attributeError using e.g. unwa beta 5e model in UVR
- a) Ensure you installed UVR Roformer patch (5.6.1), and you're not using the old 5.6 version (but 5.6.1 is reported once you open the app)
b) You could pick wrong model architecture in Install model option (so not Roformer, and generally v1), or haven’t turned on “Roformer model” option during importing the model into UVR (option present in the older beta versions).
You can go to the bottom of the models list and pick “Edit Model” to change it.
c) Alternatively the following might help in some cases:
“Edit the yaml file [of the model] from this -
training:
instruments:
- vocals
- other
target_instrument: vocals
use_amp: True
to this -
training:
instruments:
- Vocals
- Instrumental
target_instrument: Vocals
use_amp: True”
- Anjok
More troubleshooting on the “norm” issue
E.g. BS-Roformer_LargeV1 is stuck on 5%
- Decrease chunk_size to 112455 (new patches)
(pre-10# patches) “Go to MDX settings, MX23C specific, turn off default segment size and use segment size 256, it's probably filling up your VRAM”
The setting resets itself. You should be able to set it permanently in the yaml configuration file of the model at the bottom (dim_t parameter).
It might be required for AMD/Intel 4GB VRAM GPUs (or potentially even 201, although it was using Rofo patch #2).
Q: Is there a way to reset which yaml file to use? I chose the incorrect yaml file for a particular ckpt file, and now I cannot change it
A: Go to models list>Edit Model Config>Change parameters and choose the yaml from the list.
Or go to Ultimate Vocal Remover\models\MDX_Net_Models\model_data and if it was done just now, the last modified yaml in this folder will be the one corresponding to your model. You can just open that json, and edit it to write the config name located in mdx_c_configs you want to use with that model. You should find proper hashed yaml by the saved model of your choice, but if you really feel lost, you can delete all hashed yamls, so you’ll need to go through the process of choosing configs for all custom MDX and Roformer models, but remember to not delete "model_data.json" and "model_name_mapper.json" as they cover models added to Download Center or written to be recognized automatically by UVR, so it’s rather not a place you look for. Also, you can decode hashed json names corresponding to specific models here.
Layers error - for issues with Unwa 400MB model
-“First make sure you're running the latest patch. If you're on the latest patch, It might be trying to associate with an incompatible YAML”, but resetting the parameters below might be enough.
Go into the mdx_c_configs folder
Find and delete BS_Inst_EXP_VRL.yaml
Go back into the "Download Center"
Select "MDX-Net" and give it a moment.
Close the Download Center and try again.
If that doesn't work, you might have a previous json model file that's interfering:
Select the model in the MDX-Net model menu
Then select "Edit Model Config"
From the popup, click "Reset Parameters"” Anjok
Layers errors - general
a) You didn’t install the newest patch and still use e.g. beta 2 with some newer model
b) You could check Roformer v2 instead of v1 during installing of the custom model.
c) Model trainer didn't clean the weight using this Python script (you can do it by yourself)
Usage of the script (install Python, and on fresh installation execute “pip install torch” then execute):
python C:\path_of_the_script\prepare_weights_for_inference.py --checkpoint "path_to_the_input_model/model_to_clean.ckpt" --output_file "path_to_the_output_model/cleaned_model.ckpt”
d) You can just use MSST instead
TypeError: (...) freqs_per_bands
You probably set Mel-Roformer model type instead of BS-Roformer when it was necessary.
Go to MDX-Net, pick the model>Edit model config>Choose parameters.
There you should change Model type.
mlp_expansion_facfor
You probably use outdated codebase for Linux or old UVR version incompatible with some newer Roformers like the SW.
We have some workarounds for it, but at least below not confirmed to work so far:
"The first issue is that the BS Roformer implementation in UVR5 does not accept the mlp_expansion_factor parameter.
>If mlp_expansion_factor is set to 4 in the [yaml] config file, simply remove that line.” - anvuew
And parallely the second issues might exist:
"The size of tensor a (484864) must match the size of tensor b (485100) at non-singleton dimension 2"
Regarding the second issue, chunk_size should be a multiple of stft_hop_length." - anvuew
Q: Hop is 512 and chunk is 588800.
That's a ratio of 1150.
A: Then you need to find out where this 485100 comes from.
[And they didn't, although in some other case, it stopped appearing also after fixing the error below and deleting
use_torch_checkpointing line too]
The issue also exists in the old MSST versions on which Colabs are based on, so from before refactoring of the MSST code. Fixed model_utils.py like below:
= model(arr)
# The fix:
if x.shape[-1] != arr.shape[-1]:
# If the dimensions do not match, trim or fill x to size arr
import torch.nn.functional as F
diff = arr.shape[-1] - x.shape[-1]
if diff > 0:
x = F.pad(x, (0, diff)) # Pad with zeros if too short
else:
x = x[..., :arr.shape[-1]] # Trim if too long
In mdx23c_tfc_tdf_v3.py (fixed) also padding was fixed before, but IDK if it was necessary:
res = encoder_outputs.pop()
if x.shape[2:] != res.shape[2:]:
import torch.nn.functional as F
x = F.interpolate(x, size=res.shape[2:], mode='bilinear', align_corners=True)
x = torch.cat([x, res], 1)
Or alternatively “I've added a length constraint to the inverse stft function as additional safe guard for small shape mismatch (which is not present in ZF's MSST repo version or original mdx23c code). It will not help with wrong chunk_size values but maybe help with these "randomly" happening issues.” - jarredou
skip_connection
Remove that entire line from the yaml too if you started seeing this error after the above.
But it might still not help.
torch._dynamo.polyfills.fx
Some of the newer Roformer models fail with the UVR RTX 5000 patch giving that error, while the non-RTX 5000 patch works as usual.
You need to go to model’s yaml and edit the line:
use_torch_checkpoint: True
to False, or delete it.
To find out which yaml is associated with your model, scroll down the model list, and go to Edit Models Config>Change Parameters>Edit Model Param option, and the corresponding config name will be shown at the top.
Then you should be able to locate this yaml in: models\MDX_Net_Models\model_data\mdx_c_configs
Alternatively use MSST instead
[WinError 2] The system cannot find the file specified
- Sometimes 32/64 bit float output set can trigger it
- Or you can also reencode your input file and name it intput.wav, and choose wav mode.
- Setting FLAC output might also work (it seems to happen with mp3 input and output).
(can’t remember if it still exists in the latest patches)
num_bands
“Most of these models are BS[band-split]/Mel-Band Roformer. [It] can be distinguished by their config, bs uses `freqs_per_bands`, Mel-band uses `num_bands`”
So you most likely installed this model with wrong model type chosen during installation of the model. Go to Edit model config and set it to Mel-Roformer instead of BS-Roformer (or the opposite) and the non-V2 one.
System error / file not found
> Use WAV output quality
- UVR will only process files with English characters - some complicated names/paths give “System error” during separation
(can’t remember if it still exists in the latest patches - check if your UVR is up to date as well)
E.g. for RuntimeError: "Error opening 'F:/Fl studio files/ACAPELLAS\1y2mate.com - 2 AM Full Video Karan Aujla Roach Killa Rupan Bal Latest Punjabi Songs 2019(Vocals).wav': System error." your file path/file name is too long or contains some unsupported charts.
>You need to shorten/simplify the path and/or copy the input file to a different location. E.g. D:\input.wav
Python39.dll
Probably it happened after installation of some patch on (too) dirty UVR installation (probably already patched before or some older than 22_30.
You must reinstall UVR using only the latest required patches.
More troubleshooting
RuntimeError: ""
Traceback Error: "
If you have these two lines without any text at the beginning of the error log as above using AMD GPU on every attempt of separation with GPU Conversion option turned on in UVR, for all archs and models
- you probably use outdated GPU drivers and/or Windows not compatible with newer drivers.
>Go to Ultimate Vocal Remover\torch_directml and replace DirectML.dll from C:\Windows\System32\AMD\ANR (make a backup before).
>Experimentally, you can use this older 1.9.1.0 version of the library. Restart UVR after replacing the file!
If you use an incompatible library version, you’ll encounter the “Unhandled exception” startup issue. The same fix might work on Intel GPUs.
Be aware that the linked older version of the library might cause additional noise for MDX-Net v2 models like HQ_X (the issue is gone when you turn off GPU Conversion).
The same Runtime Error: “” also happens on Mac Pro 2009, at least on some older UVR versions. If updating UVR won't help, turn off GPU Conversion. It might be too old.
Theoretically all DX12 GPUs should be compatible with DirectML, but min. VRAM working with Roformers with low chunk_sizes is rather 4GB, and older NVIDIA GPUs than Maxwell might fail to switch to DirectML at least on 5.6.1 (in 5.6.0 and older, the option might be called OpenCL, and was probably deleted in Roformer patches).
- All MDX-Net v2 models (maybe beside 4 stem variants), have so called MDX noise, which can be cancelled by using Options>Advanced MDX-Net Settings>Denoise Output>Standard (or Model). Just with older directml.dll it's more noisy, and dedicated models don't work so efficiently with this more elevated noise.
- At least beta #2 Roformer update caused some stability and performance issues with other archs than Roformers for some people when specific parameters started to take more time than before.
Roll back to stable 5.6 (non-5.6.1) in these cases if necessary (but you won't have Roformers support). Possibly make a copy of the old installation. Your configuration files might be lost. You can use two installations at the same time (or at least when one, e.g. Roformer patch is installed or symlinked in the default location).
- (I think I covered that issue above more thoroughly)
Roformer models in at least patch #2 work only in “Multi-Stem” mode in UVR. Using them in Ensemble causes layers errors (you can use manual Ensemble instead).
Iirc, it’s caused by yaml config where instead of Instrumental + Vocals (with V as capital letter) there’s written other + vocals, and you need to change it. Iirc it doesn’t happen on models downloaded from Download Center as Anjok was fixing the issue, but the problem might still exist in yamls of some custom models outside the center
- If you have sudden issues with not being able to separate, try to reinstall the app, and/or possibly make sure you didn’t turn on some power saving option in your laptop. Plus, you can simply try to reopen UVR (few fail tries on incompatible DirectML.dll with your GPU driver/OS will hang UVR on “Loading Model” till you close UVR manually from Task Manager).
- When using e.g. BS-Roformer SW: “RuntimeError: "The size of tensor a (2) must match the size of tensor b (6) at non-singleton dimension 0"
Traceback Error: "”
> Replace the yaml config manually in:
C:\Users[User]\AppData\Local\Programs\Ultimate Vocal Remover\models\MDX_Net_Models\model_data\mdx_c_configs
Restart UVR, start over. Make sure you were asked to replace it. If not, the yaml was maybe wrongly picked anyway. Then edit in the Edit config menu.
_______
You’ll find more UVR troubleshooting in this section
_____
Problems fixed in newer patches
- (deprecated since patch #10 - now convert dim_t it to chunk_size) dim_t = 1101 seems to be a sweet spot in terms of speed/SDR according to measurements (although on 1 minute files); use 1120 if UVR refuses to accept 1101 in GUI (or edit yaml file)
- (deprecated since patch #10) Some Roformer configs have wrong dim_t at the bottom of the yaml by default (e.g. 256), change it at the bottom of the yaml config for better SDR (not the one at the top), e.g. to 1101 (more explanations on it later).
- (fixed in patch #10) VIP code in Roformer beta patch #2-9 (and probably #1) doesn’t work -
Download all the VIP models you need before patching older 5.6 to beta Roformer or use two installations of the UVR if you can’t use patch #10 with the fix.
- (fixed in patch #6) People experience All stems error with viperx’ 12xx models in newer versions of UVR Beta Roformer patch (patch #2 was the last confirmed to work with these older models)
- (fixed in patch #10) mlp_expansion_factor: 1 or (when mlp line is deleted from yaml) mismatch for MelBand Roformer error
You probably use older Roformer patch incompatible with newer models (e.g. #2)
It also appears when you wrongly set v2 model type.
Fixed in the beta patch #3 and #4 for all platforms
- Don't set overlap higher than 11 for 1101 dim_t (at the bottom of yaml file in the “inference” section, not above) and overlap 8 for 801 - these two are the fastest settings before stem misalignment issues occur. Otherwise, it can lead occasionally to some effects or synths missing from the instrumental stem (although some rules can be broken here with various settings). Also, the problems with clicks are alleviated with these good settings.
- In beta #2 patch, best measured SDR for both Mel and BS-Roformers is when dim_t = 1101 in the inference section of yaml config and when overlap is set to 2 in GUI (although 1 wasn’t tested, and is actually lower). But the last beta patches, all bigger overlap values are slower, so SDR might be higher with higher values.
Be aware that it will increase separation time. Maximum allowed value before error is 1801, but 1501 or 1601 depending on a model will be the max reasonable for experiments before some unwanted downsides of too high or too low dim_t appear (disappearing of some stem elements). In some specific cases, 1333 (or potentially 1301) was giving better results than 1101 or 1501, but it depended on song length - usually it happened on short fragments.
- Instruction for overlap and dim_t above applies to other Roformer models as well, and not only those in Download Center. With the instructions, you can achieve faster separation times, as you’re not forced to use the most time-consuming overlap 2 in older patches to avoid stem misalignment issues
_______________
Infos and fixes for older patch #1/2
(with matching overlap (reversed) and dim_t necessity)
- To avoid separation errors for 4GB VRAM and AMD/Intel GPUs using Roformers, set segments 32, overlap 2 and dim_t 201 with num_overlap 2 both at the bottom of yaml config in \models\MDX_Net_Models\model_data\mdx_c_configs
(dim_t 301 and overlap 3 also works, although not on all models [e.g. not for beta 3, but inst v1] and seems to be less muddy and fewer clicks appear).
dim_t 201 is not optimal setting and might lead to more occasional quiet residues, clicks or sudden volume changes (like chunk was changing every 2 seconds), although there’s no stem misalignment issue with these settings (they work both for Mel and BS Roformers). dim_t 301 with lighter models seems to be a bare minimum to avoid the majority of audible artefacts (after patch #3 dim_t 256 is allowed - “make sure you check the "Segment Default" in MDXNET23 Only Options for it to take effect”).
Using the settings above on patch #2, with GPU acceleration it will take 39m 28s for 3:28 song using 1296 model on RX 470 4GB and 18 minutes for Kim Mel-Roformer and 3:01 song.
Using HQ_4 is much faster than realtime using default settings, but even longer than accelerated Roformer, when on CPU only using old Core 2 Quad @3.6 DDR2 800MHz.
On Mac M1 using the patch above, it takes 9 minutes to process a 3-minute song using BS-Roformer (dim_t 1101, batch size 2, overlap 8) with “constant throttling”. Click
And below 4 minutes for Kim Mel-Roformer (overlap 1, dim 801). Click
- Settings working for 6GB AMD GPUs: dim_t 601 or 701 at the bottom of the yaml file and overlap 6 or 7 in GUI.
- Like I mentioned, overlap 8 can be good enough too when dim_t=801 is set (the fastest setting before SDR getting drastically reduced), at least in other cases you shouldn’t exceed 6, while 2 should provide the best quality in most cases.
- 1602 (or rather 1601) dim_t might lead to less wateriness, but turns out in cost of a bit more of vocal residues.
- “In theory, max overlap value [for Roformer separations without mentioned issues in UVR] can be known with formula:
(dim_t - 1) / 100 = Max_overlap_value
if dim_t = 801:
(801 - 1) / 100 = 8
if_dim_t = 1101:
(1101 -1) / 100 = 10 [jarredou wrote 10 here, but it’s actually 11]
Above that max value, some parts of the input will not be processed.
The lower the overlap value is, the more overlap is used, so better SDR.
- Some Rofos models still have wrong config by default, with dim_t=256, so max overlap value for that is 2. That's why I've advised to stick to overlap=2” - jarredou
So in times before dim_t was known how to be correctly set, so now overlaps can be even set to 8 now when dim_t=801 is set]).”
The same thing applies for both BS and Mel Roformers in UVR.
“audio.dim_t value is not used with roformers in ZFTurbo script, it uses audio.chunk_size and then it's parameters in the model part of config.”
- Using older Roformer beta patches for Mac M1 doesn’t allow you to choose the Roformer parameter to check for custom Roformer models and only config name can be chosen, but no confirm button is available. So the error “File "libv5/tfctdfv3.py", line 152, in __init” appears.
> Place the corresponding json file with your model from this repo into: models\MDX_Net_Models\model_data beforehand, to fix the issue.
In some cases, you may still get the same error anyway and to get rid of it, you need to edit manually model_data.json adding desired model line at the end like your custom model was downloaded from download center. On example of unwa’s beta 3:
},
"d43f93520976f1dab1e7e20f3c540825":{
"config_yaml": "config_melbandroformer_big.yaml",
"is_roformer": true
}
Additionally, you need the model at the end of model_data_mapper.json:
"model melband_roformer_big_beta3.ckpt": "config_melbandroformer_big"
}
Now copy the hash-named json file (d43f93520976f1dab1e7e20f3c540825.json for beta 3) to model_data folder.
All the three modified files for beta 3 and other models here.
If you have problems generating hash on first launch of the model and your model is not uploaded in the repo above or json is not generated then use Windows installation in VM, or ask some PC user for the config. Potentially reading Hash decoding can be helpful.
But maybe your hashed config name will be generated correctly already after you imported the model into UVR (although no confirmation button might prevent it), and now it will be enough to just place the following line like in the jsons presented above: "is_roformer": true” (so after “,” in the yaml line above).
- More in-depth - Settings per model SDR vs Time elapsed -||- (incl. dim_t and overlap evaluation for Roformers) - click or here | conclusion - made before patch #3
Model characteristics
(the list might be getting outdated, read models list at the top)
Note: E.g. unwa’s duality models v1/2 and inst v1/2/v1e are now added to UVR Beta Roformer Download Center (so you don’t have to mess with models and configs manually)
- viperx 1053 model separates drums and bass in one stem, and it's very good at it
(although now it might be better to use Mel-Roformer drums on x-minus.pro/uvronline)
“Target is drums and bass, and "other" is the rest. Despite that, it says vocals”
- Unwa released a new Inst v1e model | Colab | MSST-GUI (“The model [yaml] configuration is the same as v1”)
“The "e" stands for emphasis, indicating that this is a model that emphasizes fullness.”
- unwa inst v2 - it gets muddier than v1 at times, but it has less of noise
- unwa inst v1 - focused on instrumental stem:
model | Colab | MSST-GUI | phase fixer
"much less muddy (..) but carries the exact same UVR noise from the [MDX-Net v2] models"
But it's a different type of noise, so aufr33 denoiser won't work on it.
“you can "remove" [the] noise with uvr denoise aggr -10 or 0” although with -10 it will make it sound more muddy like Kim model and synths and bass are sometimes removed with the denoiser (~becruily). Mel-Roformer denoise might be better for it.
becruily released a Python script fixing the noise issue (execute “pip install librosa” in case of module not found error) - it sound similar to the method used for premium user on x-minus.
- unwa beta 4 Mel-Roformer (fine tune of Kim’s voc/inst model ):
https://huggingface.co/pcunwa/Mel-Band-Roformer-big/tree/main | Colab
Be aware that the yaml config has changed, and you need to download the new beta4 yaml.
“Metrics on my test dataset have improved over beta3, but are probably not accurate due to the small test dataset. (...) The high frequencies of vocals are now extracted more aggressively. However, leakage may have increased.” - unwa
“one of the best at isolating most vocals with very little vocal bleed and still doesn't sound muddy” “gives fuller vocals”. Can be a better choice on its own than some ensembles.
- unwa duality model - focused on both stems, and instrumental is similarly muddy like in beta 4
- Kim Mel-Band Roformer vocal model
It’s less muddy than 1296/1297.
(original repo - CML faster on CUDA than in UVR | model | config - place the model file to models\MDX_Net_Models and .yaml config to model_data\mdx_c_configs subfolder and “when it will ask you for the unrecognised model when you run it for the first time, you'll get some box that you'll need to tick "roformer model" and choose it's yaml” (Mac issue explained in the section above).
(simple Colab/CML inference/x-minus/MVSEP/jarredou Colab too now)
- unwa BS-Roformer finetuned a.k.a. large (further trained viperx 1297 model) download
More muddy than Kim above, a bit less of vocal residues, a bit more artificial sound.
- Mel-RoFormer Karaoke / Lead vocal isolation model files released by Aufr33 and viperx (download)
Older models in Download Center
- older viperx’ 1297 model tend to be a bit better for instrumentals, and 1296 for vocals (both more muddy than Kim and Unwa models, but “still pretty good for voice cleaning” and dealing with noise) - BS-Large model by Unwa is a fine-tune of that model.
- 1143 model is the first Mel-Roformer trained by viperx before Kim introduced changes to the config, which fixed the problem of lower SDR vs models trained on BS-Roformer. Use Kim Mel-Roformer instead
Both models struggle with saxophone and e.g. some Arabic guitars. It can still depend on a song whether these are better than even the second oldest Roformer than on MVSEP (from before viperx model got fine-tuned version). They tend to have more problems with recognizing instruments. Other than that, they're very good for vocals (although Mel-Roformer by Kim on x-minus tends to be better).
Muddy instrumentals when not ensembled with other archs.
Be aware that names of these models on UVR refer to SDR measurements of vocals conducted on private viperx dataset, not even older Synthetic dataset, instead of on multisong dataset on MVSEP, hence the numbers are higher than in the multisong chart on MVSEP.
___
Older news follow
___
- The viperx model was also added on MVSEP
- New ensembles with higher SDR were added on MVSEP
- BS-Roformer model trained by viperx was added on x-minus (it's different from the v2 model on MVSEP, and has higher SDR, it's the “1.0” one). If it's better vs V2 might depend on a song.
It struggles with saxophone and e.g. some Arabic guitars.
- (x-minus - aufr33) “I have just completed training a new UVR De-noise model. Unlike the previous version, it is less aggressive and does not remove SFX.
It was trained on a modified dataset. I reduced the noise level and made it more uniform, removed footsteps, crowd, cars and so on from the noise stems. On the contrary, the crowd is now a useful / dry signal. (...) The new model is designed mainly to remove hiss, such as preamp noise.”
For vocals that have pops or clipping crackles or other audio irregularities, use the old denoise model.
- Dango.ai updated their model, also giving some kind of demudder to the instrumentals, enhancing their results. Results might be better than MDX23C and BS-Roformer v2. Still, it’s pretty pricey (8$ for 10 separations). 5x 30 seconds fragments per IP can be obtained for free, and usually it doesn’t reset. “It’s $8 for 10 tracks x 6 minutes, all aggressiveness modes included (but vocal and inst models are separate). The entire multisong dataset for proper SDR check would cost around $133.” becruily
- Be aware that queues on https://doubledouble.top/ are much shorter for Deezer than Qobuz links. If there’s no 24 bit versions for your music, use Deezer instead.
[outdated; currently there’s no longer any MQA files on Tidal] Also, avoid Tidal and 16 bit FLACs from “Max” quality, which is slightly lossy MQA. Use 24 bit MQA from Tidal only when there’s no 24 bit on Qobuz. Most older albums under 2020 are 16 bit MQA instead of 24 bit MQA on Tidal, and are lossy compared to Deezer and Qobuz which doesn’t use MQA (so doubledouble doesn’t convert MQA to FLAC like on Tidal). MQA is only “slightly” lossy, because it affects frequencies mainly from 18kHz and up, and not greatly.
- Members of neighboring AI Hub server made a fork of KaraFan Colab updated with the new HQ_4 and InstVoc HQ2 models. It has slow separation fix applied. Click
- HQ_4 and Crowd models added to HV Colab temp fork before merge with main GH repo
- (MVSEP) “We have added longer filenames disabling option to mvsep, you can access it from Profile page
20240312034817-b3f2ef51cb-ballin_bs_roformer_v2_vocals_[mvsep.com].wav -> ballin_bs_roformer_v2_vocals.wav
Due to browser caching, you might want to hard refresh the page if you have downloaded onc”
- The ensembles for 2 and 5 stems on MVSEP have been updated with bigger SDR bag of models containing now new BS-Roformer v2 (with MDX23C, VitLarge23, and for multistem, the old demucsht_ft, deumcs_ht, demucs_6s and demucs_mmi models)
- All the Discord direct links leading to images in this document have expired. I already reuploaded some more important stuff. Please ping me on Discord if you need access to some specific image. Provide page and expired link.
- https://free-mp3-download.net has been shut down. Check out alternatives here.
New Apple Music ALAC/Atmos downloader added, but its installation is a bit twisted and subscription is required. Murglar added.
- MDX-Net HQ_4 model (SDR 15.86) released for UVR 5 GUI! Go to Models list>Download center>MDX-Net and pick HQ_4 for download. It is an improved and faster than HQ_3, trained for epoch 1149 (only in rare cases there’s more vocal bleeding, more often instrumental bleeding in vocals, but the model is made with instrumentals in mind.
Along with it, also UVR-MDX-NET Crowd HQ 1 has been added in download center.
- HQ_4 model added to the Colab:
https://colab.research.google.com/github/kae0-0/Colab-for-MDX_B/blob/main/MDX_Colab.ipynb
- New BS-Roformer v2 model released on MVSEP. It’s more aggressive model than above.
- Fixed KaraFan Colab with the fix for slow non-MDX23 models. You'll no longer stack on voc_ft using any other preset than 1, but be aware that it will take 8 minutes more to initialize. (same fix as suggested before, but w/o console, as it wasn't defined, and faster ort nightly fix doesn't work here).
Turns out, there has been an official non-nightly package released, and it works with KaraFan correctly (no need to wait 8 minutes any longer):
!python -m pip -q install onnxruntime-gpu --extra-index-url https://aiinfra.pkgs.visualstudio.com/PublicPackages/_packaging/onnxruntime-cuda-12/pypi/simple/
- (x-minus.pro) “Since Boosty is temporarily not accepting PayPal and generally working sucks, I made the decision to go back to Patreon. Please be aware that automatic charges will resume on March 22, 2024. If you have Boosty working correctly and do not intend to use Patreon, please cancel your Patreon subscription to avoid being charged.
If you wish to switch from Boosty to Patreon, please wait for further instructions in March.” Aufr33
- If you suffer from bleeding in other stem of 4 stems Ripple, beside decreasing volume by e.g. 3/4dB also “when u throw the 'other stem' back into ripple 4 track split a second time, it works pretty well [to cancel the bleeding]” if it's still not enough, put other stem through Bandlab Splitter.
- If you suffer from vocal residues using Ensemble 4 models on MVSEP.com, decrease volume of input file by -8dB “now it's silent. No more residue” usually 3 or 4dB was doing the trick for Ripple, but here it’s different. Might depend on a song too.
- Image Line “released an update for FL Studio, and they improved the stem separation and it's better, but it has quite a bit of bleeding still, but it also seems they may have improved the vocal clarity”
- (probably fixed in new HV MDX) Our newly fixed VR and newer HV MDX Colabs started to have issues with very slow initialization for some people (even 18 minutes/+ instead of normally 3). It’s probably due to very slow download of some dependencies. Possible solutions: use other Google account, use VPN, make another Google account (maybe using Polish VPN). Let us know if it happens only for some specific dependency or all of them. You can try to uncomment the ORT nightly line in mounting cell (add # before), as it triggers more dependencies to be installed, which can be slow in that case. The downside is - there won't be GPU acceleration, and one song will be processed in 6-8 minutes instead of ~20 seconds.
- New paid drum separation service:
https://remuse.online/download (might not work anymore)
It uses free drumsep model (same model hash: 9C18131DA7368E3A76EF4A632CD11551)
- MDX Colab seem to not work due to Numpy issues. I already fixed them in Similarity Colab, and hopefully reimplement the fixes elsewhere soon. VR Colab fixed too.
Tech details about introduced changes described below Similary Extractor section.
- Music AI surfaced. Paid - $25 per month or pay as you go (pricing chart). No free trial. Good selection of models and interesting module stacking feature. To upload files instead of using URLs “you make the workflow, and you start a job from the main page using that custom workflow” [~ D I O ~].
Allegedly it’s made by Moises team, but the results seem to be better than those on Moises.
“Bass was a fair bit better than Demucs HT, Drums about the same. Guitars were very good though. Vocal was almost the same as my cleaned up work. (...) I'd say a little clearer than mvsep 4 ensemble. It seems to get the instrument bleed out quite well, (...) An engineer I've worked with demixed to almost the same results, it took me a few hours and achieve it [in] 39 seconds” Sam Hocking
- “I just got an email from Myxt saying they're going to limit stem creation to 1 track per month. For creator plan users (the $8 a month one) and 2 per month for the highest plan.
So I may assume with that logic, they're gonna take it away for free users?”
- (probably fixed) For all jarredou's MDX23 v. 2.3 Colab fork users:
“Components of VitLarge arch are hosted on Huggingface... when their maintenance will be finished it will work again. I can't do anything about it in the meantime.”
2.2 and 2.1 and MVSEP.com 4-8 models ensemble (premium users) should work fine.
- Ripple now has fade in and clicking issues fixed. Also, there's less bleeding in the other stem (but Bas Curtiz’ trick for -3dB/-4dB input volume decreasing can be still necessary).
“Ripple’s lossless outputs are weird, some stems like the drums are semi full band (kicks go full band, snares not etc) and the “other” stem looks like fake full band”. These fixes are applied also for old versions of the app.
Also, the lossless option fixes to some extend the offset issue so it's more similar to input now, but not identical (lossless option might require updating). Also no more abrupt endings
Ripple = better than CapCut as of now (and fullband).
plus Ripple fixed the click/artifacts using cross-fade technique between the chunks.
- ViperX currently doesn't plan to release his BS-Roformer model
- New “uvr de-crowd (beta)” model added on x-minus. Seems to provide better results than the MVSEP model. Also, an MDX arch model version is planned for training.
“At minimum aggressiveness value, a second model is now used, which removes less crowd but preserves other sounds/instruments better.”
- Ripple seems to have a lossless export option now. “First make sure the app is updated then click the folder then click the magnet icon then export and change it to lossless”
- Seems like CapCut now has added separation inside Android Capcut app in unlocked Pro version
https://play.google.com/store/apps/details?id=com.lemon.lvoverseas (made by ByteDance)
Seems like there is no other Pro variant for this app.
At least unlocked version on apklite.me have a link to regular version, so it doesn't seem to be Pro app behind any regional block. But -
"Indian users - Use VPN for Pro" as they say, so similar situation like we had on PC Capcut before. Can't guarantee that unlocked version on apklite.me is clean. I've never downloaded anything from there.
- Mega, GDrive and direct link support for input files added on MVSep. If you want to apply MVSep algorithm to result of other algorithm, you can use "Direct link" upload and point https link on separated audio-file on MVSep.
- If you have an issue with Demucs module not found in e.g. MDX23 v.2.3 Colab (now fixed there and also in VR Colab), here's a solution:
“In the installation code, I added `!pip install samplerate==0.1.0` right before the `!pip install -r requirements.txt &> /dev/null` and I managed to get all the dependencies from the requirements.txt installed properly.” (derichtech15)
- If you repost your images or files from Discord elsewhere while cutting link after "ex=" for all new posted files, it will make your files expire pretty soon (17.02.24). If you leave the full link with "ex=" and so on, it won't expire so fast, but who knows if not later.
So far, all the old Discord images shared elsewhere with "ex=" cut, work (also in incognito without Discord logged in), but it's not certain that it will be that way forever.
Discord announced in the end of 2023, that they'll update their mechanisms of sharing links, so they'll expire after some time when they're shared, to avoid some security vulnerabilities allowing scams. Or they just want to offload the servers.
- OpenVINO™ AI Plugins for Audacity 3.4.2 64-bit introduced.
4 stems separation, noise suppression, Music Style Remix - uses Stable Diffusion to alter a mono or stereo track using a text prompt, Music Generation - uses Stable Diffusion to generate snippets of music from a text prompt, Whisper Transcription - uses whisper.cpp to generate a label track containing the transcription or translation for a given selection of spoken audio or vocals.
Not bad results. They use Demucs.
- For people with low VRAM GPUs (e.g. 4GB or less), you can test out Replay app, which provides voc_ft model and tends to crash less than UVR. Sadly, the choice of models is much smaller, but it has some de-reverb solution. Screenshot
- Latest MVSep changes:
1) All ensembles now have option to output intermediate waveforms from independent algorithms + additional max_mag, min_mag.
2) Ensemble All-In now includes DrumSep results extracted from Drum stem.
- resemble-enhance (GH) model added on x-minus in denoise mode. It can work better than the latest denoise model on x-minus. It is intended only for vocals. For music use UVR De-noise model on x-minus.
- (fixed in kae, 2.1, 2.2 [and KaraFan irc] Colabs) All Colabs using MDX-Net models are currently very slow. GPU acceleration is broken and separations now only work on CPU with onnxruntime warnings.
To work around the issue, go to Tools>Command palette>Use fallback runtime version (while it's still available).
Downgrading CUDA to 11.8 version fixes the issue too, but it takes 9 minutes in order to install that dependency, so it’s faster to use fallback runtime till it’s still available. After that period, just execute this line after initialisation cell:
console('apt-get install cuda-11-8') and GPU acceleration will start to work as usual.
>“Better fix [than CUDA 11.8] until final version is released, using that onnxruntime-gpu nightly build for cuda12:
!python -m pip install ort-nightly-gpu --index-url=https://aiinfra.pkgs.visualstudio.com/PublicPackages/_packaging/ort-cuda-12-nig
htly/pypi/simple/
(no need to install cuda 11.8)” jarredou
In case of credential issues you can try out this package instead:
!python -m pip -q install onnxruntime-gpu --extra-index-url https://aiinfra.pkgs.visualstudio.com/PublicPackages/_packaging/onnxruntime-cuda-12/pypi/simple/
- LarsNet model was added on MVSep. It's used to separate drums tracks into 5 stems: kick, snare, cymbals, toms, hihat. Source: https://github.com/polimi-ispl/larsnet
It’s worse than Drumsep as it uses Spleeter-like architecture, but “at least they have an extra output, so they separate hihats and cymbals.”. Colab
“Baseline models don't seem better quality than drumsep, but the provided checkpoints are trained with oly 22 epochs, it doesn't seem much. (and STEMGMD dataset was limited by the only 10 drumkits), so it could probably be better with better dataset & training”
“ it separates the toms so much better [than Drumsep]”
Similar situation as with Drumsep - you should provide drums separated from e.g. Demucs model.
- Captain FLAM from KaraFan asks for some help due to some recent repercussions.
You can support him on https://ko-fi.com/captain_flam
- To preserve instruments which are counted as vocals by other MDXv2 models in KaraFan, use these preset 5 modified settings (dca100fb8).
- Added more remarks from testing these settings against sax preset and others.
- drumsep added on MVSEP!
(separation of drums from e.g. Demucs 4 stem or “Ensemble 8 models”/+)
- New Bandid Plus model added on MVSEP
“I trained BandIt for vocals. But it's too far away from MDX23C” -ZFTurbo
“I loved this bandit plus model!! It has great potential.”
- UVR De-noise model by FoxJoy added on x-minus. It’s helpful for light noise, e.g. vinyl. (de-reverb and de-echo are up already)
New MDX de-noise model is in the works and beta model was also added!
“the instruments in the background are preserved much better than the FoxJoy model”
It works for hiss, interference, crackle, rustles and soft footsteps, technical noise.
- New hifi-gan-bwe Colab fork made by jarredou:
https://colab.research.google.com/github/jarredou/hifi-gan-bwe/blob/main/HIFIGAN_BWE.ipynb
- New AI speech enhancer - https://www.resemble.ai/introducing-resemble-enhance
- Reason 12.5 (a DAW) was released with VST3 plugin support
- jazzpear94 “I made a new model with a modified version of my SFX and Music dataset with the addition of other/ambient sound and speech. It's a multistem model and should even work in UVR GUI as it is MDX23C.
Note: You may want to rename the config to .yaml as UVR doesn't read .yml and I didn't notice till after sending. Renaming it fixes that, however”
“You put config in models\mdx_net_models\model_data\mdx_c_configs. Then when you use it in UVR it'll ask you for parameters, so you locate the newly placed config file.”
“Keep in mind that the cinematic model focus is mainly on sfx vs instruments
voice stems are supplemental. Usually I remove voices first”
- https://github.com/karnwatcharasupat/bandit
Better SDR for Cinematic Audio Source Separation (dialogue, effect, music) than Demucs 4 DNR model on MVSEP (mean SDR 10.16>11.47)
- "Demucs+CC_Stereo_to_5.1" - it's a script where you can convert Stereo 2.0 to 5.1 surround sound. Full discussion about script. They use MVSep to get steams and after use script on them.
- Colab by jazzpear96 for using ZFTurbo's MSS training script. “I will add inference later on, but for now you can only do the training process with this!”
- New djay Pro 5.0 has “very good realtime stems with low CPU” Allegedly “faster and better than Demucs, similar” although “They are not realtime, they are buffered and cached.” it uses AudioShake. It can be better for instrumentals than UVR at times.
- AudiosourceRE Demix Pro new version has lead/backing vocals separation
- New crowd MSX23C model added on MVSEP (applause, clapping, whistling, noise) (and got updated by the time 5.57 -> 6.06; added hollywood laughts, old models also available)
- VitLarge23 model on MVSEP got updated (9.78>9.90 for instrumentals)
- MelBand RoFormer (9.07 for vocals) model added on MVSEP for testing purposes
“The model is really good at removing the hi-hat leftovers. These e.g. in the Jarredou colab sometimes when you can hear the hi-hats from the acapella. And Melband roformer can almost remove all the hi-hat leftovers from the acapella.”
“are the stems not inverted result? for me it sounds like there is insane instrument loss in the instrumental stem and vocals loss in the vocal stem, yet there is no vocal bleed in instrumental stem and vice versa” “I also think that the vocals are surprisingly clean considering the instrumentals sound quite suppressed but also clean”
- Goyo Beta plugin for dereverb stopped working on December 2nd (as it required internet connection and silent authorization on every initialization). They transitioned to paid Supertone Clear. They send BETA29 coupon over emails (with it, it’s $29).
- New MVSep-MDX23 Colab Fork v2.3 by jarredou published under new Colab link here
Now it has Vitlarge23 model (previously used exclusively on MVSEP) instead of HQ3-Instr, also improved BigShifts and MDXv2 processing.
Doesn't seem to be better than RipX which is better in preserving some instruments, and also removes vocals completely
- Check out new Karaoke recommendations (dca100fb8)
- Dango.ai finally received English web interface translation
- New SFX model based on Mel roformer was released by jazzpear94. More info
- User friendly Colab made by jarredou and forked by jazzpear94 with new feature. In case of some problems, use WAV file.
- Seems like Ripple got updated, "it sounds a lot better and less muddied" doesn’t seem to give better results for all songs, though. Might be similar case with Capcut too.
- Hit 'n' Mix RipX DAW Pro 7 released. For GPU acceleration, min. requirement is 8GB VRAM and NVIDIA 10XX card or newer (mentioned by the official document are: 1070, 1080, 2070, 2080, 3070, 3080, 3090, 40XX, so with min. 8GB VRAM). Additionally, for GPU acceleration to work, exactly “Nvidia CUDA Toolkit v.11.0” is necessary. Occasionally, during transition from some older versions, separation quality of harmonies can increase. Separation time with GPU acceleration can decrease from even 40 minutes on CPU to 2 minutes on decent GPU.
- UVR BVE v2 beta has been updated on x-minus
“It now performs better on songs with 2 people singing the lead
No longer separates the second lead along with it”
-dca100fb8 found out new settings for KaraFan which give good results for some difficult songs (e.g. Juice WRLD) for both instrumental and acapella. It’s now added as preset 5.
Debug mode and God mode can be disabled, as it's like that by default.
"It's like an improved version of Max Spec ensemble algorithm [from UVR]"
Processing time for 6:16 track on medium setting is 22 minutes.
- New MDX23C model added exclusively on MVSEP:
vocals SDR 10.17 -> 10.36
instrum SDR 16.48 -> 16.66
Also ensemble 4 got updated by new model (10.32>10.44 for vocals)
- For some people using mitmproxy scripts for Capcut (but not everyone), they “changed their security to reject all incoming packet which was run through mitmproxy. I saw the mitmproxy log said the certificate for TLS not allowed to connect to their site to get their API. And there are some errors on mitmproxy such as events.py or bla bla bla... and capcut always warning unstable network, then processing stop to 60% without finish.” ~hendry.setiadi
“At 60% it looks like the progress isn't going up, but give it idk, 1 min tops, and it splits fine.” - Bas
-ZFTurbo published his training code:
https://github.com/ZFTurbo/Music-Source-Separation-Training
"It gives the ability to train 5 types of models: mdx23c, htdemucs, vitlarge23, bs_roformer and mel_band_roformer.
I also put some weights there to not start training from the beginning."
It contains checkpoint of e.g. 1648 (1017 for vocals) MDX23C model to train it further.
Be aware that the older bs_roformer implementation is very slow to train IRC.
Vitlarge23 “is running 2 times faster than MDX models, it's not the best quality available, but it's the fastest inference”
“change the batch size in config tho
I think zfturbo sets the default config suited for a single a6000 (48gb)
and chunksize”
-"A small update to the backing vocals extractor [on X-Minus]
Now you can more accurately specify the panning of the lead vocal." ~Aufr33 Screen
- IntroC created a script for mitmproxy for Capcut allowing fullband output, by slowing down the track. Video
- Jazzpear created new VR SFX model. Sometimes it’s better, sometimes it’s worse than Forte’s model. Download
For UVR 5.x GUI, use these parameters (irc same as Forte):
User input stem name: SFX
Do NOT check inverse stem!
1band sr44100 hl 1024
- Now KaraFan should work locally on 4GB GTX GPUs (e.g. laptop 1060), on presets 2 or 3, and with chunk 500K, speed can be slowest. Download on GitHub the Code > ZIP
-Bas Curtiz' new video on how to install and use Capcut for separation incl. exporting:
https://www.youtube.com/watch?v=ppfyl91bJIw
and saving directly as FLAC, although the core source of FLAC is still AAC in this case:
https://www.youtube.com/watch?v=gEQFzj6-5pk
"It's a bit of a hassle to set it up, but do realize:
- This is the only way (besides Ripple on iOS) to run ByteDance's model (best based on SDR).
- Only the Chinese version has these VIP features; now u will have it in English
- Exporting is a paid feature (normally); now u get it for free
The instructions displayed in the video are also in the YouTube description."
Capcut normalizes the input, so you cannot use Bas’ trick to decrease volume by -3dB like in Ripple to workaround the issue of bleeding (unless you trick out the CapCut, possibly by adding some loud sound in the song with decreased volume, something like presented here).
- (fixed) KaraFan Colab will be fixed on 27th at morning.
- There’s a workaround for people not able to split using Capcut. The app discriminate based on country (poor/rich) and paywalls Pro option.
The video demonstration for below
0. Go offline.
1. Install the Chinese version from capcut.cn
2. Use these files copied over your current Chinese installation, and don’t use English patch.
3. Open CapCut, go online after closing welcome screen, happy converting!
4. Before you close the app, go offline again (or the separation option will be gone later).
Before reopening the app, go offline again, open the app, close welcome screen, go online, separate, go offline, close. If you happen to missed that step, you need to start from the beginning of the instruction.
(replacing SettingsSDK folder no longer works after transition from 4.6 to 4.7, it freezes the app)
FYI - the app doesn’t separate files locally.
- Bas Curtiz found out that decreasing volume of mixtures for Ripple by -3dB eliminates problems with vocal residues in instrumentals in Ripple. Video.
This is the most balanced value, which still doesn't take too many details out of the song due to volume attenuation.
Other good values purely SDR-wise are -20dB>-8dB>-30dB>-6dB>-4dB> /wo vol. decr.
The method might be potentially beneficial for other models and probably work best for the loudest tracks with brickwalled waveforms.
- Stable 5.6 OpenCL (DirectML) version of UVR 5 GUI for Windows
Supporting AMD and Intel GPUs acceleration but no Roformers yet
Mac: https://github.com/Anjok07/ultimatevocalremovergui/releases/
(newer beta Roformer [with “roformer” in the installer name] supports both DirectML and CUDA out of the box already; for Mac M1 click).
- For CUDA (NVIDIA GPUs) - non-OpenCL installer in the name from here:
https://github.com/Anjok07/ultimatevocalremovergui/releases/download/v5.6/UVR_v5.6.0_setup.exe
(Following based on previous OpenCL build)
8GB VRAM for 3:00/3:30 tracks using MDX23C HQ model with 12GB VRAM probably enough for 5:00 track which is more than in CUDA.
Now the issue should be mitigated, and less memory crashes should occur.
Ensembles might require more memory due to memory allocation issues not met in CUDA before. Also, VRAM is fully freed only after closing the application.
Acceleration for only Demucs 2 (and 1?) arch on AMD is not supported. All others archs should work.
- Be aware that there was also full MPS (GPU) acceleration introduced for Mac M1 for all MDX-NET Original Models (HQ3, etc.), all MDX23C Models, all Demucs v4 models (no VR models acceleration on GPU). So don’t use Windows in VM to run UVR anymore, but separate using dmg installer from releases section (ARM). GPU acceleration is 3x faster than separation took on CPU before.
____
- “MDX23C-InstVoc HQ 2 is out as a VIP model [for UVR 5]! It's a slightly fine-tuned version of MDX23C-InstVoc HQ. The SDR is a tiny bit lower, but I found that it leaves less vocal bleeding.” ~Anjok
It’s not always the case, sometimes it can be even the opposite, but as always, all can depend on specific song.
- jarredou’s MDX23 2.2 Colab should allow separating faster, and also longer files now (tech details)
- All-in ensemble added for premium users of MVSEP - it has vocals, vocals lead, vocals back, drums, bass, piano, guitar, other. Basically 8 stems (and from drums stem you can further separate single percussion instruments using drumsep - up to 4 instruments, so it will give 10 stems in total).
- https://www.capcut.cn/ (outdated section: read)
Is a new Windows app which contains Ripple/SAMI-Bytedance inst/vocal model (not 4 stems like in Ripple).
“At the moment the separation is only available in Chinese version which is jianyingpro, download at capcut.cn [probably here - it’s where you’re redirected after you click “Alternate download link” on the main page, where download might not work at all]
Separation doesn't require sign up/login, but exporting does, and requires VIP.
Separated vocal file is encrypted and located in C:\Users\yourusername\AppData\Local\JianyingPro\User Data\Cache\audioWave”
The unencrypted audio file in AAC format is located at \JianyingPro Drafts\yourprojectname\Resources\audioAlg (ends with download.aac)
Drag and drop it in Audacity or convert to WAV (https://cloudconvert.com/aac-to-wav)
“To get the full playable audio in mp3 format a trick that you can do is drag and drop the download.aac file into capcut and then go to export and select mp3. It will output the original file without randomisation or skipping parts”
“Trying out Capcut, the quality seems the same as the Ripple app (low bitrate mp3 quality)
at least the voice leftover bug is fixed, lol”
Random vocal pops from Ripple are fixed here.
Also, it still has the same clicks every 25 seconds as before in Ripple.
Some people cannot find the settings on this screen in order to separate. Maybe it’s due to lack of Chinese IP, or Chinese regional settings in Windows, but logging wasn’t necessary from what someone told.
- Looks like the guitar model on MVSEP can pick up piano better than the available there piano model in lots of cases (isling)
- AudioSep has been released
https://github.com/Audio-AGI/AudioSep
(separate anything you describe)
https://replicate.com/cjwbw/audiosep?prediction=j7dsrvtbyxfm3gjax3vfzbf7py
(use short fragments as input)
https://colab.research.google.com/github/badayvedat/AudioSep/blob/main/AudioSep_Colab.ipynb (basic Colab)
https://huggingface.co/spaces/badayvedat/AudioSep (it’s down)
"so far it's ranged from mediocre to absolutely horrible from samples I've tried"
"So far[,] it does [a] great job with crowd noise/cheering."
Didn't pick piano.
Output is mono 32kHz. Where input is 30s, the output can be 5s.
- UVR started to process slower for some people using Nvidia 532 and 535 drivers (at least Studio ones on at least W11). More about the issue. Consider rolling back to 531.79.
“Took 10 seconds to run Karaoke 2 on a full song (~5[]mins), with the latest drivers it took like 20 minutes”. The problem may occur once you reboot your system.
- AMD GPU acceleration has been introduced in the official UVR repo under a new branch on GH. Beta as exe patch will be released in the following days. Currently, it supports only MDX-Net, but not MDX23C, and Demucs 4 models (not 3) and VR arch (5.0, but not 5.1).
Currently, GPU memory is not clearing, so you need a lot of VRAM in order to use ensembles.
- (x-minus) "Added additional download buttons when using UVR BVE model.**
Now you can download:
- song without backing vocals
- backing vocals
- instrumental without vocals
- all vocals" Anjok
- MacOS UVR versions should be fixed now - redownload the latest 5.6 patches. GPU processing on M1 is fully functioning with MacOS min. Monterey 12.3/7 (only VR models will crash with GPU processing). It’s very fast for the latest MDX23C fullband model - 11 minutes vs 1 hour on CPU previously.
- Cyrus version of MedleyVox Colab with chunking introduced, so you don't need to perform this step manually
https://colab.research.google.com/drive/1StFd0QVZcv3Kn4V-DXeppMk8Zcbr5u5s?usp=sharing
“Run the 1st cell, upload song to folder infer_file, run 2nd cell, get results from folder results = profit”
“one annoying thing is that is always converts the output to mono 28k”
- Separation times since the UVR 5.6 update increased double for some people. Almost the same goes to RAM usage.
Having lots of space on your system disk or additional partition assigned for pagefile can be vital in fixing some crashes, especially for long tracks. Be aware that CPU processing tends to crash less, but it's much slower in most cases.
"I realized that with 2-3h long audio files, I was able to use Demucs, after I added another 32GB of RAM. In total my system got 64GB and I increased the swap file to 128GB, which is located on an NVMe drive... so just in case the 64GB RAM are not enough, which I experienced with the "Winds" model, it's not crashing UVR, instead using the SWAP."
- Segments set to default 256 instead of 512 is ⅓ faster for the new MDX23C fullband model at least for 4GB cards. But it's still very slow on such RTX 3050 mobile variant (20 minutes for 3:40 song).
- Sometimes inverting vocals with mixture using MDX23C instead of using instrumental output can give better results and vice versa.
“Differences were more significant with D1581 [than fullband], but secondary vocals stem has "a bit" higher score” (click). Generally inversion of these MDX23C models (but not spectral) was giving sometimes better results.
- MedleyVox Colab preconfigured to use with Cyrus model
Newer model epochs can be found here:
https://huggingface.co/Cyru5/MedleyVox/tree/main
Q: What is isrnet?
A: It's basically just another model that builds on top of what I've built so far that performs better. That's the surface level explanation, at least.
- Settings for v2.2.2 Colab
If you stuffer from some vocal residues, try out these settings
BigShifts_MDX: 0
overlap_MDX: 0.65
overlap_MDXv3: 10
overlap demucs: 0.96
output_format: float
vocals_instru_only: disabled
Also, you can manipulate with weights.
E.g. different weight balance, with less MDXv3 and more VOC-FT.
- As an addition to AI-killing tracks section, and in response to deletion of "your poor results" channel, there was recently created a Gsheet with your problematic tracks to fill in. It is open to everyone to contribute.
- Video tutorial by Bas Curtiz how to install MedleyVox (based on Vinctekan fixed source). Cyrus trained a model. MD serves to separation of various singers from a track. It sometimes does a better job than BVE models in general.
Sadly, it has 24kHz output sample rate, but AudioSR works pretty good for upscaling the results.
https://github.com/haoheliu/versatile_audio_super_resolution
https://replicate.com/nateraw/audio-super-resolution
https://colab.research.google.com/drive/1ILUj1JLvrP0PyMxyKTflDJ--o2Nrk8w7?usp=sharing
Be aware that it may not work with full length songs - you might need to divide them into smaller 30 seconds pieces.
- "Ensemble 4/8 algorithms were updated on MVSep with new VitLarge23 model. All quality metrics were increased:
Multisong Vocals: 10.26 -> 10.32
Multisong Instrumental: 16.52 -> 16.63
Synth Vocals: 12.42 -> 12.67
Synth Instrumental 12.12 -> 12.38
MDX23 Leaderboard: 11.063 -> 11.098
I added Ensemble All-In algorithm which includes additionally piano, guitar, lead/back vocals. Piano and guitar has better metrics comparing to standard models, because they are extracted from high quality "other" stem. Lead/back vocals also has slightly better metrics.
piano: 7.31 -> 7.69
guitar: 7.77 -> 8.95" ZFTurbo
- New vocal model added on MVSEP:
"VitLarge23" it's based on new transformers arch. SDR wise (9.78 vs 10.17) it's not better than MDX23C, but works "great" for ensemble consisting of two models with weights 2, 1.
- MVSEP-MDX23-Colab fork v2.2.2 is out.
It is now using the new InstVocHQ model instead of D1581:
https://github.com/jarredou/MVSEP-MDX23-Colab_v2/ (dead)
Memory issues with 5:33 songs fixed (even 19 minutes long with 500K chunks supported)
It should be slightly faster than the previous version, as the extra processing for the fullband trick is not needed anymore with the new model.
Q: Why is "overlap_MDX" set to 0.0 by default in MVSEP-MDX23-Colab_v2 ?
A: because it's a "doublon" with MDX BigShifts (that is better)
- Stable final version of UVR v5.6.0 has been released along with MDX23C fullband model (the same as on MVSEP) - SDR is 10.17 for vocals & 16.48 for instrumentals.
It’s called MDX23C-InstVoc HQ.
https://github.com/Anjok07/ultimatevocalremovergui/releases/
Be aware it’s taking much more time to process a song with it, then all previous models. Also, it doesn’t require volume compensation set. It can leave more vocal residues than HQ_3 models for some songs. On the other hand, it can give very good results with song with “super dense mix like Au5 - Snowblind” but also for older tracks like Queen - March Of The Black Queen (always caused issues, but it gave the best result so far, although still lot of BV is missed).
Performance:
- 3:30 track with HQ_3 takes up to 24 minutes on i3-3217u while the new model takes 737 minutes (precisely 1:34 vs 41:00 for 15 seconds song).
- RTX 3060 12 GB - takes around 15 minutes to process a 25 minutes file with the new model.
- GTX 1080 Ti took about 4 minutes to process, about a 5 min 30 song
- If you upgraded from beta, Matchering might not work correctly. In order to fix the error:
Go to the Align tool.
Select another option under "Volume Adjustment", it can be anything.
Now, matchering should work. The fix may not apply for Linux installations.
- KaraFan original Colab seems to work now (v. 3.1) but one track with default settings takes 30 minutes for 3:37 track on free T4 (the last files processed are called Final) and it can get you disconnected from runtime quick (especially if you miss some multiple captcha prompts). V. 3.1 can have more vocal residues than in 1.x version and even more than in HQ_3 model on its own.
You might want to consider using older versions of KF with Kubinka Colab.
- Now 3.2 version was released with less vocal residues.
As mentioned before, after runtime disconnection error, output folder still constantly populated with new files, while progress bar is not being refreshed after clicking close or even after closing your tab with Colab opened.
- "Image-Line the company that made Fl Studio 21 took to instagram announcing a beta build that allows the end users to separate stems from the actual program itself, this is in beta and isn’t final product"
People say it's Demucs 4, but maybe not ft model and/or with low parameters applied or/and it's their own model.
"Nothing spectacular, but not bad."
"- FL Studio bleeds beats, just like Demucs 4 FT
- FL Studio sounds worse than Demucs 4 FT
- Ripple clearly wins"
- Org. KaraFan Colab with v. 3.0 should work with the large GPU option disabled (now done by default).
- You may be experiencing issues with KaraFan 3.0 alpha (e.g. lack of 5_F-music with which the result was better before), and using Kubinka Colab which uses the older version for now has some problems with GPU acceleration. Maybe the previous KF commit will work or even the one before (2.x is used here).
- New UVR beta patches for Windows/Mac/M1 at the bottom of the release note
https://github.com/Anjok07/ultimatevocalremovergui/releases/
Usually check for newer versions above, but this one currently fixes long error on using the new BVE model
- “The new BVE (Background Vocal Extractor) model [in UVR 5 GUI] has been released!
To use the BVE model, please make sure you use the UVR_Patch_9_18_23_18_50_BETA patch (Mac). Remember, it's designed to be used in a chain ensemble, not on its own. It's better to utilize it via "Vocal Splitter Options". ~Anjok”
Using Lead vocal placement = stereo 80% is still only available on X-Minus only. UVR GUI doesn't support this yet - it’s for the situation when your main vocals are confused with backing vocals.
- In the latest UVR GUI beta patch, vocal stems of MDX instrumental models have polarity flipped. You might want to flip it back in your DAW.
- Investigating KaraFan shapes issue > link
- New piano and guitar models added on MVSEP. Use other stem from e.g. “Ensemble 8 models” or MDX23 Colab or htdemucs_ft for better results.
- To separate electric and acoustic guitar, you can run a song (e.g. other stem) through the Demucs guitar model and then process the guitar stem with GSEP (or MVSEP model instead of one of these).
Gsep only can separate electric guitar so far, so the acoustic one will stay in the "other" stem.
- New UVR beta patch implements chain ensemble from x-minus for splitting backing and lead vocals. To use it:
1. Enable "Help Hints" (so you can see a description of the options),
2. Go to any option menu
3. Click the "*Vocal Splitter Options*"
4. From there you will see the new chain ensemble options.
Patch (patching from the app may cause startup issues)
- "New MDX23C model improved on [MVSEP] Leaderboard from 10.858 up to 11.042"
- "For those of you who were running into errors related to missing *"msvcp140d.dll"* and *"VCRUNTIME140D.dll"* after installing the latest patch, it's been fixed." -Anjok
- The UVR's latest beta 9 patch causes startup issue for lots of people on even clean Windows 10. No fix for it. Copying libraries manually or installing all possible redistributables doesn't work. In such case, use beta 8 patch.
- If you see an error that you're disconnected from KaraFan Colab, it can still separate files in the background and consume free "credits" till you click Environment>Terminate session. It happens even if you close the Colab.
So, you can see your GDrive output folder still constantly populated with new files, while progress bar is not being refreshed after error of runtime disconnection or even after Closing your tab with Colab.
- KaraFan got updated to 1.2 (eg. model picking was added). Deleting your old KaraFan folder on GDrive can be necessary to avoid an error now in Colab.
- KaraFun - next version of MDX23 fork (originally developed by ZFTurbo, enhanced and forked by jarredou) has been created by Captain FLAM (with jarredou’s assistance on tweaks).
Official Colab (video guide in case of problems)
Colab forked by Kubinka (can show error now after 1.2 update)
GUI for offline use: https://github.com/Captain-FLAM/KaraFan/tree/master
It gives very clean instrumentals with much less of consistent vocal residues than in MDX23 2.0-2.2 and Ripple/Bytedance.
(might have been changed) You can also disable SRS there to get a bit cleaner result, but in cost of more vocal residues. How detestable it will be without SRS, depends on a track - e.g. if it has heavy compressed modern vocals and lots of places with not busy mix (when not a lot of instruments play). Disabled SRS adds a substantial amount of information above 17.7kHz.
One of our users had problems caused seemingly by empty Colab Notebooks folder which he needed to delete. Could have been something else they did too, though.
- New epoch of new BVE model has been added to x-minus
“In some parts the new BVE is better, in some it's worse. Still a great model”
> To get better results, you can downmix the result to mono and repeat the separation
- For people having issues with Boosty x-minus payment:
https://boosty.to/uvr/posts/5d88402e-9eb1-4046-a00a-cf8b09e27561
- Sometimes for instrumental residues in vocals, AIs for voice recorded with home microphone can be used (e.g. Goyo [now Supertone Clear], or even Krisp, RTX Voice, AMD Noise Suppression, Elgato Wave Link 3.0 Voice Focus or Adobe Podcast as a last resort) it all depends on type of vocals and how destructive the AI can get.
- Izotope Ozone 11 has been released. It’s 1200$ for Advanced Edition. It’s the only version possessing Spectral Recovery. Music Rebalance is said to have Demucs instead of Spleeter now.
https://www.izotope.com/en/products/ozone.html
- Acon Digital has released Remix, their first plug-in capable of real-time separation to five stems: Vocals, Piano, Bass, Drums, and Other.
“Just listened to the demo, not great but still”
- RemFX for detection and removal of the following effects: chorus, delay, distortion, dynamic range compression, and reverb. Huggingface (currently stopped working) | Samples
The Colab is slow while downloading checkpoints from zenodo (400KB/s for 1GB file out of 6), later it stopped working.
Outputs in at least Huggingface are mono, may not work in every case, the website in general doesn't work well with big files, keep them short, 0-30 seconds.
Sometimes 30 seconds is still not enough on Colab and it throws OutOfMemoryError.
It's not better than our dereverb model in UVR.
To fix Colab:
“speechbrain lib API was totally changed in recent 1.0.0 version, it's working if you downgrade it:
!pip install speechbrain==0.5.16”
OG repo for running locally.
- Beta UVR patch also released for x86_64 & M1 Macs:
“If you have any trouble running the application, and you've already followed the "MacOS Users: Having Trouble Opening UVR?" instructions here, try the following:
Right-click the "Ultimate Vocal Remover" file and select "Show Package Contents".
Go to -> Contents -> MacOS ->
Open the "UVR" binary file.”
In case of further issues, check this out:
https://www.youtube.com/watch?v=HQsazeOd2Iw&feature=youtu.be
Looks like e.g. with Denoise Lite models it can ask for parameters. Set 4band_v3 and 16 channels, press yes on empty window.
“The Mac beta is not stable yet.” - Anjok
-"The new beta [UVR] patch has been released! I made a lot of changes and fixed a ton of bugs. A public release that includes the newest MDX23 model will be released very soon. Please see the change log via the following message - https://discord.com/channels/708579735583588363/785664354427076648/1145622961039101982"
Patch:
-"I found a way to bypass the free sample limits of Dango.ai. With VPN and incognito, when the limit appears, change the date on the computer or other device (I set the next day) and close and re-open the incognito tab. Sometimes it can show network error, in such case restart the VPN and re-enter in incognito again" Tachoe Bell
- Bas' guide to change region to US for Ripple on iOS
- Another way to use Ripple without Apple device
Sign up at https://saucelabs.com/sign-up
Verify your email, upload this as the IPA: https://decrypt.day/app/id6447522624/dl/cllm55sbo01nfoj7yjfiyucaa
Rotating puzzle captcha for TikTok account can be tasking due to low framerate. Some people can do it after two tries, others will sooner run out of credits, or completely unable to do it.
- Every 8 seconds there is an artifact of chunking in Ripple. Heal feature in Adobe Audition works really well for it:
https://www.youtube.com/watch?v=Qqd8Wjqtx-8
-The same explained on RX 10 example and its Declick feature:
https://www.youtube.com/watch?v=pD3D7f3ungk
- Ripple/SAMI Bytedance's API was found. If you're Chinese, you can go through it easier.
The sami-api-bs-4track (the one with 10.8696 SDR Vocals) - you need to pass the Volcengine facial/document recognition apparently only available to Chinese people
https://www.volcengine.com/docs/6489/72011
We already evaluated its SDR, and it even scored a bit better than Ripple itself.
This is the Ripple audio uploading API:
https://github.com/bitelchux/TikTokUploder/blob/2a0f0241a91b558a7574e6689f39f9dd9c39e295/uploader.py
there's a sample script on the volcengine SAMI page
"API from volcengine only return 1 stem result from 1 request, and it offers vocal+inst only, other stems not provided. So making a quality checker result on vocal + instrument will cost 2x of its API charging
something good is that volcengine API offers 100 min free for new users"
API is paid 0.2 CNY per minute.
It takes around 30 seconds for one song.
It was 1.272 USD for separating 1 stem out MVSEP's multisong dataset (100 tracks x 1 minute).
- (outdated) Using Ripple on an M1 remote machine turned out to be successful but very convoluted.
https://discord.com/channels/708579735583588363/708579735583588366/1143710971798507520
-It is possible that "a particular song that an older version of mdx23 (mdx23cmodel3.ckpt) has a much better extraction than D1581 and the current 4 model ensemble on MVSEP for preserving the instruments (also organ-like instruments)"
-Seems like Google raised Colab limit for free users from 1 hour to 5 hours. It depends on a session, but in most cases you should be able to perform tasks taking above 4 hours now.
-How to change region to US in Apple App Store to make "Ripple - Music Creation Tool" (SAMI-Bytedance) work.
https://support.apple.com/en-gb/HT201389
https://www.bestrandoms.com/random-address-in-us
Or use this Walmart address in Texas, the number belongs to an airport.
Do it in App Store (where you have the person-icon in top right).
You don't have to fill credit cards details, when you are rejected,
reboot, check region/country... and it can be set to the US already.
Although, it can happen for some users that it won't let you download anything forcing your real country.
"I got an error because the zip code was wrong (I did enter random numbers) and it got stuck even after changing it.
So I started from the beginning, typed in all the correct info, and voilà"
If ''you have a store credit balance; you must spend your balance before you can change stores''.
It needs (an old?) a sim card to log your old account out if necessary.
- Long awaited app made by Bytedance with one of their SAMI variants from MDX23 competition which holds top of our MVSEP leaderboard was published on iOS and for US region only
(with separate possibility to sign up for beta testing, also not for people outside US, and the app is in the official store already anyway, but it was before official release - at the end of June, so it's older news).
It's a multifunctional app for audio editing, which also contains a separation model.
It's free, called:
"Ripple - Music Creation Tool"
https://apps.apple.com/us/app/ripple-music-creation-tool/id6447522624
The app requires iOS 14.1
(it's only for iOS).
Output files are 4 stems 256kbps M4A (320 max).
Currently, the best SDR for public model/AI, but it gives the best results for vocals in general. For instrumentals, it rather doesn’t beat paid Dango.ai (and rather not KaraFan too).
"My only thought is trying an iOS Emulator, but every single free one I've tried isn't far-fetched where you can actually download apps, or import files that is"
Sideloading of this mobile iOS app is possible on at least M1 Macs.
"If you're desperate, you can rent an M1 Mac on Scaleway and run the app through that for $0.11 an hour using this https://github.com/PlayCover/PlayCover"
IPA file:
https://www.dropbox.com/s/z766tfysix5gt04/com.ripple.ios.appstore_1.9.1_und3fined.ipa?dl=0
"been working like a dream for me on an M1 Pro… I've separated 20+ songs in the last hour"
"bitrise.com claims to have M1s and has a free trial"
Scaleway method:
https://cdn.discordapp.com/attachments/708579735583588366/1146136170342920302/image.png
“keep in mind that the vm has to be up for 24 hours before you can remove it, so it'll be a couple bucks in total to use it”
"I used decrypted ipa + sideloadly
seems that it doesn't have internet access or something"
So far, Ripple didn't beat voc_ft (although there might be cases when it's better) and Dango. Samples we got months ago are very similar to those from the app, also *.models files have SAMI header and MSS in model files (which use their own encryption), although processing is probably fully reliable on external servers as the app doesn't work offline (also model files are suspiciously small - few megabytes, although it's specific for mobilenet models). It's probably not the final iteration of their model, as they allegedly told someone they were afraid that their model will leak, but better than the first iteration judging by SDR with even lossy input files.
Later they told that it’s different model than the one they previously evaluated, and that time it was trained with lossy 128kbps files due to some “copyright issues”.
Most importantly, it's the good for vocals, also cleaning vocal inverts, and surprisingly good for e.g. Christmas songs, (it handled hip-hop, e.g. Drake pretty well). It's better for vocals than instrumentals due to residues in other stem - bass is “so” good, drums also decent. Vocals can be used for inversion to get instrumentals, and it may sound clean, but rather not as good as what 2 stem option or 3 stem mixdown gives.
Other stem residues appear due to the fact they told the other stem is taken from the difference of all remaining stems - they didn’t train the other stem model to save on separation time.
"One thing you will notice is that in the Strings & Other stem there is a good chunk of residue/bleed from the other stems, the drum/vocal/bass stems all have very little to no residue/bleed" doesn't exist in all songs.
It's fully server-based, so they may be afraid of heavy traffic publishing the app worldwide, and it's not certain that it will happen.
Thanks to Jorashii, Chris, Cyclcrclicly, anvuew and Bas.
Press information:
https://twitter.com/AppAdsai/status/1675692821603549187/photo/1
https://techcrunch.com/2023/06/30/tiktok-parent-bytedance-launches-music-creation-audio-editing-app/
Beta testing
- Following models added on MVSep:
UVR-De-Echo-Aggressive
UVR-De-Echo-Normal
UVR-DeNoise
UVR-DeEcho-DeReverb
They are all available under the "Ultimate Vocal Remover HQ (vocals, music)" option (MDX FoxJoy MDX Reverb Removal model is available as a separate category).
- If you looked for possibility to pay for Dango using Alipay - they recently introduced the possibility to link foreign cards, and if that option fails (sometimes does), you can open 6 months “tourcard”, and open new later if necessary, but only Visa, Mastercard, Diners Club and JCB cards are supported to top tourcard up
https://ltl-beijing.com/alipay-for-foreigners/
- Dango no longer supports Gmail email accounts
- New piano model added on MVSEP. SDR-wise it’s better than GSep, but GSep is probably also using some kind of processing in order to get better separation results, but e.g. Dango instrumentals can be inverted to get just vocals despite the fact they claim to use some recovery technology.
- arigato78 method for main vocals
-Captain Curvy method for instrumentals added in instrumentals models list section (the top link)
- For canceling room reverb check out:
Reverb HQ
then
De-echo model (J2)
- Sometimes vox_ft can pick up SFX
- Install UVR5 GUI only in the default location picked by the installer. Otherwise, you might get python39.dll error on startup. If you see that error after installing the beta patch, reinstall the whole app.
- Few of our users finally evaluated sonically new dango.ai 9.0 models. Turns out the models are not UVR's (or no longer), and actually give pretty close results to original instrumentals, but not so good vocals.
"It's slightly better but still voc_ft keeps more reverb/delays
but again, it's 99% close, Dango has maybe more noise reduction" maybe even less instrumental residues (can be a result of noise reduction).
"A bit cleaner than voc_ft in terms of having synths/instruments, but they do sound a bit filtered at times. [In] overall it's close tho"
"I discovered Dango's conservative mode keeps instrumentals even fuller, but might introduce some background vocals
still quite better than what we have.
I'm still surprised how it's so clean, as if not having vocal residues like any other MDX model. Sometimes the Dango sounds like a blend of VR's architecture, but I'm probably wrong, it could be the recovery technology" - becruily
https://tuanziai.com/vocal-remover/upload
You must use the built-in site translate option in e.g. Google Chrome, because it's Chinese.
On Android, it may not work correctly. In case of further issues, use Google Translate or one of Yandex apps with image to text translators.
You are able to pay for it using Alipay outside China.
Dango redirects to Tuanziai site - it's the same.
https://tuanziai.com/encouragement
Here you might get 30 free points (for 2 samples) and 60 paid points (for 1 full songs) "easily".
Dango.ai scores bad in SDR leaderboards due to recovery algorithms applied. Similar situation probably like in GSep.
- New BVE model on X-Minus for premium users. One of, if not the best so far. It uses voc_ft as a preprocessor.
"BVE sounds good for now but being an (u)vr model the vocals are soft (it doesn’t extract hard sounds like K, T, S etc. very well)"
"Pretty good, if still [in] training. Seems to begin a phrase with a bit of confusion between lead and backing, but then kicks in with better separation later in the phrase. Might just be the sample I used, though."
- Jarredou published the final 2.2 version of MDX23 Colab (don't confuse it with MDX23C single models v3 arch) - gives more vocal residues than 2.0/2.1, but better SDR. Now it has SRS trick, bigshifts, new fine-tuning, separated overlap parameters for MDX, MDXv3 and Demucs models, and also possess one narrowband MDX23C model D1581 among other MDX ones, which states a new set of models now (also said to use VOC-FT Fullband SRS instead of UVR-MDX-Instr-HQ3, although HQ3 is still listed during processing). You can also use faster optional 2 stem only output (demucs_ft vocal stem is used here only). Float parameter returns WAV 32-bit. Don’t set overlap v3 to more than 10, or you’ll get error. It can be way more frequent with odd values.
Changing weights added: “For residues, I would first try a different weight balance, with less MDXv3 and more VOC-FT, as model D1581, and current MDXv3 models in general tend to have more residues than VOC-FT.”
- New "8K FFT full band" model published on MVSEP. Currently, a better score than only 2.2 Colab above from commonly available solutions, although more vocal residues than current default on MVSEP at least in some cases, and “voice sounded more natural [in default] than the new 10 SDR model” but in some problematic songs it can even give the best results so far.
"Sometimes 8K FFT model is false detect the vocals, in the vocal stem synth was treated as vocal. On instrumental stem, mostly are blur result compared with 12K FFT. But 12K FFT seems to be some vocal residue but very less heard (like a whisper) and happened for several songs, not all songs."
- "The karaoke ensemble works best with isolated vocals rather than the full track itself" Kashi
- Center isolation method further explained in Tips to enhance separation, step 19
- VR Kara models freeze on files over ~6 minutes in UVR beta 2 (GTX 1080).
>Divide your song into two parts.
- New public dataset published by Moises (MoisesDB). There are some problems with downloading it now, and it’s 82,7GB and link expires during downloading after 600 seconds. Not enough for 30MB/s, but good for 10Gbps one. Moises team works on the issue. Probably it's fixed already.
- RipX inside the app uses UVR for gathering stems now. Consider also comparing its stem cleanup feature to RX 10 debleed in RX Editor.
- “RipX is badass for removing residues and harmonics from vocals. The ability to remove harmonics & BGVs using RipX is amazing but is very tedious but so far so good” (Kashi)
- Sometimes using vocal model like voc_ft on the result from instrumental model might give less vocal residues or sometimes even none (Henry)
- mvsep1.ru from now on, contains a content of mvsep.com, so without MDX23/C and login features, while mvsep.com has the richer content of mvsep1.ru
The old leaderboard link has changed and is now:
https://mvsep1.ru/quality_checker/leaderboard2.php?sort=instrum
- old domain is also fixed now, redirecting leaderboard links.
If you’re uploading in quality checker is stopped, clear your browser and start over.
- Dereverb and denoiser for VR arch is not compatible with any VR Colab and manual installation of such model will fail with errors. It requires modifying nets and layers. More
- New best ensemble (all Avg/Avg)
(read entries details on the chart for settings - they can have very time-consuming parameters and differ in that aspect)
#1 MDX23C_D1581 + Voc FT | #2 MDX23C_D1581 + Inst HQ3 + Voc FT | #3
MDX23C_D1581 + Inst HQ3 + Voc FT
Be aware that above can sound noisy/have vocal leaks at times; consider using HQ_3 or kim inst then, also:
- The best ensembles so far in Kashi's testing for general use:
Kim Vocal 2 + Kim FT other + Inst Main + 406 + 427 + htdemucs_ft avg/avg, or:
Voc FT, inst HQ3, and Kim FT other (kim inst)
“This one's much faster than the first ensemble and sometimes produces better results”
It all depends on a song. Also, sometimes "running one model after another in the right order can yield much better results than ensembling them".
- Disable "stem combining" for vocal inverted against the source. Might be less muddy, possibly better SDR.
It's there in MDX23C because now the new arch supports multiple stems separation in one model file.
- Disabling "match freq cutoff" in advanced MDX settings seems to fix issues with 10kHz cutoff in vocals of HQ3 model.
- New explanations on Demucs parameters added in Demucs 4 section
(shifts 0, overlap 0.99 won in SDR vs shifts 1, overlap 0.99 and even shifts 10, overlap 0.95)
- "Last update of Neutone VST plugin has now a Demucs model to use in realtime in a DAW
(it's a 'light' version of Demucs_mmi)
https://neutone.space/models/1a36cd599cd0c44ec7ccb63e77fe8efc/
It doesn't use GPU, and it's configured to be fast with very low parameters, also the model is not the best on its own. It doesn't give decent results, so it's better to stick to other realtime alternatives (see document outline)
- Turns out that with a GPU with lots of VRAM e.g. 24GB, you can run two instances of UVR, so the processing will be faster. You only need to use 4096 segmentation instead of 8192.
SDR difference between overlap 0.95 and 0.99 for voc_ft MDX model in (new/beta) UVR is 0.02.
0.8 seems to be the best point for ensembles
12K segmentation performed worse than 4K SDR-wise
- Recommended balanced values between quality and time for 6GB graphic cards in the latest beta:
VR Architecture:
Window Size: 320
MDX-Net:
Segment Size: 2752 (1024 if it’s taking too long)
Overlap: 0.7-/0.8
Demucs:
Segment: Default
Shifts: 2 (def)
Overlap: 0.5
(exp. 0.75,
def. 0.25)
"Overlap can reduce/remove artifacts at audio chunks/segments boundaries, and improve a little bit the results the same way the shift trick works (merging multiple passes with slightly different results, each with good and bad).
But it can't fix the model flaws or change its characteristics"
“Best SDR is a hair more SDR and a shitload of more time.
In case of Voc_FT it's more nuanced... there it seems to make a substantial difference SDR-wise.
The question is: how long do u wanna wait vs. quality (SDR-based quality, tho)”
- A script with guide for separating multiple speakers in a recording added
- If you're stuck at 5% of separation in UVR beta, try to divide your audio into smaller pieces (that's beta's regression)
- A new separation site appeared, giving seemingly better results than Audioshake:
“Guitar stem seems better than Demucs, piano maybe too. Drums sound like Spleeter. Vocal bleeds in most of the stems, or not vocals are picked up, so they end up in the synths. But that's just from one song test” becruily
- Drumsep Colab now has GPU acceleration and much better max quality optional settings
- 1620 MDX23C model added on x-minus. Opposing the model on UVR, it's fullband and not released yet (16.2 SDR).
"Even if the separations have more bleeding than VOC-FT (and it's an issue), the voice sound itself is much fuller, "in your face" compared to VOC-FT, that I now find it like blurry sounding compared to MDXv3 models.
I think that's why the new MDXv3 models are scoring better despite having more bleeding (at the moment, like I said before, trainers/finetuners have to get familiar with new arch, and that will probably help with that new bleed issue)."
- New MDX23C model added on MVSEP (better SDR - 16.17)
- UVR beta patch 2 repairing no audio issue with GPU separation on the GTX 1600 series using MDX23C arch. Fixes some other bugs too.
- Narrowband MDX23C vocal model (MDX23C_D1581 a.k.a. model_2_stem_061321) trained by UVR team has been released. SDR is said to be better than voc_ft (but the latter was evaluated with older non-beta patch). Be aware that CPU processing returns errors for MDX23C models, at least on some configs (“deserialize model on CUDA” error). Fullband models will be released in a few weeks (and as it was usually before, on x-minus first for a few weeks later). Download (install beta patch first and drop it into the MDX-Net models folder). The patch is for only Windows now, with an upcoming Mac patch planned later. For Linux, there's probably a source of the patch already out.
MDX23C_D1581 parameters are set up with its yaml config file and its n_fft value is 12288, not 7680. It has cutoff at 14.7khz (while VOC-FT cutoff is 17.5khz)
- "(Probably all) models are stereo and can't handle mono audio. You have to create a fake stereo file with the same audio content on the L and R channel if the software doesn't make it by itself." Make sure that the other channel is not empty when isolation is executed - it can produce silent bleeding of vocals in the opposite channel (happens in e.g. MDX23 and GSEP, and errors with mono in MDX-Net)
- ”For Unbound local” error while you do anything in UVR since the new model installation, you might be forced to rollback the update
- Clear the Auto-Set Cache in the MDX-Net menu if you set wrong parameter and end up with error
- Pitch shift is the same as soprano mode except in the GUI beta you can choose how many semitones to pitch the conversion
- Dango.ai released a 9.0 model. We received a very positive report on it so far.
- UVR beta patch released. Potentially new SDR increases with the same models.
Added segmentation, overlap for MDX models, batch mode changes.
Soprano trick added. Basically, you can set it by semi-tones.
Support for MDX-NET23 arch. For now, it uses only basic models attached by Kuielab (low SDR, so don't bother for now), but UVR team already trained their own model for that arch, which will be released later, and a few weeks after x-minus and MVSep. And it's performing well already. Wait™. Don't exceed an overlap 0.93-0.95 for MDX models, it's getting tremendously long with not much of a difference, 0.8 might be a good choice as well. Also, segments can ditch the performance AF. 2560 might be still a high but balanced value.
Sadly, it looks like max mag for single models is no longer available - you can use it only under Ensemble Mode for now.
Q: What is Demucs Pre-process model?
A: You can process the input with another model that could do a better job at removing vocals for it to separate into the other 3 stems
Beta UVR patch link
- "Post-Process [for VR] has been fixed, the very end bits of vocals don't bleed anymore no matter which threshold value is used"
- New BVE model will be ready at the beginning of August (Aufr33).
- MDX23C by ZFTurbo model(s) added on mvsep.com. They're trained by him using the new 2023 MDX-Net V3 arch.
Slightly worse SDR than MDX23 2.1 Colab on its own.
Might be good for rock, the best when all three models are weighted/ensembled)
- MDX23C ensemble/weighted available on mvsep.com for premium users (best SDR for public 2 stem model).
It might still leave some instrumental residues in vocals of some tracks (which can be cleared with MDX-UVR HQ_3 model) but it can be also vice versa - the same issue as kim vocal models, where the vocals are slightly left in the instrumentals [vs e.g. MDX23 2.1 free of the issue]
On some Modern Talking and CC tracks it can give the best results so far).
- If you have problems with “Error when uploading file” on MVSEP, use VPN. Similar issues can happen for free X-Minus for users in Turkey.
- lalal.ai cooperation with MVSEP was fake news. Go along.
- As for Drumsep, besides in fixed Colab, you can also use it (the separation of single percussion instruments from drums stem) in UVR GUI. How to do this:
Go to UVR settings and open the application directory.
Find the folder "models" and go to "demucs models" then "v3_v4"
Copy and paste both the .th and .yaml files, and it's good to go.
Overlap above 0.6 or 0.7 becomes placebo, at least for dry track, with no effects.
- Drumsep benefits from shifts a lot (you can use even 20).
- For better results, test out potentially also -6 semitones in UVR beta, or with 31183Hz sample rate with changed tempo.
12 semitones from 44100Hz is 22050 and should be rather less usable in most cases, the same for tempo preservation on.
- If you have a long band_net error log while using DeNoise model by Fox Joy in UVR, reinstall the app.
- It can happen that every second separation using MDX Colab will fail due to memory issues, at least with Karaoke 2 model.
- New fine-tuned vocal model added to UVR5 GUI download center and HV Colab (slightly better SDR than Kim Vocal 2) it's called "UVR-MDX-Net-Voc_FT" and is narrowband (because it's based on previous models).
- Audioshake 3 stem model is added to https://myxt.com/ for free demo accounts. Unfortunately, it has WAVs with 16kHz cutoff which Audioshake normally doesn't have. No other stem. Results, maybe slightly better than Demucs.
Might be good for vocals.
- Spectralayers 10 received an update of an AI, and they no longer use Spleeter, but Demucs 4, and they now also good kick, snare, cymbals separation too. Good opinions so far. Compared to drumsep sometimes it's better, sometimes it's not. Versus MDX23 Colab V2, instrumentals sometimes sound much worse. “SpectraLayers is great for taking Stems from UVR and then carrying on separating further and editing down. (...) Receives a GPU processing patch soon”
- (? some) MDX Colabs started causing errors of insufficient driver version.
> "As a temp workaround you can go to "Tools" in the main menu, and "Command Palette", and search for "Use fallback runtime version", and click on it, this will restart the notebook with the previous Ubuntu version in Colab, and things should works as they were before (at least till mid July or earlier [how it was once] where it is currently scheduled to be deleted)" probably it will be fixed.
X: Some people have an error that fallback runtime is unavailable.
- New v2 version of ZFTurbo's MDX23 Colab released by jarreadou (now also with denoiser off memory fix added). Now it should have less bleeding in general.
It includes models changed for better ones (Kim Vocal 2 and HQ_3), volume compensation, fullband of vocals, higher frequency bleeding fix. It all manifests in increased SDR.
Instrum is inverted of vocals stem
Instrum2 is the sum of drums+bass+other stems (I used to prefer it, but most people rarely see any difference between both, and it also depends on specific fragments, although instrum gets better SDR and is less muddy, so it’s rather better to stick with instrum)
If your separation ends up instantly with path written below, you wrongly wrote it in the cell.
Simply remove the `file - name.flac` at the end and leave only path leading to a file.
It's organized in a way that it catches all files within that path/folder.
Suggestion: go to drive.google.com and create a folder `input`,
and drop the tracks you want to process in there.
When the process is done, delete them, and add others you want to process.
Overlap large and small are the main settings, higher values = slightly higher score, but way longer processing.
Colab doesn't allow much higher value for chunk size, but you can try little higher ones and see when it crashes because of memory. Higher chunk size give better results.
- Updated inference with voc_ft model (Colab v2.1 has denoiser now on, but updated inference not and is essentially what 2.2 currently is).
- Volume compensation fine-tuning - it is in line 359 (voc_ft), 388 (for ensembling the vocals stem), 394 (for HQ_3 instrumental stem inversion).
- chunk_size = 500000 will fail with 5:30 track, decrease it to at least 300K in such case.
Overlap 0.8 is a good balance between duration and quality.
- In case of system error wav not found, simply retry separation.
Nice instruction how to use the Colab.
The v. 2.1 Colab was firstly evaluated with lower parameters, hence it received slightly worse SDR. Then it was evaluated again and got better score than v2.
WiP Colabs
- 2.2 Beta 1 (no voc_ft yet)
- 2.2 Beta 1.5
- 2.2 Beta (1.5.1, inference with voc_ft, replace in the Colab above; no fine-tuning)
- v2.2 beta 2/3 (working inference) (MDX bigshifts, overlap added, fine-tuning, no 4 stems > experimental, no support for now, 22 minutes for vocals only, mdx: bsf 21, ov 0.15, 500k, 5:30 track)
- v2.2 (w/ voc_ft inference) pre beta 3 w/o MDX v3 yet - comment out both bigshifts in the cell - they won’t work
- current beta link (WiP, might be unstable at times; e.g. here for 19.07 bigshifts doesn’t work, and you need to look for working inference in history or delete the two bigshifts references in the cell; doesn’t seem that MDX v3 model is here yet)
In general -
MDX23 is quite an improvement over htdemucs_ft (...).
Drum stem makes htdemucs_ft sound like lossy in comparison, absolutely beautiful
Bass is significantly more accurate, identifies and retains actual bass guitar frequencies with clarity and accuracy
"Other", equally impressive improvement over htdemucs_ft, much more clarity in guitars"
And problems with vocals they originally described are probably fixed in V2 Colab.
- “I just added 2 new denoise models that were made by FoxJoy. They are both very good at removing any residual noise left by MDX-Net models. You can find them both in the "Download Center". - Anjok
Be aware that they're narrowband (17.7kHz cutoff). Good results.
To download models from Download Center -
In UVR5 GUI, click the tools icon > click Download Center tab > Click radio button of VR architecture > click dropdown > select the model > hit Download button > wait for it to download... Profit.
- New MDX-UVR “HQ_3” model released in UVR5 GUI! The best SDR for a single instrumental model so far. Model file (but visiting download center is enough). On X-Minus I think too.
-HQ_3 model added to MDX Colab (old)
-HV just made a new version of her own updated MDX Colab with all new models, including HQ_3. It lacks e.g. Demucs 2 for Instrumentals of vocal models, but in return it allows using YouTube and Deezer links for lossless tracks, with providing ARL, and allows specifying manually more than one file name to process at the same time. Also, for any new models in the future, there's optional input for model settings, to bypass parameters of parameters autoloader. IRC, the Colab stores its files in different path, so be aware about it when uploading tracks for separations on GDrive.
- she has added volume compensation in new revision (they’re applied automatically for each model)
In previous MDX Colabs there were also min, avg, max, and chunks, but they're gone in HV Colab.
- HV also made a new VR Colab which irc, now don’t clutter all your GDrive, but only downloads models which you use (but without VR ensemble) and probably might work without GDrive mounting, but it lacks VR ensemble.
- New MDX models added to both variants of MVSep (Kim inst, Vocal 1/2, Main [vocal model], HQ_2)
- ZFTurbo’s MDX23 code now requires less GPU memory. “I was able to process file on 8 GB card. Now it's default mode.”: 6GB VRAM is not enough. Lowering overlaps (e.g. 500000 instead of 1000000) or chunking track manually might be necessary in this case. Also now you can control everything from options: so you can set chunk_size 200000 and single ONNX. It can possibly work with 6GB VRAM that way.
Overlap large and small - controls overlap of song during processing. The larger value the slower processing but better quality (both)
If you have fail to allocate memory error, use --large_gpu parameter
Sometimes turning off use large GPU and reducing chunk size from 1000000 to 500000 helps
- Models/AIs of the 1st and 2nd place winners in MDX23 music challenge (ByteDance’s and quickpepper947’s) sadly won’t be released to the public (at least won’t be open-sourced). Maybe in June, ByteDance will be released as an app in worse quality.
Judging by the few snippets we had:
"the vocal output, yes, better than what can be achieved right now by any other model, it seems.
the instrumental output... meh. I can hear vocals in it, on a low volume level." but be aware that improved their model by the time by a lot.
- MDX23 4 stem model and source code with dedicated app by ZFTurbo (3rd place) was released publicly with the whole AI and instructions how to run it locally. No longer requires minimum 16GB VRAM Nvidia GPU. It even has a neat GUI (3rd place in leaderboard C, better SDR than demucs ft). You can still use the model online on mvsep1.ru (now mvsep.com).
The command:
"conda install -c intel icc_rt"
SOLVES the LLVM ERROR
For above, you can get less vocal residues by replacing the Kim Vocal 1 model there manually by newer Kim Vocal 2 and kim inst by and Kim Inst with UVR Inst HQ 292 (“full 292 is a lot more aggressive than kim_inst”).
jarredou forked it with better models and settings already.
Short technical summary of ZFTurbo about what is under the hood and small paper.
From what I see in the code, it uses inverted vocals output for instrumentals from - Demucs ft, with - hdemucs_mmi, and - Kim vocal 1 and - Kim inst (ft other). More explanations in MDX23 dedicated section of this doc.
- jarreadou made a Colab version of ZFTurbo MDX23:
"(It's working with `chunk_size = 500000` as default, no memory error at this value after few tests with Colab free)
Output files are saved on Colab drive, in the "results" folder inside MVSep installation folder, not in *your* GDrive."
On 19.05 its SDR was tested, and had better score for instrumentals than UVR5 ensemble for that time being. Currently not, but there are new versions of the Colab planned.
- ByteDance-USS was released with Colab by jazzpear. It works better than zero-shot for SFX and “user-friendly wise” while zero-shot stil better for instruments.
"https://www.dropbox.com/sh/fel3hunq4eb83rs/AAA1WoK3d85W4S4N5HObxhQGa?dl=0
Queries for ByteDance USS taken from the DNR dataset. Just DL and put these on your drive to use them in the Colab as queries."
QA section added.
- The modified MDX Colab - now with automatic models downloading (no more manual GDrive models installations) and Karaoke 2 model.
> Separate input for 3 models parameters added, so you don’t need to change models.py every time you switch to some other model. Settings for all models listed in Colab. From now on, it uses reworked main.py and models.py (made by jarredou) downloaded automatically. Don’t replace models.py from packages with models from here now. Now denoiser optionally added!
- MDX Colab with newer models is now reworked to use with current Python 3.10 runtime which all Colabs now use.
- Since 28.04 lots of Colabs started having errors like "onnxruntime module not found". Probably only MDX Colab (was) affected.
(not needed anymore)
> "As a temp workaround you can go to "Tools" in the main menu, and "Command Palette", and search for "Use fallback runtime version", and click on it, this will restart the notebook with the previous python version, and things should works as they were before"
- OG MDX HV Colab is (also) broken due to torch related issues (reported to HV). To fix it, add new code row with:
!pip install torch==1.13.1
below mounting and execute it after mounting
> or use fixed MDX Colab with newer models and fix added (now with also old Karaoke models).
- While using OG HV VR Colab, people are currently encountering issues related to librosa. The issues are already reported to HV (the author of the Colab).
> use this fixed VR Colab for now (04.04.23). (the issue itself was fixed by uncommenting librosa line and setting 0.9.1 version - deleted "#" before the lines in Mount to Drive cell, now also fresh installation issues are fixed - probably the previous fix was based on too old HV Colab revision). VR Colab is not affected by May/April runtime issues.
- If you have fast CPU, consider using it for ensemble if you have only 4GB VRAM, otherwise you can encounter more vocal residues in instrumentals. 11GB VRAM is good enough, maybe even 8GB.
- New Kim's instrumental "ft other" model. Already added to UVR's download center with parameters.
Manual settings - dim_f = 3072 n_fft = 7680 https://drive.google.com/drive/folders/19-jUNQJwols7UyuWO5PWWVUlJQEwpn78
(Unlike HQ models, it has cutoff, but better SDR than even inst3/464, added to Colab)
- Anjok (UVR5) "I released an additional HQ model to the Download Center today. **UVR-MDX-NET Inst HQ 2** (epoch 498) is better at removing long drawn out vocals than UVR-MDX-NET Inst HQ 1." It has already evaluated slightly better SDR vs HQ_1 both for vocals and instrumentals (HQ_1 evaluation was made once more since introducing Batch Mode which slightly decreases SDR for only single models vs previous versions incl. beta, but mitigates an issue when there are sudden vocal pop-ins using <11GB VRAM cards)
- Anjok (UVR5, non-beta) “So I fixed MDX-Net to always use Batch Mode, even when chunks are on. This means setting the chunk and margin size will solely be for audio output quality. Regardless of PC specs, users will be able to set any chunk or margin size they wish. Resource usage for MDX-Net will solely depend on Batch Size.”
Edit. Batch size set to default instead of chunks enabled on 11GB cards for ensemble achieves better SDR, but separation time is longer.
- Public UVR5 patch with batch mode and final full band model was released (MDX HQ_1)
- 293/403 and 450/498 (HQ_1 and 2) full band MDX-UVR models added to Colab and (also in UVR) (PyTorch fix added for Colab)
- Wind model (trumpet, sax) beside x-minus, added also to UVR5 GUI
You'll find it in UVR5 in Download Center -> VR Models -> select model 17
(10 seconds of audio separated with Wind model, from a 7-min track, takes 29 minutes to isolate on a 3rd gen i7 - might be your last resort if it crashes your 4GB VRAM GPU as some people reported)
- (x-minus/Aufr33) "1. **Batch mode** is now enabled. This greatly speeds up processing without degrading quality.
2. The **b.v.** models have been renamed to **kar**.
3. A new **Soprano voice** setting has been added for songs with the high-pitched vocals.
*This only works with mdx models so far.*"
It slows down the input file similarly to the method we described in our tip section below.
- New MDX23 vocal model added to beta MVSEP site.