RuntimeError: ""
Traceback Error: "
If you have these two lines without any text at the beginning of the error log as above using AMD GPU on every attempt of separation with GPU Conversion option turned on in UVR, for all archs and models
- you probably use outdated GPU drivers and/or Windows not compatible with newer drivers.
>Go to Ultimate Vocal Remover\torch_directml and replace DirectML.dll from C:\Windows\System32\AMD\ANR (make a backup before).
>Experimentally, you can use this older 1.9.1.0 version of the library. Restart UVR after replacing the file!
If you use an incompatible library version, you’ll encounter the “Unhandled exception” startup issue. The same fix might work on Intel GPUs.
Be aware that the linked older version of the library might cause additional noise for MDX-Net v2 models like HQ_X (the issue is gone when you turn off GPU Conversion).
The same Runtime Error: “” also happens on Mac Pro 2009, at least on some older UVR versions. If updating UVR won't help, turn off GPU Conversion. It might be too old.
Theoretically all DX12 GPUs should be compatible with DirectML, but min. VRAM working with Roformers with low chunk_sizes is rather 4GB, and older NVIDIA GPUs than Maxwell might fail to switch to DirectML at least on 5.6.1 (in 5.6.0 and older, the option might be called OpenCL, and was probably deleted in Roformer patches).
- All MDX-Net v2 models (maybe beside 4 stem variants), have so called MDX noise, which can be cancelled by using Options>Advanced MDX-Net Settings>Denoise Output>Standard (or Model). Just with older directml.dll it's more noisy, and dedicated models don't work so efficiently with this more elevated noise.
- At least beta #2 Roformer update caused some stability and performance issues with other archs than Roformers for some people when specific parameters started to take more time than before.
Roll back to stable 5.6 (non-5.6.1) in these cases if necessary (but you won't have Roformers support). Possibly make a copy of the old installation. Your configuration files might be lost. You can use two installations at the same time (or at least when one, e.g. Roformer patch is installed or symlinked in the default location).
- (I think I covered that issue above more thoroughly)
Roformer models in at least patch #2 work only in “Multi-Stem” mode in UVR. Using them in Ensemble causes layers errors (you can use manual Ensemble instead).
Iirc, it’s caused by yaml config where instead of Instrumental + Vocals (with V as capital letter) there’s written other + vocals, and you need to change it. Iirc it doesn’t happen on models downloaded from Download Center as Anjok was fixing the issue, but the problem might still exist in yamls of some custom models outside the center
- If you have sudden issues with not being able to separate, try to reinstall the app, and/or possibly make sure you didn’t turn on some power saving option in your laptop. Plus, you can simply try to reopen UVR (few fail tries on incompatible DirectML.dll with your GPU driver/OS will hang UVR on “Loading Model” till you close UVR manually from Task Manager).
- When using e.g. BS-Roformer SW: “RuntimeError: "The size of tensor a (2) must match the size of tensor b (6) at non-singleton dimension 0"
Traceback Error: "”
> Replace the yaml config manually in:
C:\Users[User]\AppData\Local\Programs\Ultimate Vocal Remover\models\MDX_Net_Models\model_data\mdx_c_configs
Restart UVR, start over. Make sure you were asked to replace it. If not, the yaml was maybe wrongly picked anyway. Then edit in the Edit config menu.
_______
You’ll find more UVR troubleshooting in this section
_____
Problems fixed in newer patches
- (deprecated since patch #10 - now convert dim_t it to chunk_size) dim_t = 1101 seems to be a sweet spot in terms of speed/SDR according to measurements (although on 1 minute files); use 1120 if UVR refuses to accept 1101 in GUI (or edit yaml file)
- (deprecated since patch #10) Some Roformer configs have wrong dim_t at the bottom of the yaml by default (e.g. 256), change it at the bottom of the yaml config for better SDR (not the one at the top), e.g. to 1101 (more explanations on it later).
- (fixed in patch #10) VIP code in Roformer beta patch #2-9 (and probably #1) doesn’t work -
Download all the VIP models you need before patching older 5.6 to beta Roformer or use two installations of the UVR if you can’t use patch #10 with the fix.
- (fixed in patch #6) People experience All stems error with viperx’ 12xx models in newer versions of UVR Beta Roformer patch (patch #2 was the last confirmed to work with these older models)
- (fixed in patch #10) mlp_expansion_factor: 1 or (when mlp line is deleted from yaml) mismatch for MelBand Roformer error
You probably use older Roformer patch incompatible with newer models (e.g. #2)
It also appears when you wrongly set v2 model type.
Fixed in the beta patch #3 and #4 for all platforms
- Don't set overlap higher than 11 for 1101 dim_t (at the bottom of yaml file in the “inference” section, not above) and overlap 8 for 801 - these two are the fastest settings before stem misalignment issues occur. Otherwise, it can lead occasionally to some effects or synths missing from the instrumental stem (although some rules can be broken here with various settings). Also, the problems with clicks are alleviated with these good settings.
- In beta #2 patch, best measured SDR for both Mel and BS-Roformers is when dim_t = 1101 in the inference section of yaml config and when overlap is set to 2 in GUI (although 1 wasn’t tested, and is actually lower). But the last beta patches, all bigger overlap values are slower, so SDR might be higher with higher values.
Be aware that it will increase separation time. Maximum allowed value before error is 1801, but 1501 or 1601 depending on a model will be the max reasonable for experiments before some unwanted downsides of too high or too low dim_t appear (disappearing of some stem elements). In some specific cases, 1333 (or potentially 1301) was giving better results than 1101 or 1501, but it depended on song length - usually it happened on short fragments.
- Instruction for overlap and dim_t above applies to other Roformer models as well, and not only those in Download Center. With the instructions, you can achieve faster separation times, as you’re not forced to use the most time-consuming overlap 2 in older patches to avoid stem misalignment issues
_______________
Infos and fixes for older patch #1/2
(with matching overlap (reversed) and dim_t necessity)
- To avoid separation errors for 4GB VRAM and AMD/Intel GPUs using Roformers, set segments 32, overlap 2 and dim_t 201 with num_overlap 2 both at the bottom of yaml config in \models\MDX_Net_Models\model_data\mdx_c_configs
(dim_t 301 and overlap 3 also works, although not on all models [e.g. not for beta 3, but inst v1] and seems to be less muddy and fewer clicks appear).
dim_t 201 is not optimal setting and might lead to more occasional quiet residues, clicks or sudden volume changes (like chunk was changing every 2 seconds), although there’s no stem misalignment issue with these settings (they work both for Mel and BS Roformers). dim_t 301 with lighter models seems to be a bare minimum to avoid the majority of audible artefacts (after patch #3 dim_t 256 is allowed - “make sure you check the "Segment Default" in MDXNET23 Only Options for it to take effect”).
Using the settings above on patch #2, with GPU acceleration it will take 39m 28s for 3:28 song using 1296 model on RX 470 4GB and 18 minutes for Kim Mel-Roformer and 3:01 song.
Using HQ_4 is much faster than realtime using default settings, but even longer than accelerated Roformer, when on CPU only using old Core 2 Quad @3.6 DDR2 800MHz.
On Mac M1 using the patch above, it takes 9 minutes to process a 3-minute song using BS-Roformer (dim_t 1101, batch size 2, overlap 8) with “constant throttling”. Click
And below 4 minutes for Kim Mel-Roformer (overlap 1, dim 801). Click
- Settings working for 6GB AMD GPUs: dim_t 601 or 701 at the bottom of the yaml file and overlap 6 or 7 in GUI.
- Like I mentioned, overlap 8 can be good enough too when dim_t=801 is set (the fastest setting before SDR getting drastically reduced), at least in other cases you shouldn’t exceed 6, while 2 should provide the best quality in most cases.
- 1602 (or rather 1601) dim_t might lead to less wateriness, but turns out in cost of a bit more of vocal residues.
- “In theory, max overlap value [for Roformer separations without mentioned issues in UVR] can be known with formula:
(dim_t - 1) / 100 = Max_overlap_value
if dim_t = 801:
(801 - 1) / 100 = 8
if_dim_t = 1101:
(1101 -1) / 100 = 10 [jarredou wrote 10 here, but it’s actually 11]
Above that max value, some parts of the input will not be processed.
The lower the overlap value is, the more overlap is used, so better SDR.
- Some Rofos models still have wrong config by default, with dim_t=256, so max overlap value for that is 2. That's why I've advised to stick to overlap=2” - jarredou
So in times before dim_t was known how to be correctly set, so now overlaps can be even set to 8 now when dim_t=801 is set]).”
The same thing applies for both BS and Mel Roformers in UVR.
“audio.dim_t value is not used with roformers in ZFTurbo script, it uses audio.chunk_size and then it's parameters in the model part of config.”
- Using older Roformer beta patches for Mac M1 doesn’t allow you to choose the Roformer parameter to check for custom Roformer models and only config name can be chosen, but no confirm button is available. So the error “File "libv5/tfctdfv3.py", line 152, in __init” appears.
> Place the corresponding json file with your model from this repo into: models\MDX_Net_Models\model_data beforehand, to fix the issue.
In some cases, you may still get the same error anyway and to get rid of it, you need to edit manually model_data.json adding desired model line at the end like your custom model was downloaded from download center. On example of unwa’s beta 3:
},
"d43f93520976f1dab1e7e20f3c540825":{
"config_yaml": "config_melbandroformer_big.yaml",
"is_roformer": true
}
Additionally, you need the model at the end of model_data_mapper.json:
"model melband_roformer_big_beta3.ckpt": "config_melbandroformer_big"
}
Now copy the hash-named json file (d43f93520976f1dab1e7e20f3c540825.json for beta 3) to model_data folder.
All the three modified files for beta 3 and other models here.
If you have problems generating hash on first launch of the model and your model is not uploaded in the repo above or json is not generated then use Windows installation in VM, or ask some PC user for the config. Potentially reading Hash decoding can be helpful.
But maybe your hashed config name will be generated correctly already after you imported the model into UVR (although no confirmation button might prevent it), and now it will be enough to just place the following line like in the jsons presented above: "is_roformer": true” (so after “,” in the yaml line above).
- More in-depth - Settings per model SDR vs Time elapsed -||- (incl. dim_t and overlap evaluation for Roformers) - click or here | conclusion - made before patch #3
Model characteristics
(the list might be getting outdated, read models list at the top)
Note: E.g. unwa’s duality models v1/2 and inst v1/2/v1e are now added to UVR Beta Roformer Download Center (so you don’t have to mess with models and configs manually)
- viperx 1053 model separates drums and bass in one stem, and it's very good at it
(although now it might be better to use Mel-Roformer drums on x-minus.pro/uvronline)
“Target is drums and bass, and "other" is the rest. Despite that, it says vocals”
- Unwa released a new Inst v1e model | Colab | MSST-GUI (“The model [yaml] configuration is the same as v1”)
“The "e" stands for emphasis, indicating that this is a model that emphasizes fullness.”
- unwa inst v2 - it gets muddier than v1 at times, but it has less of noise
- unwa inst v1 - focused on instrumental stem:
model | Colab | MSST-GUI | phase fixer
"much less muddy (..) but carries the exact same UVR noise from the [MDX-Net v2] models"
But it's a different type of noise, so aufr33 denoiser won't work on it.
“you can "remove" [the] noise with uvr denoise aggr -10 or 0” although with -10 it will make it sound more muddy like Kim model and synths and bass are sometimes removed with the denoiser (~becruily). Mel-Roformer denoise might be better for it.
becruily released a Python script fixing the noise issue (execute “pip install librosa” in case of module not found error) - it sound similar to the method used for premium user on x-minus.
- unwa beta 4 Mel-Roformer (fine tune of Kim’s voc/inst model ):
https://huggingface.co/pcunwa/Mel-Band-Roformer-big/tree/main | Colab
Be aware that the yaml config has changed, and you need to download the new beta4 yaml.
“Metrics on my test dataset have improved over beta3, but are probably not accurate due to the small test dataset. (...) The high frequencies of vocals are now extracted more aggressively. However, leakage may have increased.” - unwa
“one of the best at isolating most vocals with very little vocal bleed and still doesn't sound muddy” “gives fuller vocals”. Can be a better choice on its own than some ensembles.
- unwa duality model - focused on both stems, and instrumental is similarly muddy like in beta 4
- Kim Mel-Band Roformer vocal model
It’s less muddy than 1296/1297.
(original repo - CML faster on CUDA than in UVR | model | config - place the model file to models\MDX_Net_Models and .yaml config to model_data\mdx_c_configs subfolder and “when it will ask you for the unrecognised model when you run it for the first time, you'll get some box that you'll need to tick "roformer model" and choose it's yaml” (Mac issue explained in the section above).
(simple Colab/CML inference/x-minus/MVSEP/jarredou Colab too now)
- unwa BS-Roformer finetuned a.k.a. large (further trained viperx 1297 model) download
More muddy than Kim above, a bit less of vocal residues, a bit more artificial sound.
- Mel-RoFormer Karaoke / Lead vocal isolation model files released by Aufr33 and viperx (download)
Older models in Download Center
- older viperx’ 1297 model tend to be a bit better for instrumentals, and 1296 for vocals (both more muddy than Kim and Unwa models, but “still pretty good for voice cleaning” and dealing with noise) - BS-Large model by Unwa is a fine-tune of that model.
- 1143 model is the first Mel-Roformer trained by viperx before Kim introduced changes to the config, which fixed the problem of lower SDR vs models trained on BS-Roformer. Use Kim Mel-Roformer instead
Both models struggle with saxophone and e.g. some Arabic guitars. It can still depend on a song whether these are better than even the second oldest Roformer than on MVSEP (from before viperx model got fine-tuned version). They tend to have more problems with recognizing instruments. Other than that, they're very good for vocals (although Mel-Roformer by Kim on x-minus tends to be better).
Muddy instrumentals when not ensembled with other archs.
Be aware that names of these models on UVR refer to SDR measurements of vocals conducted on private viperx dataset, not even older Synthetic dataset, instead of on multisong dataset on MVSEP, hence the numbers are higher than in the multisong chart on MVSEP.
___
Older news follow
___
- The viperx model was also added on MVSEP
- New ensembles with higher SDR were added on MVSEP
- BS-Roformer model trained by viperx was added on x-minus (it's different from the v2 model on MVSEP, and has higher SDR, it's the “1.0” one). If it's better vs V2 might depend on a song.
It struggles with saxophone and e.g. some Arabic guitars.
- (x-minus - aufr33) “I have just completed training a new UVR De-noise model. Unlike the previous version, it is less aggressive and does not remove SFX.
It was trained on a modified dataset. I reduced the noise level and made it more uniform, removed footsteps, crowd, cars and so on from the noise stems. On the contrary, the crowd is now a useful / dry signal. (...) The new model is designed mainly to remove hiss, such as preamp noise.”
For vocals that have pops or clipping crackles or other audio irregularities, use the old denoise model.
- Dango.ai updated their model, also giving some kind of demudder to the instrumentals, enhancing their results. Results might be better than MDX23C and BS-Roformer v2. Still, it’s pretty pricey (8$ for 10 separations). 5x 30 seconds fragments per IP can be obtained for free, and usually it doesn’t reset. “It’s $8 for 10 tracks x 6 minutes, all aggressiveness modes included (but vocal and inst models are separate). The entire multisong dataset for proper SDR check would cost around $133.” becruily
- Be aware that queues on https://doubledouble.top/ are much shorter for Deezer than Qobuz links. If there’s no 24 bit versions for your music, use Deezer instead.
[outdated; currently there’s no longer any MQA files on Tidal] Also, avoid Tidal and 16 bit FLACs from “Max” quality, which is slightly lossy MQA. Use 24 bit MQA from Tidal only when there’s no 24 bit on Qobuz. Most older albums under 2020 are 16 bit MQA instead of 24 bit MQA on Tidal, and are lossy compared to Deezer and Qobuz which doesn’t use MQA (so doubledouble doesn’t convert MQA to FLAC like on Tidal). MQA is only “slightly” lossy, because it affects frequencies mainly from 18kHz and up, and not greatly.
- Members of neighboring AI Hub server made a fork of KaraFan Colab updated with the new HQ_4 and InstVoc HQ2 models. It has slow separation fix applied. Click
- HQ_4 and Crowd models added to HV Colab temp fork before merge with main GH repo
- (MVSEP) “We have added longer filenames disabling option to mvsep, you can access it from Profile page
20240312034817-b3f2ef51cb-ballin_bs_roformer_v2_vocals_[mvsep.com].wav -> ballin_bs_roformer_v2_vocals.wav
Due to browser caching, you might want to hard refresh the page if you have downloaded onc”
- The ensembles for 2 and 5 stems on MVSEP have been updated with bigger SDR bag of models containing now new BS-Roformer v2 (with MDX23C, VitLarge23, and for multistem, the old demucsht_ft, deumcs_ht, demucs_6s and demucs_mmi models)
- All the Discord direct links leading to images in this document have expired. I already reuploaded some more important stuff. Please ping me on Discord if you need access to some specific image. Provide page and expired link.
- https://free-mp3-download.net has been shut down. Check out alternatives here.
New Apple Music ALAC/Atmos downloader added, but its installation is a bit twisted and subscription is required. Murglar added.
- MDX-Net HQ_4 model (SDR 15.86) released for UVR 5 GUI! Go to Models list>Download center>MDX-Net and pick HQ_4 for download. It is an improved and faster than HQ_3, trained for epoch 1149 (only in rare cases there’s more vocal bleeding, more often instrumental bleeding in vocals, but the model is made with instrumentals in mind.
Along with it, also UVR-MDX-NET Crowd HQ 1 has been added in download center.
- HQ_4 model added to the Colab:
https://colab.research.google.com/github/kae0-0/Colab-for-MDX_B/blob/main/MDX_Colab.ipynb
- New BS-Roformer v2 model released on MVSEP. It’s more aggressive model than above.
- Fixed KaraFan Colab with the fix for slow non-MDX23 models. You'll no longer stack on voc_ft using any other preset than 1, but be aware that it will take 8 minutes more to initialize. (same fix as suggested before, but w/o console, as it wasn't defined, and faster ort nightly fix doesn't work here).
Turns out, there has been an official non-nightly package released, and it works with KaraFan correctly (no need to wait 8 minutes any longer):
!python -m pip -q install onnxruntime-gpu --extra-index-url https://aiinfra.pkgs.visualstudio.com/PublicPackages/_packaging/onnxruntime-cuda-12/pypi/simple/
- (x-minus.pro) “Since Boosty is temporarily not accepting PayPal and generally working sucks, I made the decision to go back to Patreon. Please be aware that automatic charges will resume on March 22, 2024. If you have Boosty working correctly and do not intend to use Patreon, please cancel your Patreon subscription to avoid being charged.
If you wish to switch from Boosty to Patreon, please wait for further instructions in March.” Aufr33
- If you suffer from bleeding in other stem of 4 stems Ripple, beside decreasing volume by e.g. 3/4dB also “when u throw the 'other stem' back into ripple 4 track split a second time, it works pretty well [to cancel the bleeding]” if it's still not enough, put other stem through Bandlab Splitter.
- If you suffer from vocal residues using Ensemble 4 models on MVSEP.com, decrease volume of input file by -8dB “now it's silent. No more residue” usually 3 or 4dB was doing the trick for Ripple, but here it’s different. Might depend on a song too.
- Image Line “released an update for FL Studio, and they improved the stem separation and it's better, but it has quite a bit of bleeding still, but it also seems they may have improved the vocal clarity”
- (probably fixed in new HV MDX) Our newly fixed VR and newer HV MDX Colabs started to have issues with very slow initialization for some people (even 18 minutes/+ instead of normally 3). It’s probably due to very slow download of some dependencies. Possible solutions: use other Google account, use VPN, make another Google account (maybe using Polish VPN). Let us know if it happens only for some specific dependency or all of them. You can try to uncomment the ORT nightly line in mounting cell (add # before), as it triggers more dependencies to be installed, which can be slow in that case. The downside is - there won't be GPU acceleration, and one song will be processed in 6-8 minutes instead of ~20 seconds.
- New paid drum separation service:
https://remuse.online/download (might not work anymore)
It uses free drumsep model (same model hash: 9C18131DA7368E3A76EF4A632CD11551)
- MDX Colab seem to not work due to Numpy issues. I already fixed them in Similarity Colab, and hopefully reimplement the fixes elsewhere soon. VR Colab fixed too.
Tech details about introduced changes described below Similary Extractor section.
- Music AI surfaced. Paid - $25 per month or pay as you go (pricing chart). No free trial. Good selection of models and interesting module stacking feature. To upload files instead of using URLs “you make the workflow, and you start a job from the main page using that custom workflow” [~ D I O ~].
Allegedly it’s made by Moises team, but the results seem to be better than those on Moises.
“Bass was a fair bit better than Demucs HT, Drums about the same. Guitars were very good though. Vocal was almost the same as my cleaned up work. (...) I'd say a little clearer than mvsep 4 ensemble. It seems to get the instrument bleed out quite well, (...) An engineer I've worked with demixed to almost the same results, it took me a few hours and achieve it [in] 39 seconds” Sam Hocking
- “I just got an email from Myxt saying they're going to limit stem creation to 1 track per month. For creator plan users (the $8 a month one) and 2 per month for the highest plan.
So I may assume with that logic, they're gonna take it away for free users?”
- (probably fixed) For all jarredou's MDX23 v. 2.3 Colab fork users:
“Components of VitLarge arch are hosted on Huggingface... when their maintenance will be finished it will work again. I can't do anything about it in the meantime.”
2.2 and 2.1 and MVSEP.com 4-8 models ensemble (premium users) should work fine.
- Ripple now has fade in and clicking issues fixed. Also, there's less bleeding in the other stem (but Bas Curtiz’ trick for -3dB/-4dB input volume decreasing can be still necessary).
“Ripple’s lossless outputs are weird, some stems like the drums are semi full band (kicks go full band, snares not etc) and the “other” stem looks like fake full band”. These fixes are applied also for old versions of the app.
Also, the lossless option fixes to some extend the offset issue so it's more similar to input now, but not identical (lossless option might require updating). Also no more abrupt endings
Ripple = better than CapCut as of now (and fullband).
plus Ripple fixed the click/artifacts using cross-fade technique between the chunks.
- ViperX currently doesn't plan to release his BS-Roformer model
- New “uvr de-crowd (beta)” model added on x-minus. Seems to provide better results than the MVSEP model. Also, an MDX arch model version is planned for training.
“At minimum aggressiveness value, a second model is now used, which removes less crowd but preserves other sounds/instruments better.”
- Ripple seems to have a lossless export option now. “First make sure the app is updated then click the folder then click the magnet icon then export and change it to lossless”
- Seems like CapCut now has added separation inside Android Capcut app in unlocked Pro version
https://play.google.com/store/apps/details?id=com.lemon.lvoverseas (made by ByteDance)
Seems like there is no other Pro variant for this app.
At least unlocked version on apklite.me have a link to regular version, so it doesn't seem to be Pro app behind any regional block. But -
"Indian users - Use VPN for Pro" as they say, so similar situation like we had on PC Capcut before. Can't guarantee that unlocked version on apklite.me is clean. I've never downloaded anything from there.
- Mega, GDrive and direct link support for input files added on MVSep. If you want to apply MVSep algorithm to result of other algorithm, you can use "Direct link" upload and point https link on separated audio-file on MVSep.
- If you have an issue with Demucs module not found in e.g. MDX23 v.2.3 Colab (now fixed there and also in VR Colab), here's a solution:
“In the installation code, I added `!pip install samplerate==0.1.0` right before the `!pip install -r requirements.txt &> /dev/null` and I managed to get all the dependencies from the requirements.txt installed properly.” (derichtech15)
- If you repost your images or files from Discord elsewhere while cutting link after "ex=" for all new posted files, it will make your files expire pretty soon (17.02.24). If you leave the full link with "ex=" and so on, it won't expire so fast, but who knows if not later.
So far, all the old Discord images shared elsewhere with "ex=" cut, work (also in incognito without Discord logged in), but it's not certain that it will be that way forever.
Discord announced in the end of 2023, that they'll update their mechanisms of sharing links, so they'll expire after some time when they're shared, to avoid some security vulnerabilities allowing scams. Or they just want to offload the servers.
- OpenVINO™ AI Plugins for Audacity 3.4.2 64-bit introduced.
4 stems separation, noise suppression, Music Style Remix - uses Stable Diffusion to alter a mono or stereo track using a text prompt, Music Generation - uses Stable Diffusion to generate snippets of music from a text prompt, Whisper Transcription - uses whisper.cpp to generate a label track containing the transcription or translation for a given selection of spoken audio or vocals.
Not bad results. They use Demucs.
- For people with low VRAM GPUs (e.g. 4GB or less), you can test out Replay app, which provides voc_ft model and tends to crash less than UVR. Sadly, the choice of models is much smaller, but it has some de-reverb solution. Screenshot
- Latest MVSep changes:
1) All ensembles now have option to output intermediate waveforms from independent algorithms + additional max_mag, min_mag.
2) Ensemble All-In now includes DrumSep results extracted from Drum stem.
- resemble-enhance (GH) model added on x-minus in denoise mode. It can work better than the latest denoise model on x-minus. It is intended only for vocals. For music use UVR De-noise model on x-minus.
- (fixed in kae, 2.1, 2.2 [and KaraFan irc] Colabs) All Colabs using MDX-Net models are currently very slow. GPU acceleration is broken and separations now only work on CPU with onnxruntime warnings.
To work around the issue, go to Tools>Command palette>Use fallback runtime version (while it's still available).
Downgrading CUDA to 11.8 version fixes the issue too, but it takes 9 minutes in order to install that dependency, so it’s faster to use fallback runtime till it’s still available. After that period, just execute this line after initialisation cell:
console('apt-get install cuda-11-8') and GPU acceleration will start to work as usual.
>“Better fix [than CUDA 11.8] until final version is released, using that onnxruntime-gpu nightly build for cuda12:
!python -m pip install ort-nightly-gpu --index-url=https://aiinfra.pkgs.visualstudio.com/PublicPackages/_packaging/ort-cuda-12-nig
htly/pypi/simple/
(no need to install cuda 11.8)” jarredou
In case of credential issues you can try out this package instead:
!python -m pip -q install onnxruntime-gpu --extra-index-url https://aiinfra.pkgs.visualstudio.com/PublicPackages/_packaging/onnxruntime-cuda-12/pypi/simple/
- LarsNet model was added on MVSep. It's used to separate drums tracks into 5 stems: kick, snare, cymbals, toms, hihat. Source: https://github.com/polimi-ispl/larsnet
It’s worse than Drumsep as it uses Spleeter-like architecture, but “at least they have an extra output, so they separate hihats and cymbals.”. Colab
“Baseline models don't seem better quality than drumsep, but the provided checkpoints are trained with oly 22 epochs, it doesn't seem much. (and STEMGMD dataset was limited by the only 10 drumkits), so it could probably be better with better dataset & training”
“ it separates the toms so much better [than Drumsep]”
Similar situation as with Drumsep - you should provide drums separated from e.g. Demucs model.
- Captain FLAM from KaraFan asks for some help due to some recent repercussions.
You can support him on https://ko-fi.com/captain_flam
- To preserve instruments which are counted as vocals by other MDXv2 models in KaraFan, use these preset 5 modified settings (dca100fb8).
- Added more remarks from testing these settings against sax preset and others.
- drumsep added on MVSEP!
(separation of drums from e.g. Demucs 4 stem or “Ensemble 8 models”/+)
- New Bandid Plus model added on MVSEP
“I trained BandIt for vocals. But it's too far away from MDX23C” -ZFTurbo
“I loved this bandit plus model!! It has great potential.”
- UVR De-noise model by FoxJoy added on x-minus. It’s helpful for light noise, e.g. vinyl. (de-reverb and de-echo are up already)
New MDX de-noise model is in the works and beta model was also added!
“the instruments in the background are preserved much better than the FoxJoy model”
It works for hiss, interference, crackle, rustles and soft footsteps, technical noise.
- New hifi-gan-bwe Colab fork made by jarredou:
https://colab.research.google.com/github/jarredou/hifi-gan-bwe/blob/main/HIFIGAN_BWE.ipynb
- New AI speech enhancer - https://www.resemble.ai/introducing-resemble-enhance
- Reason 12.5 (a DAW) was released with VST3 plugin support
- jazzpear94 “I made a new model with a modified version of my SFX and Music dataset with the addition of other/ambient sound and speech. It's a multistem model and should even work in UVR GUI as it is MDX23C.
Note: You may want to rename the config to .yaml as UVR doesn't read .yml and I didn't notice till after sending. Renaming it fixes that, however”
“You put config in models\mdx_net_models\model_data\mdx_c_configs. Then when you use it in UVR it'll ask you for parameters, so you locate the newly placed config file.”
“Keep in mind that the cinematic model focus is mainly on sfx vs instruments
voice stems are supplemental. Usually I remove voices first”
- https://github.com/karnwatcharasupat/bandit
Better SDR for Cinematic Audio Source Separation (dialogue, effect, music) than Demucs 4 DNR model on MVSEP (mean SDR 10.16>11.47)
- "Demucs+CC_Stereo_to_5.1" - it's a script where you can convert Stereo 2.0 to 5.1 surround sound. Full discussion about script. They use MVSep to get steams and after use script on them.
- Colab by jazzpear96 for using ZFTurbo's MSS training script. “I will add inference later on, but for now you can only do the training process with this!”
- New djay Pro 5.0 has “very good realtime stems with low CPU” Allegedly “faster and better than Demucs, similar” although “They are not realtime, they are buffered and cached.” it uses AudioShake. It can be better for instrumentals than UVR at times.
- AudiosourceRE Demix Pro new version has lead/backing vocals separation
- New crowd MSX23C model added on MVSEP (applause, clapping, whistling, noise) (and got updated by the time 5.57 -> 6.06; added hollywood laughts, old models also available)
- VitLarge23 model on MVSEP got updated (9.78>9.90 for instrumentals)
- MelBand RoFormer (9.07 for vocals) model added on MVSEP for testing purposes
“The model is really good at removing the hi-hat leftovers. These e.g. in the Jarredou colab sometimes when you can hear the hi-hats from the acapella. And Melband roformer can almost remove all the hi-hat leftovers from the acapella.”
“are the stems not inverted result? for me it sounds like there is insane instrument loss in the instrumental stem and vocals loss in the vocal stem, yet there is no vocal bleed in instrumental stem and vice versa” “I also think that the vocals are surprisingly clean considering the instrumentals sound quite suppressed but also clean”
- Goyo Beta plugin for dereverb stopped working on December 2nd (as it required internet connection and silent authorization on every initialization). They transitioned to paid Supertone Clear. They send BETA29 coupon over emails (with it, it’s $29).
- New MVSep-MDX23 Colab Fork v2.3 by jarredou published under new Colab link here
Now it has Vitlarge23 model (previously used exclusively on MVSEP) instead of HQ3-Instr, also improved BigShifts and MDXv2 processing.
Doesn't seem to be better than RipX which is better in preserving some instruments, and also removes vocals completely
- Check out new Karaoke recommendations (dca100fb8)
- Dango.ai finally received English web interface translation
- New SFX model based on Mel roformer was released by jazzpear94. More info
- User friendly Colab made by jarredou and forked by jazzpear94 with new feature. In case of some problems, use WAV file.
- Seems like Ripple got updated, "it sounds a lot better and less muddied" doesn’t seem to give better results for all songs, though. Might be similar case with Capcut too.
- Hit 'n' Mix RipX DAW Pro 7 released. For GPU acceleration, min. requirement is 8GB VRAM and NVIDIA 10XX card or newer (mentioned by the official document are: 1070, 1080, 2070, 2080, 3070, 3080, 3090, 40XX, so with min. 8GB VRAM). Additionally, for GPU acceleration to work, exactly “Nvidia CUDA Toolkit v.11.0” is necessary. Occasionally, during transition from some older versions, separation quality of harmonies can increase. Separation time with GPU acceleration can decrease from even 40 minutes on CPU to 2 minutes on decent GPU.
- UVR BVE v2 beta has been updated on x-minus
“It now performs better on songs with 2 people singing the lead
No longer separates the second lead along with it”
-dca100fb8 found out new settings for KaraFan which give good results for some difficult songs (e.g. Juice WRLD) for both instrumental and acapella. It’s now added as preset 5.
Debug mode and God mode can be disabled, as it's like that by default.
"It's like an improved version of Max Spec ensemble algorithm [from UVR]"
Processing time for 6:16 track on medium setting is 22 minutes.
- New MDX23C model added exclusively on MVSEP:
vocals SDR 10.17 -> 10.36
instrum SDR 16.48 -> 16.66
Also ensemble 4 got updated by new model (10.32>10.44 for vocals)
- For some people using mitmproxy scripts for Capcut (but not everyone), they “changed their security to reject all incoming packet which was run through mitmproxy. I saw the mitmproxy log said the certificate for TLS not allowed to connect to their site to get their API. And there are some errors on mitmproxy such as events.py or bla bla bla... and capcut always warning unstable network, then processing stop to 60% without finish.” ~hendry.setiadi
“At 60% it looks like the progress isn't going up, but give it idk, 1 min tops, and it splits fine.” - Bas
-ZFTurbo published his training code:
https://github.com/ZFTurbo/Music-Source-Separation-Training
"It gives the ability to train 5 types of models: mdx23c, htdemucs, vitlarge23, bs_roformer and mel_band_roformer.
I also put some weights there to not start training from the beginning."
It contains checkpoint of e.g. 1648 (1017 for vocals) MDX23C model to train it further.
Be aware that the older bs_roformer implementation is very slow to train IRC.
Vitlarge23 “is running 2 times faster than MDX models, it's not the best quality available, but it's the fastest inference”
“change the batch size in config tho
I think zfturbo sets the default config suited for a single a6000 (48gb)
and chunksize”
-"A small update to the backing vocals extractor [on X-Minus]
Now you can more accurately specify the panning of the lead vocal." ~Aufr33 Screen
- IntroC created a script for mitmproxy for Capcut allowing fullband output, by slowing down the track. Video
- Jazzpear created new VR SFX model. Sometimes it’s better, sometimes it’s worse than Forte’s model. Download
For UVR 5.x GUI, use these parameters (irc same as Forte):
User input stem name: SFX
Do NOT check inverse stem!
1band sr44100 hl 1024
- Now KaraFan should work locally on 4GB GTX GPUs (e.g. laptop 1060), on presets 2 or 3, and with chunk 500K, speed can be slowest. Download on GitHub the Code > ZIP
-Bas Curtiz' new video on how to install and use Capcut for separation incl. exporting:
https://www.youtube.com/watch?v=ppfyl91bJIw
and saving directly as FLAC, although the core source of FLAC is still AAC in this case:
https://www.youtube.com/watch?v=gEQFzj6-5pk
"It's a bit of a hassle to set it up, but do realize:
- This is the only way (besides Ripple on iOS) to run ByteDance's model (best based on SDR).
- Only the Chinese version has these VIP features; now u will have it in English
- Exporting is a paid feature (normally); now u get it for free
The instructions displayed in the video are also in the YouTube description."
Capcut normalizes the input, so you cannot use Bas’ trick to decrease volume by -3dB like in Ripple to workaround the issue of bleeding (unless you trick out the CapCut, possibly by adding some loud sound in the song with decreased volume, something like presented here).
- (fixed) KaraFan Colab will be fixed on 27th at morning.
- There’s a workaround for people not able to split using Capcut. The app discriminate based on country (poor/rich) and paywalls Pro option.
The video demonstration for below
0. Go offline.
1. Install the Chinese version from capcut.cn
2. Use these files copied over your current Chinese installation, and don’t use English patch.
3. Open CapCut, go online after closing welcome screen, happy converting!
4. Before you close the app, go offline again (or the separation option will be gone later).
Before reopening the app, go offline again, open the app, close welcome screen, go online, separate, go offline, close. If you happen to missed that step, you need to start from the beginning of the instruction.
(replacing SettingsSDK folder no longer works after transition from 4.6 to 4.7, it freezes the app)
FYI - the app doesn’t separate files locally.
- Bas Curtiz found out that decreasing volume of mixtures for Ripple by -3dB eliminates problems with vocal residues in instrumentals in Ripple. Video.
This is the most balanced value, which still doesn't take too many details out of the song due to volume attenuation.
Other good values purely SDR-wise are -20dB>-8dB>-30dB>-6dB>-4dB> /wo vol. decr.
The method might be potentially beneficial for other models and probably work best for the loudest tracks with brickwalled waveforms.