Google Doc

- (no longer necessary) Fork of UVR GUI and How to install - support for AMD and Intel GPUs appeared (works only for VR and MDX architectures), Besides W11, also W10 confirmed working, MDX achieves speeds of i5-4460s using 6700 XT, while for VR, speeds are v. fast and comparable to CUDA, so CPU processing might be slower in VR, but for MDX you might want to stick with the official UVR5 GUI.

Page 2 of 28 · Edit this page in Google Docs ↗

- Batch mode seems to fix problems with vocal popping using low chunks values in MDX models, and also enhance separation quality while eliminating lots of out of memory issues. It decreases SDR very slightly for single models, and increases SDR in ensemble.

- (outdated) New beta MDX model “Inst_full_292” without 14.7kHz cutoff released (performs better than Demucs 4 ft). If the model didn’t appear on your list in UVR 5 GUI, make sure you’ve redeemed your code https://www.buymeacoffee.com/uvr5/vip-model-download-instructions

Or use Colab.

Newer epochs available for paid users of https://x-minus.pro/ai?hp&test-mdx

- To use Colabs in mobile browsers, you probably no longer need to switch your browser to PC Mode first.

News section continues in older news/update logs

General reading advice

- If you found this document elsewhere (e.g. as PDF), here is always up-to-date version of the doc

- If you have anything to add to this doc, ping me @deton24 on our Discord server from the footer, but rather refrain from PMing directly if not necessary. Every time you request writing privileges via GDoc, God kills a cat. Don't click “ask for privileges”!

- You can use the (the outdated) Table of content section, but better go to Options and show “Document outline” to see an up-to-date clickable table of content. If you don't have Google Docs installed, and you opened the doc in a mobile browser and no Table of content option appear, use the old Table of content or go to options of the mobile browser and run the site in PC mode to show document outline (but it's better to have Google Docs installed on your phone instead, as it’s more convenient in use).

- When you click from elsewhere on a GDoc link with “heading” string in the URL redirecting to a specific section, it should show the first page, and redirect after a few seconds

- Once you load the document and reach the document outline, redirection to any entry might not work at the very beginning till all the pages will be loaded. It takes some time. Tap a few times.

- Sometimes you cannot scroll down the list of headings on the phone in PC mode. Then you need to tap on the scroll bar in the very left, but it might suddenly look buggy all highlighted, but working nevertheless.

- Downloaded .docx will have a similar document outline as in the GDoc (but more messy - with all headers used in the GDoc). If you have an error on attempt of opening the .docx file on Windows, go to RBM>Properties and check Unlock below Attributes.

- Even on a powerful phone with lots of RAM, GDoc app can occasionally crash, esp. while clicking on a specific section before the doc is fully loaded. Even deleting the app and reopening won’t help for occasional crashes. The document became too big to handle properly by Google service (both mobile app and online).

- Be aware that the document can hang for a while on the attempt of accessing a specific section of the document - it doesn't happen often on a PC browser - PC is the most stable form of reading the doc online. At least on a decent PC (so not C2Q, but even a decade old i7 is usually fine). But it can be stable on Android phone too (e.g. Snapdragon 700 series instead of old 400 series). Google’s app support for old 32-bit ROMs in e.g. Android 9 and older is terrible.

- Search and navigating through the document outline works faster when you download the doc as .pdf or .docx, but in the latter you’ll have access to the document outline on the left like in GDoc (if not, press CTRL+F>Headings, or check View>Navigation window).

- When visiting the online version of this doc, you can paste whole phrase when searching instead of single letters to avoid severe stuttering during using search function online.

- Use the search bar in Google Documents, not the one from the browser - the browser’s search won’t find everything unless all the pages were shown before - the doc is huge.

- Sometimes if you search for a specific keyword in the mobile app and the result doesn't show up, you need to go to the document outline, and open its last section and search again (so the whole document will be loaded first, otherwise you won't get all the search results in some cases). But it might happen mostly if you use the wrong search function.

- Make sure you've joined our Discord server to open some of the Discord links attached below (those without any file extension at the end).

But I think now it might be no longer necessary.

- Download links from Discord with file extensions at the end no longer work, but I reuploaded most of the important links already. If you need to download from previously shared Discord link anyway:
1) Join our Discord server via invitation at the top of the document 2) Delete file name from the link 3) Open our Discord server in the browser 4) Leave everything in the link before the first slash and delete the rest (so channels\xxxx\ 5) paste two identifiers divided by slashes afterwards, but without file name (so channels\xxxx\xxxx\xxxx - where the last two are taken from inactive file link) *) If you paste offline link in any channel on the source server, the link will work again

- If you have a crash on opening the doc in the app, e.g. on Android - reset the app cache and data. Keep the app updates or find some old version (e.g. even from period when your phone was released or uninstall all updates if it's GDoc is your stock app

- If it loads 4 minutes/infinitely in the doc app, update your Google Docs app and reset the app cache/data, e.g. if you started to have crashes after the app update.

- You can share a specific section of this document by opening it on PC or on a mobile browser set in PC mode by clicking on one of the sections in the document outline (or hyperlinks leading to specific sections). Now it will add a reference to the section in the link in your address bar, which you can copy and paste, so opening this link will redirect someone straight to the section after opening the link (in some specific cases, some people won’t be redirected, but in fact, you only need to wait a few seconds after the first page of the doc has been shown, and then the proper redirect starts).

Some headers not referred to in the outline are also set in a way that when you click them, the address bar will change with a link leading to that specific section. Not all headers present in the doc are shown in the outline to preserve better readability.

 - In the GDoc app sometimes you need to tap “wait” a few times when the app freezes. Afterwards, searching will start working all the time (at least till the next time). The doc is huge and the GDoc app on at least low-end Androids is cursed (desktop version on PC behaves the most stable, as long as decent phones). You've been warned.

- If you feel overwhelmed by the doc size, theoretically you can load the doc into Google Gemini or Google NotebookLM and ask questions from there, but I encourage befriending with the document outline and the content of the interesting section yourself - asking the AI chat for the best models leads to hallucinating of the LLM and providing list of outdated separation models from this doc. Also, they all miserably fail with generating model and config links from even cut fragments of the GDoc. They also cannot edit the document directly (unless you paste the text, but it will rather delete all the formatting and hyperlinks which are essential to the task), maybe Office 365 with paid subscription is capable of editing docs with AI directly already.

- Descriptions on the list of models are usually shortened compared to the news section information added when the model was released. You can use search with the model name for more possible descriptions.

- Published audio demos of models pasted from MVSEP get online after a while, so most links to audio files from there will be offline after a week or more.

- Without GDoc app installed, or when in not PC mode of the browser, if you you open links to this document ending with e.g. “#heading=h.hk34hc4d1ah7” or similar, you won't be redirected to a specific section of this document referred to in the heading. Redirections from such links don't work in the mobile version of the GDoc site.

To be redirected after a moment to proper section from outside links with “heading” in the URL, you should open these links with GDoc app installed, or on PC browser, or mobile browser with PC Mode turned on (in Chrome that option appears when you open a page already).

- (fixed in April) In February 2026 it started to happen that Chrome on Windows was closing after opening the document. It helped to reopen it a few times, maybe along with uninstalling offline GDoc extension and maybe after going to chrome://settings/content/all?sort=data-stored > Google > docs.google.com (I just did all of these). The issue was recurring even when you just disabled the extension (it was able to re-enable itself).

- If you click on any hotlink redirecting to a specific part of this document from the mobile version of GDoc in the browser, you won't be able to show options to display the document outline after switching your browser into PC mode - it will remain in the mobile layout. It's because redirections in mobile versions have their own linking scheme adding to the site address - you need to delete its ending or reopen the doc.

- If loading bar stutters suddenly very much on opening the GDoc in the app, most likely it will crash after a longer time, so then just reopen the doc - once your OS freezes/closes the app due to inactivity, task switching, or esp. the phone I sleep mode, reopening the app will fail in many cases (but sometimes it will redirect you to the last place you read)

- Besides me, jarredou (Discord: rigo2), dca100fb8 (both since 23 May 25) and isling (since 12 Feb 26), currently no one else has writing privileges to this document, although they were reluctant to be active editors, and were granted the privileges as the last resort for possible cases of my longer absence in the future.


             (I’m trying to keep the following list always updated with the Last updates/news section at the top)

Everyone asks which service and/or model is the best for instrumentals, vocals or stems. The answer is - we have listed a few models and services which behave the best in most cases, but the truth is - the result also strictly depends on the genre, specific song, and how aggressive and heavily processed vocals it has. Also, how much distortion instruments have, style of mixing, etc. Sometimes one specific album gets the best results with one specific tool/AI/model, but there might be some exceptions for specific songs, so just feel free to experiment with each, to get the best result possible using various models, ensembles and services/AIs from those listed. SDR on MVSEP in the multisong dataset doesn't always reflect bleeding well. That’s why we introduced bleedless and fullness metrics for evaluation of the models as well. Although they’re not perfect either. Pay attention to also other metrics featured in MVSEP evaluation results on the site:
“-l1_freq = bleedless (higher is cleaner)

-aura_mrstft = fullness (higher is fuller)” - becruily, “more perceptually relevant than SDR for musical content imo” - gilliaan
You’ll read more about it in this section.

“Some people don't realize that if you want something to sound as clean as possible, you'll have to work for it. Making an instrumental/acapella sounding good takes time and effort. It's not something that can be rushed. Think of it like (...) love to a woman. You wouldn't want to just rush through it, would you? Running your song through different models/algorithms, then manually filtering, EQing, noise/bleed removing the rest is a start. You can't just run a song through one of these models and expect it to immediately sound like this” - rAN

Sometimes you might want to combine results of specific models in specific song fragments.
If the song is too muddy, you might want to use demudder in newer UVR patches and/or use some free AIs like AudioSR, Apollo or other, to further enhance the result.
If you’re still not happy, you might want to manually mix separated song stems and/or master it using plugins or AI mastering services (more about it here).

Sometimes you might be capable of creating a loop out of the fade ins/outs/intros/outros so you could totally refrain from using AI separation in the key fragments of the song, so you could just only use the separation result as a reference to arrange the song as it was, and only fill missing fragments with it.

A good starting point is to have a lossless song (the result will be a bit less muddy after separation).

Now, from free separation AIs/models, to get a decent instrumental/vocals, you can use the models below, starting from those at top, which are usually the best in most cases. But be aware that every song might work differently with different models - find the best for your song. Also, various headphones and speakers might be more or less sensitive to show bleeding in your songs - lots of the time it will be imminent without using the phase fixer for most instrumental models, esp. with high fullness metric, in songs with dense mix.

Models placement on the list usually depends on model versatility, opinions of people on our Discord, frequency of use, and metrics (sometimes specific) in comparison to different models. Sometimes it depends on just deliberate or subjective choices depending on the case. Older, outperformed models are usually lower on the list, or moved out from the main section.


The best models

for specific stems

There's no such thing like the best model. It depends on a song, even in specific genre, mixing, effects, etc.
You need to test the best models posted at the top here, and see what fits the best for your song or cut it into pieces and/or check ensembles. For constant vocal buzzing check phase fixer/swapper in e.g. UVR>Tools (and use bleedless inst result of a vocal model as reference/source). Create an account using MVSEP, or you’ll have a very long queue.

- Most models here are Mel-Roformer and BS-Roformer model type in the compatible UVR version - not v2 model type
(there's only one V2 model so far); don’t confuse it with e.g. v2 versions/iterations of models below, which are just their names
- Check the reading about SDR and fullness metric. Evaluations are made on the multisong dataset on MVSEP (table can be sorted by also fullness/bleedless and other metrics, once you open a result, fullness/bleedless metrics are shown too, excluding old results)

TL;DR of the doc (shortened models list)

2 stems:

> for instrumentals (click here for inst. ensembles, or here for vocal models)

- Model names starting with MVSEP can be used only on MVSEP (no download links available)
- Metrics for Roformers are provided mostly for overlap 2 (50% on MVSEP)

Good all-rounders from various categories (balanced, fullness, bleedless):

* MVSep Becruily/ZFTurbo BS-Roformer high instrum fullness 124 bands

Only for premium MVSep users (demanding architecture).
Inst. fullness: 34.76, bleedless: 44.29, SDR: 18.47
It’s a paywalled model due to increased separation time from the model architecture. Iirc 1 credit per minute.
“It’s a fine-tune of ZFTurbo's model [BS-Roformer 124 bands], so unfortunately private” - becruily
“It’s dual model so vocals are also fullness, but whether they’re better or not than the original level 1-3 I don’t know”

Perfect mix of fullness and bleedless - dynamic64

Q: How does it compare to deux and HyperACE, any insight?

A: “I like it more”, “I feel like it functionally replaces the BS Rofo 124 fullness model” - -||-
“The instrumental fullness model makes cleaner acapellas than every other model

and the instrumentals sound so perfect its like they were phase inverted with the acapellas” - nymphi/nimfia

“For sure [it] sounds better than deux, but it has some difficulties keeping some instruments in instrumental compared to deux which is excellent at it (...) But yeah the model is amazing in how it sounds. Very high quality sounding especially with phase fix (...) [which] is optional.

I am still (...) going to use deux for some bands like Radiohead tho. Keeps more instruments as I said” - dca100fb8

- Becruily dual Mel-Roformer “deux” model (its instrumental stem)

Compatible with the latest UVR Roformer patch (including the RTX 5000 one) | makidanyee’s Colab | uvronline.app | mvsep.com (tips for both sites)

Inst. fullness 34.25, bleedless 41.36, SDR 17.55
“The best fullness model and does not even need phase fix. SOTA even (...) my favourite fullness model” - dca100fb8

“The best fullness to bleedless compromise” - becruily

On uvronline it uses vocal stem as phase fixer reference, and it happens automatically, so you don't have to click on the correct phase button (previously it was available only in premium, not sure how now). “First, phase inversion is performed, resulting in a muddy instrumental. But the result is only used for phase correction. The second stem, which is available for download/listening, is more full.”

Consider using MDX23C Phantom Center model by gilliaan as a preprocessor for better results.

If you have severe cross-bleeding with some specific songs, proceed with other models listed below, or see ensembles. Deux notes:

“More bleedless than resurrection inst (...) on a song with mostly piano, sounds considerably less muddy than resurrection inst… maybe overall slightly more noise... but the noise isn't bothersome”, (...) deux just seems to make some things super muddy and/or break them up... like an instrument here is just really wobbly with deux, probably because it doesn't have the fullness required.

HyperACE sounds better in areas here, but noisy in others… I guess you really gotta use all 3 models if you want a great result”

Resurrection inst seems to be also fuller at times, but while being also noisier, and when the song has the noise in it intentionally, Resurrection inst tend to pick it up. - Rainboomdash.

“Love the new model, besides it being on par with v1e in terms of fullness (on some tracks it's even fuller) without the noise tradeoff, it's also noticeably better in terms of bleedlessness as well, it managed to completely remove some faint choir/vocals pretty much all other models weren't able to remove. Most definitely my favorite model so far.” - Shintaro

“I think V1e+ is still fuller, sounds more open” - rainboomdash

Doesn't work well with VHS recordings - MarsEverythingTech

Doesn’t eat flute - mohammedmehditber

More sensitive to keep some tiny sounds over gaboxflowersv10 - sakkuhantano

“Worse than HyperACE v2 inst to remove scratching and beatbox” - dca

“I do recommend trying a higher chunk_size for instrumentals with deux.

I find 661500 works for a lot of songs, 749700 for a good amount of others.

Higher=more fullness, less bleedless (...) 705600 might be a better default setting -

highest SDR and more chance it won't be noisy” - rainboomdash (it complies with jarredou's measurements of the model with different chunks). For vocals, the default 570K is well balanced. “going beyond 882000 may sharply decrease SDR and other metrics.” - makidanyee

The stem named “other” on Colab is the muddiest, and the most bleedless, reminding vocal models results.

Some people like using it with overlap 8.

Trained “almost from scratch”

- Unwa bs_roformer_inst_HyperACEv2 model (incompatible with UVR, use MSST) makidanyee’s Colab | MVSEP | nextgen.uvronline

Inst. fullness 38.03, bleedless 37.87, SDR 17.40

You need to replace bs_roformer.py in MSST for it (for v2_voc and v2_inst, .py is the same).

“Definitely way better than v1e+. But I have also noticed it's more noisy” - rainboomdash

“Picks up vocals better than the first model, but along with the vocals, it muffles other instruments, such as a synthesizer.” - Halif
Much slower than Unwa BS-Resurrection inst.

The aura_mrstft score was further improved, and the SDR also increased. (...)

Incidentally, this new model outperformed v1e+ on all metrics. (...) holds the highest aura_mrstft score on the instrumental side of the Multisong dataset.” - unwa

https://mvsep.com/quality_checker/entry/9475

“Kept the chants in [one song], resurrection inst isn't really much better, though...

It takes a high vocal fullness model to extract these (like fv7beta1, even fv7beta2 isn't enough), maybe that's why the instrumental models are getting confused... Surprised it's still very much there with resurrection inst…

I'm surprised, HyperACEv2 fixed vocal bleed on one song compared to the previous version. Does seem like vocal bleed with HyperACE v2 is a bit better. It's not removed, just quieter in areas where it did happen (...) I've found HyperACE works extremely well with a lot of acoustic songs, as it handles it well and has low vocal bleed

unlike deux which has a lot of vocal bleed due to the song not being very loud” - rainboomdash

Good at preserving SFX when phase fixed - wancitte (check e.g. becruily vocal or BS 2025.07 on MVSEP instrumental result as reference for phase fixer)

Doesn't remove vocoder - stray_kids_filters

- Unwa BS-Roformer Resurrection inst (yaml) | a.k.a. “unwa high fullness inst" on MVSEP | uvronline.app/x-minus.pro | MK Colab | Kaggle | UVR (don’t confuse with Resurrection vocals variant)

Inst. fullness: 34.93, bleedless: 40.14, SDR: 17.25

MVSEP BS 2025.07 works as a reference for phase fix with 3000/5000 settings.

Consider using MDX23C Phantom Center model by gilliaan as a preprocessor for better results.

Only 200MB. Some people might prefer it over V1e+, although it’s more muddy.

“use if the others [below] are noisy”

Models working for phase fixer (to alleviate the noise) are only BS-Roformer 1296/1297 by viperx and BS Large V1 by unwa, but generally the model might require phase fixing less than other models here - dca

“One of my favorite fullness inst models ATM. Sounds like v1e to me, but cleaner. Especially with guitar/piano where v1e tended to add more phase distortion, I guess that's what you'd call it lol. This model preserves their purity better IMO” - Musicalman

“I like Resurrection inst for segments of piano, a lot of other models are too noisy there (...) I also needed to turn the overlap up for piano” (from 2 to 8). FNO was less noisy for it, but “the hit to fullness was extremely apparent” - rainboomdash

“The way it sounds, is indeed the best fullness model, it's like between v1e and v1e+, so not so noisy and full enough, though it creates problems with instruments gone in the instrumental sadly, but apparently it seems Roformer inst models will always have problems with instruments it seems, seems like a rule. (...) Instrument preservation (...) is between v1e and v1e+” - dca100fb8

“it seems to just nip some bits of random instruments like saxophone or guitar whereas v1e+ leaves them intact.” - dennis777

“In some songs leaves vocal residues. It is heard little but felt” - Fabio

“Almost loses some sounds that v1e+ picks up just fine” - neoculture

Mushes some synths a bit in e.g. trap/drill tune compared to inst Mel-Roformers like INSTV7/Becruily/FVX/inst3, but the residues/vocal shells are a bit quieter, although the clarity is also decreased a bit. Kind of a trade.

BS 2025.07/BS 2024.04/BS 2024.08/SW removes less noise than viperx models for phase fixer.

- Unwa Mel-Roformer V1e+ model (yaml) | UVR (guide) | MVSEP | x-minus/uvronline | MK Colab | SESA Colab | Huggingface / 2 | Kaggle

Inst. fullness: 37.89, bleedless: 36.53, SDR: 16.65

*) Phase fixer Colab (e.g. with FT3 as source or becruily voc)/UVR>Tools, or on x-minus with becruily vocal model used as reference model (premium) - for less noise/script.

*) introC script to get rid of vocal leakage in this model

*) “If you use Gabox Mel denoise/debleed model | yaml | MK Colab | on mixture, then put the “denoised” inst stem of that into unwa inst v1e+ you get a very clean result with good fullness and very little noise” - 5b. But it can’t remove vocal residues, just vocal noise.

Might also sound interesting when using as target in Phase fixer, and with source set as Becruily inst model (overlap 50/chunk_size 112455 was used; very slow - gustownis).

Or bigbeta5e as a source to get rid of vocal residues - santilli_

(single model inference descriptions below)

“strange leakage [robot-like] in the vocal-only section with no instrumentation” - Unwa

”less noise than v1e (probably due to different loss function), but it’s also less full, “somewhere between v1e and v1”

“sometimes a detail piece of instrumental sound was lost, while on becruily inst [below] can pick that sound”. Might be too strong for MIDI sounds - kittykatkat_uwu.

Problems with broken lead-ins not happening in instv7 and v1e. Some issues with cymbals bleed in vocals - dca.
Better than v1+. ”has fewer problems with quiet vocals in instrumentals than the V1+, “issues with harmonica, saxophone, electric guitar and synth seem to have been fixed” - dca100fb8.
“has this faint pitched noise whenever vocals hit in dead silence, you may need to manually cut it out.” - dynamic64. Check out also BS_ResurrectioN later below, it’s like v1e++ (more fullness).

- Unwa BS-Roformer-HyperACE inst (a.k.a. v1) | MK Colab | uvronline.app (doesn’t work in UVR)

Inst., fullness 36.91, bleedless 38.77, SDR 17.27

“sounding just like v1e+ after phase fix, but straight out of one single model

(...)  quite bleedy, but honestly, it's a fair price to pay, I guess” - santilli_

Although for some people it can be even on pair with v1e+ bleed-wise, so check it out too (more fullness).

JFYI - there was no v1 of the voc variant.

Note: It uses its own inference script. “You can use this model by replacing the MSST repository's models/bs_roformer.py with the repository's bs_roformer.py.”

To not affect functionality of other BS-Roformer models by it, you can add it as new model_type by editing utils/settings.py and models/bs_roformer/init.py here (thx anvuew).

For error while installing py file for HyperACE model in Sucial’s WebUI:

from models.bs_roformer.attend import Attend

ModuleNotFoundError: No module named 'models'"

The fix: “SUC-DriverOld/MSST-WebUI use the name "modules" and ZFTurbo/Music-Source-Separation-Training use the name "models". And Unwa's  bs_roformer.py that you replace with, also use "models". So you'll have to do some coding and symlink to make it work.” - fjordfish

Metrically less fullness than v1e+: 37.89, but more bleedless: 36.53, SDR: 16.65 (v1e+).

While using locally, consider changing overlap from default 4 to 2 in the yaml of the model. The difference won’t be really noticeable for most people, but it will be faster.

“Currently, this model holds the highest aura_mrstft score on the instrumental side of the Multisong dataset. (...)

This weight is based on the following weights. Thank you, anvuew!” - unwa

“Does seem like HyperACE is picking up more instruments than v1e+

does seem like slightly worse vocal bleed overall (still need to test this more, though)... haven't encountered the super tinny vocal bleed like v1e+, at least

still fails to pick up that brass instrument on one song... Not really any worse than v1e+, though (...) resurrection inst does sound more muddy, but also a lot less noise... which makes sense... IDK, a little muddy for my tastes.

I did find one song/spot and resurrection inst was on par with HyperACE in picking up the wind instrument, v1e+ lost it for a bit.

I have found in the past that resurrection inst generally picks up more instruments than v1e+ (...) fullness of HyperACE is much closer to v1e+ than resurrection inst (...) it gets pretty staticy compared to v1e+ [on some drums] (...) v1e+ does this to a lot less extent

it's not super common, though… (...) I'm very confident in saying BSHyperACE  picks up more stuff than v1e+.

Resurrection inst does pick it up much better than v1e+, but I think it's still too quiet

resurrection inst really does just pick up so much more instruments, despite having a lot less fullness” - rainboomdash

”fullness that is comparable to v1e+, but has significant more vocal crossbleeding in instrumental than BS Roformer Resurrection Inst, but still less than v1e+ and v1e” - dca100fb8

“I’ve found myself using a mix of both HyperACE v1 and deux.

I'm doing this literally on a Joy Division track right now. Convert both. Invert. Spectral down the residuals just enough to make the buzz less audible without going full Deux. I'm glad to have both models.

Convert the original track with each of the two models. Invert one against the other to get all the extra 'vocal residue' that the HyperACE model has compared to the deux model. Then use spectral editing to lower the volume of the vocal residue and mixing it back into the Deux model to create a kind of 'inbetween' result. Usually, I leave the higher frequencies alone (above 5 or 7k or so).” - CC Karaoke

- Gabox inst_gaboxFlowersV10 Mel-Roformer (yaml) | MK Colab | nextgen.uvronline or via special link (for free/premium accounts) | MVSEP

Inst. fullness 37.12, bleedless 36.68, SDR 16.95 

For UVR RTX 5000 patch and torch._dynamo error, edit “use_torch_checkpoint: True” to False (or delete the line).

All of these metrics are better than Inst_FV8b).

The yaml is unique from the previous ones. Deux model fine-tune.
For less buzzing, use phase fix with becruily voc model - prodbyluke (phase fixer Colab), but ft2 bleedless instead seems to be even more effective than even becruily voc/inst setting - Squid.

Chunk size for the flowers model in the Colab should be 352800 (yaml value, might be not evaluated).

“a huge improvement over what Gabox has accustomed us to! I got amazing results. It truly achieves a stability that previous models lacked; the clarity is much greater in these.” - billieoconnell.

It might have louder constant buzzing vs even inst_fv4 without phase fixer, but the FlowersV10 fixes some issues of crossbleeding existing in previous models, and also with some distorted vocals and screams (dtn, dca100fb8, Tobias51).

- unwa BS-Roformer-Leap Xe instrumental

https://huggingface.co/pcunwa/BS-Roformer-Leap/tree/main/Xe | Colab

Instrumental fullness 35, bleedless 41.46, SDR: 17.53

It's 90-band. A bit better bleedless/fullness, but a bit worse SDR vs Deux (potentially a bit worse at recognizing instruments).

“Big issue (...) is BV bleed, sometimes (...) kept almost intact (...) Very good model in how it sounds like (...) Fullness is very good, sounds fuller than deux and HyperACE, but less full than Flowers. Haven't found any tracks where it's muddy. (...) most of the tracks I tested it on, it really has almost no bothering noise (...) phase fix doesn't need to be always applied (...) quite song-dependent, which is not the case with deux (...) For keeping instruments, it's better than HyperACE , but worse than Deux. (...) Considering SFX vocals sometimes [they’re] kept and sometimes [they’re] removed, so I don't think it has been improved since previous models. Deux was also not good at it, while Flowers shines for this.” - dca100fb8

“Sometimes leaves a bit of faint "crust" behind but still really fantastic stuff here” - pixelrino

“Mids sound a tiny bit muffled, but there's a lot of hf noise from about 7 kHz upward. I can hear what it's going for, it's trying to prevent the annoying kind of muddiness you get with a really bleedless model, but it's putting a lot of noise in a region I'm sensitive to.” - Musicalman

“a test of the new loss function.

It looks like it still needs some improvement.” - unwa

- Gabox Inst_GaboxFv8 v2 Mel-Roformer (yaml) | MK Colab (don’t confuse with Inst_FV8 or INSTV8)

Inst. fullness: 33.21, bleedless: 40.73, SDR: 16.57

Usually referred to as just Inst_GaboxFv8 without v2. Since its release, the checkpoint has been updated on 11.05.25 (same file name), metrics have changed (updated above).

It’s not the one on uvronline.
“Good result for bleedless instead, fullness went down instead of up a little.”
Might be an interesting competitor to Unwa inst v2 which is muddier.

Inst. fullness: 35.57, bleedless: 38.06, SDR: 16.51 are the metrics of the old v1 model.


_____

- Gabox Inst_GaboxFv7z Mel Roformer (yaml) | MK Colab | nextgen.uvronline/x-minus.pro

Inst. fullness: 29.96, bleedless: 44.61, SDR: 16.62

(duplicate from the above)

Becruily vocal used for phase fixer on x-minus.pro/uvronline (premium feature).

“Focusing on the less amount of noise, keeping fullness”

“the results were similar to INSTV7 but with less noise” but “the drums are totally fine with this model ”- neoculture

“it seems to capture some vocals better” - Gabox

In some songs, “it leaves a lot of reverb or noise from the vocals. unva v1e+  a little better” - GameAgainPL

“[one of the] best bleedless, good fullness, almost noiseless” - Aufr33

- Gabox Inst_FV8 (yaml) a.k.a. “Mel-Roformer by Gabox V8” on nextgen.uvronline (test) | MK Colab. Previously exclusive uvronline model via special link (free/premium)

Experimental; don’t confuse with “InstGaboxFv8”.

It gives decent results too. “it's on the higher bleedless side” - Rainboomdash (although less bleedless than fv7b and fv7z).

Previous epoch of the FV8B model, previously only on uvronline.

- MVSep SCNet vocals model: SCNet XL IHF (high instrum fullness by bercuily).

Inst. fullness 32.31, inst. bleedless 38.15, SDR 17.20

“One of my favorite instrumental models, Roformer-like quality.

For busy songs it works great, for trap/acoustic etc. Roformer is better due to SCNet bleed” - becruily

“bring[s] such near perfect instrumentals”

vs the previous XL models “It's high fullness version for instrumental prepared by becruily.”

It can also be an insane vocal model too.

Still good, less frequently used models

- Gabox INSTV7 a.k.a. Inst_GaboxV7 Mel-Roformer (yaml) | MVSEP | MK Colab | Huggingface / 2  | “F”, for fullness, “V” for version.

Inst. fullness: 33.22, bleedless: 40.71, SDR: 16.51
*) Phase fixer Colab/UVR’s Phase Swapper (for less noise; e.g. with FT3 by Unwa vocal model as source).

“I hear less noise compared to v1e, but it has a worse bleedless metric” and might be less full.

It might still have too much noise like v1e for some people, but less.
“Relatively full/noisy model. Fvx [below] is a sort of middle ground between v3 and v7.”
More fullness than V6, but vs v1e, sometimes “leaves noises throughout the song, sometimes vocal remnants in the verse of the song, and some instruments are erased.”
Less muddy than Mel 2024.10 on MVSEP, and V7 doesn’t preserve vocal chops/SFX.

Gabox claims it's not the same as V7 on uvronline, but FV7z, despite both appear on the list while using a special link for: free/premium.

- Becruily’s inst model | MK Colab | Huggingface / 2 | on MVSEP a.k.a. Mel-Roformer “high fullness” | uvronline via special link for: free/premium (scroll down) or nextgen.uvronline.app

Inst. fullness 33.98, inst. bleedless 40.48, SDR 16.47
or on x-minus/uvronline (with optional phase correction feature in premium) | UVR
*) It's an older model, but it’s great to get less vocal residues when used in phase fixer Colab (also in UVR>Tools) and “becruily's vocals as source and inst as target”.
If it’s too muddy, consider using a result of nosier single model as Inst_GaboxFv8_v1 (or some fullness models) as reference for Matchering in e.g. UVR.

Alone, Becruily’s inst is as clean as unwa’s v1, but has less noise, and it can also be got rid well by:
*) Mel denoise and/or Roformer bleed suppressor by unwa/97chris. That model “removed some of the faint vocals that even the bleed suppressor didn't manage to filter out” before”. Doesn’t require phase fix. Try out denoising on a mixture first, then use the model.

On its own, the inst model correctly removes SFX voices. The instrumental model pulled out more adlibs than the released vocal model variant, when it can pull out nothing.
Currently, the only model capable of keeping vocal chops.
“Struggles a lot with low passed vocals”

More instruments correctly recognized as instruments and not vocals, although not as much as Mel 2024.10 & BS 2024.08 on MVSEP, but still more than unwa’s inst v1e/v1/v2.

- If you use lower dim_t like 256 (or maybe also corresponding chunk_size) on weaker GPUs, these are the first Mel inst models to have muddy results with it.

- In the phase fixer you can experiment with “using becruily's vocals as source and inst as target, and changing high frequency weight from 0.8 to 2 makes for impressive results” you can do it automatically after separation in this Colab (santilli_ suggestion).
Using Kim Mel FT2 as source instead might be more problematic as it tends to be more harmful to instruments, and in noise removal both are similar (dca).

- To demudd the results from phase fixer, you can use Matchering and a well sounding fragment of single instrumental model separation with high fullness metric (e.g. 7N) as a reference and becruily inst/voc phase fixed result set as target (e.g. in UVR>Tools>Matchering). It will have less bleeding than models with low bleedless metric, but still fuller than phase-fixed results (more here and here). Phase fixer can also be used in a standalone Python script or in the latest UVR. Matchering can be used in Colab or songmastr or locally (it’s very lightweight and doesn’t require a GPU).

- Gabox inst_fv4 Mel-Roformer (yaml) | MK Colab

Inst. fullness 39.40, bleedless 33.49, SDR 16.44

Don’t confuse it with inst_fv4noise - the regular variant was never released before (and with voc_fv4).

“Seems to be erasing a xylophone instrument. Does sound not too noisy and not muddy, I like it. (...) A little noisy with piano (I split the song up and process with resurrection inst there). (...) Does have some issues that resurrection inst doesn't have, but it doesn't sound muddy! It usually works great. (...) In my opinion, fv4 still has vocal traces, I don't know if in all of its songs and v1e plus doesn't have them, but the noise can bother you even though it's not much. Does have more vocal bleed at times. I think a lot of what I thought was vocal bleed was a synth, it did a pretty good job... There was one segment on a song where it caught vocal residues, though” - rainboomdash

It has more constant buzzing than becruily inst, but it’s fuller (might be a no-go if you’re sensitive to it e.g. using headphones).

- Gabox Inst_FV8b Mel-Roformer (yaml) | MK Colab
Inst. fullness: 35.05, bleedless: 36.90, SDR 16.59.

Muddier than V1e+ (37.89), but cleaner. Some people might prefer it over INSTV7.

“Preserves its volume stability to the original sound of the songs, it does not go down or lose strength, which is the most important thing, it manages to capture clear vocal chops, the voice is eliminated to 99 or 100% depending on its condition, it captures the entire instrumental and when making a mix it remains like the original that with other models the volume was lowered,” - Billie O’Connell

Recent bleedless models

- Gabox Inst_GaboxFv7z Mel Roformer (yaml) | MK Colab | uvronline.app/x-minus.pro

Inst. fullness: 29.96, bleedless: 44.61, SDR: 16.62

(duplicate from the above)

Becruily vocal used for phase fixer on x-minus.pro/uvronline (premium feature).

“Focusing on the less amount of noise, keeping fullness”

“the results were similar to INSTV7 but with less noise” but “the drums are totally fine with this model ”- neoculture

“it seems to capture some vocals better” - Gabox

In some songs, “it leaves a lot of reverb or noise from the vocals. unva v1e+  a little better” - GameAgainPL

“[one of the] best bleedless, good fullness, almost noiseless” - Aufr33

Despite the metrics, it can have more constant buzzing than fv7b.

- Gabox inst_fv7b Mel Roformer (yaml) | MK Colab

Inst. fullness 27.07, bleedless 47.49, SDR 16.71

(duplicate from the above)

Fullness worse than even most vocal Mel-Roformers (incl. BS-RoFormer SW and Mel Kim OG model).

“on the fuller side, somewhere around inst v1e+, maybe a tiny bit below. The main thing I notice is it captures more instruments than v1e+, but isn't muddy like [HyperACE] (which also captures more instruments)

can be a little on the noisy side sometimes... but it at least isn't muddy and sounds natural (...) I'd still ensemble if you want the noise reduced - rainboomdash (src)

In some lo-fi beats, it can keep muffling hi-hats by being too muddy.

0) Rifforge by mesk model final | MK Colab

Can be even more destructive for drums than the two above, muddier, but it has even less buzzing by straight up picking noise existing in the mixture along with vocals. It’s slow and big. Generally dedicated to metal, achieving better results there, but can work for e.g. hip-hop (along with caveats above).

- Unwa BS-Roformer-Inst-FNO | MK Colab

Inst. fullness: 32.03, bleedless: 42.87, SDR: 17.60

Incompatible with UVR, install MSST, then read model instructions here (requires modifying bs_roformer.py file in MSST, also models_utils.py for PyTorch newer than 2.6).

Actually similar results to BS-Resurrection inst model above, less fullness.

Some people even prefer Gabox BS_ResurrectioN instead.

“Very small amount of noise compared to other fullness inst models, while keeping enough fullness IMO. I don't even know if a phase fix is needed. Maybe it's still needed a little bit.” dca

“seems less full than the Resurrection, which I would expect given the MVSEP [metric] results. (...) I'd say it's roughly comparable to Gabox inst v7”

“I replaced the MLP of the BS-Roformer mask estimator with FNO1d [Fourier Neural Operator], froze everything except the mask estimator, and trained it, which yielded good results. (...) While MLP is a universal function approximator, FNO learns mappings (operators) on function spaces.”

“(The base weight is Resurrection Inst)”

- Unwa BS-Roformer-Large-Inst a.k.a. bs_large_v2_inst (there was no v1) | MK Colab

Fullness: 32.06, bleedless: 43.95, SDR: 17.61

It uses a custom bs_roformer.py file attached - just replace it in the MSST installation (UVR is incompatible).

Training details:

"Instead of increasing the depth to 16, I added a four-layer TransformerBlock to the MaskEstimator." 238MB weight.

“sometimes was good and other times it completely missed the separations”, more noise than flowersv10 model - mesk (spectrogram example with lacking fragments missed by the model). “It leaves residue in some places” - fabio06844

“Just looking at metrics, fullness is ranked about the same as (...) FNO” - rainboomdash

Lower fullness models 

(if you find the ones above too muddy, but here you get more noise)

0) Gabox inst_gabox3 (yaml) | MK Colab | Huggingface / 2 | Phase fixer Colab

Inst. fullness 37.69, bleedless 35.93, SDR 16.50
Actually worse fullness than v1e+ (37.89), and lower bleedless (36.53).

When used with Unwa’s beta 6 as reference for phase fixer (thx John UVR), slightly less muddy results than phase-fixed Becruily inst-voc results, but also slightly more vocal residues and a bit more inconsistent sound, fluctuations across the whole separation at times.

0) Gabox INSTV7N | MK Colab | Huggingface / 2 

Inst. fullness 36.83, bleedless 35.47, SDR: 16.65

More noisy than INSTV7; “it's [even] closer to v7 than inst3”

- Inst_GaboxFv8 v1 Mel-Roformer (yaml) | fork of makidanyee’s Colab
Inst. fullness: 35.57, bleedless: 38.06, SDR: 16.51
The OG link to the model changed to the v2 variant of the model, but the old link to the v1 was retrieved above
(the model name remained the same and if you used the regular non-v1 in the Colab before it, the ckpt won’t be replaced) .

It has “v1+ metallic noise” - Gabox. And even nasty residues louder than buzzing vs the v2 variant.

VS V1e+ “A bit cleaner-sounding and has less filtering/watery artifacts. Both models are prone to very strange vocal leakage [“especially in the chorus”].

And because Fv8 can be so clean at times, the leakage can be fairly obvious. For now, my vote is for Fv8, but I'll still probably be switching back and forth a lot. Still has ringing” - Musicalman. Although, you might still prefer it over V1e+.

Might have some “ugly vocal residues” at times (Phil Collins - In The Air Tonight) - 00:46, 02:56 - dca.

“Sometimes V1e+ has vocal residues which sound like you were speaking through a fan/low quality mp3” - dca

”Seems to pick up some instruments better” Gabox.

__

0) SCNet XL model called “very high fullness” | MVSEP

Inst. fullness 34.04, bleedless 35.15, SDR 16.60

It might work better than Roformers for less noisy/loud/busy mixes or genres like alt-pop, orchestral tracks with choir, sometimes giving more full results than even v1e, but at the cost of more noise. Might struggle with some vocal reverbs or effects.

"One thing with SCNet. If you process the song twice, once with the phase inverted, re-invert the phase and then mix together you will cancel most of the noise" - Dry Paint Dealer Undr
“Very hit or miss. When they're good they're really good but when they're bad there's nothing you can do other than use a different model”
Compared to the high fullness variant, more crossbleeding of vocals in instrumentals (along with SCNet XL basic model). Some songs which sound full enough even with basic SCNet XL (and HF variant) while others will sound muddy (dca)

“has a lot of noise/bleed, and I haven't found the best way to get rid of it, but it does tend to pick up harmonies and subtle BGV that other models don't.” dynamic64

0) MVSEP SCNet XL high fullness

Inst. fullness 31.95, bleedless 34.06, SDR 17.26

“I have a few examples where it's better than v1e+

Sometimes there is too much residue but most of the time it's fine” dca
“Really loving the way SCnet high fullness [variant] handles lower frequencies, below 2K [let’s] say. Roformers are better with the transients up high, but decay on guitars/keys on the SCnet is more natural”

“seems to also confuse less "difficult" instruments for vocals”

“I noticed classic SCNet XL preserves more instruments than the high fullness one, but has more vocal crossbleeding in instrumental compared to high fullness

So if you want instrument preservation use SCNet XL 1727 but if you want less crossbleeding of vocals in instrumental use SCNet XL high fullness

I ignore the very high fullness one because it has too much vocal residue” dca

(regular SCNet XL moved below)

_______

- Gabox BS_ResurrectioN model | yaml | MK Colab
“It is a fine-tune of BS Roformer Resurrection Inst but with higher fullness (like v1e for example), it needs [MVSEP’s] BS 2025.07 (as a source/reference) phase fix
I requested it because I found some songs where Resur Inst was producing muddy instrum results (...) I requested it not just for me because I saw other people were looking for something like v1e++” - dca
Higher fullness (but with more noise)

(sorted by fullness)

0) Gabox INSTV6N (N for noise/fullness) | yaml | MK Colab | SESA | Huggingface / 2 
Inst. fullness: 41.68 (more than v1e), bleedless: 32.63, SDR: 16.35'

N - “noisier but fuller”. To get rid of noise in INSTV6N, use Gabox denoise/debleed model (yaml) on mixture first, then use INSTV6N - “for some reason it gives cleaner results” (Gabox), but it can’t remove vocal residues.

Some people find it having less noise vs v1e and more fullness.
Also, it has more fullness vs INSTV6, and more noise, but some people might still prefer v1e.

“v1e sounds like an "overall" noise on the song, while v6n kind of mixes into it.

v6n also sounds like two layers, one of noise that's just there. And the other one mixes into the song somehow. Using the phase swap barely makes it any better than phase swapping with v1e though” - vernight
Also Kim model for phase swap seems to give less noise than unwa ft2 bleedless

“Comparing V6N with v1e and couldn't hear a fullness difference despite the metrics being approx 39 for v1e and 41 for V6N” - dca

“my all-time favorite” - ezequielcasas

0) Gabox INSTV5N | yaml | SESA
Inst. fullness 40.47, bleedless 32.73, SDR 16.41
“Despite (...) having more fullness in metrics, didn't sound fuller than v1e. And I found that weird” - dca

0) Gabox inst_Fv4Noise | yaml | MK Colab | SESA Colab | Huggingface / 2

Inst. fullness 40.40, bleedless 28.57, SDR 15.25

Can be better than INSTV6 for some people, but overkill for others. Bigger fullness metric than even v1e.

“Despite v4's significant amount of noise, it seems to be the only model [till 8 February] that gave me a fuller sounding result compared to v1e that's actually perceivable by my ears.” - Shintaro

“although the fullness metric increases when there is more noise, it doesn't always mean it's a better instrumental — an example of this is the fv4noise metrics” - Gabox

0) Unwa Inst V1e (don’t confuse with newer +/plus variant above) | yaml (from v1)

Inst. fullness 38.87, bleedless 35.59, SDR 16.37

MK Colab | MSST-GUI | UVR instructions | Huggingface / 2 | uvronline via special link for: free/premium (scroll down) or nextgen.uvronline | MVSEP

One of the first Mel Kim model fine-tunes trained with instrumental (other) target. High fullness metric, noisy at times and on some songs. To alleviate it, it can be used with automated phase fixer Colab or UVR>Tools (Kim Mel as reference removes more noise than 2024.10 vs muddier Unwa v1/2 on their own; optionally use VOCALS-MelBand-Roformer by Becruily or unwa's kim ft; you can also use FT2 as reference, but it “cuts instruments” vs FT3 which can be rather better alternative). Optionally, in Phase Fixer you can set 420 for low and 4200 for high or 500 for both and Mel-Kim model for source; and bleed suppressor (by unwa/97chris) to alleviate the noise further (e.g. phase fixer on its own works better with v1 model to alleviate the residues). Besides the default UVR default 500/5000 and Colab default 500/9000 values, you could potentially “even try like 200/1000 or even below for 2nd value.”  “I would say that the more noisy the input is, the lower you have to set the frequency for the phase fixer.” Later you can ensemble the result with BS 2025 on MVSep or the BS-Roformer SW inst result Max Spec/FFT, it can even surpass Deux+HyperAce v2 ensemble, and have less noise and sound fuller (dca).

V1e on its own might catch more instruments and vocals than INSTV6N. Even fuller model with more noise is instfv4noise below by Gabox.

“The "e" stands for emphasis, indicating that this is a model that emphasizes fullness.”
“However, compared to v1, while the fullness score has increased, there is a possibility that noise has also increased.” “lighter compared to v2.” Like other unwa’s models, it can struggle with flute, sax and trumpet (unlike Mel 2024.10, and BS 2024.08 on MVSEP respectively - you can max ensemble all the three as a fix [dca100fb8]). Also, sometimes unwa's big beta5e can retrieve missing instruments vs v1e when those two above fails. “The fullness of v1e is constant, on every song I tried” flowersv10 is almost as full” - dca.
Possible residues of dual layer vocals from suno songs.

0) inst_gaboxFv3 | yaml | Huggingface / 2 - F for fullness

inst. fullness 38.71, inst. bleedless 35.62 (“F” stands for fuillness) | Inst SDR 16.43
Like v1e when it comes to fullness, but less bleeding.

Vs v1e “it's slightly better with some instruments”, It might pick up an entire sax in the vocal stem.

It doesn't have that weird fullness noise that fullness models produce, but still gives pretty full results and the phase swapper (with big beta 6 as reference) gets rid of that weird buzzing sound” John UVR

0) Gabox experimental “fullness.ckpt” inst Mel-Roformer (yaml).

Inst. fullness: 37.66, bleedless: 35.53, SDR: 15.91

“this isn't called fullness.ckpt for nothing.” - Musicalman

0) Gabox inst_gaboxFv1 a.k.a. model F / fullness (yaml) | MK Colab | Huggingface / 2
Inst. fullness: 37.29, bleedless: 37.18, SDR: 16.47 (or 37.26/37.19)

It has more unpleasant buzzing than the FlowersV10, but both have differently sounding buzzing too, but still the latter has generally less.

_____________________________________

Sorted by the biggest instrumental fullness metric on the list:

INSTV6N (41.68)>INSTV5N (40.47)>inst_Fv4Noise (40.40)>Inst V1e (38.87)>Inst Fv3 (38.71)>INSTV7N (36.83).
While V1e+ (37.89) might be already muddy in some cases.

Sorted by bleedless metric here

_____________________________________

Lower bleedless models/balanced

Still less noise even when using without phase fixer

0) Unwa BS-Roformer Resurrection inst (yaml) | a.k.a. “unwa high fullness inst" on MVSEP | uvronline.app/x-minus.pro | MK Colab | UVR (don’t confuse with Resurrection vocals variant)

Inst. fullness: 34.93, bleedless: 40.14, SDR: 17.25

(duplicate from the above, because it fits metrically and categorization-wise here, more info above)

0) Unwa Mel-Roformer inst v1 (yaml) | MK Colab | UVR installation | MVSEP | uvronline via special link for: free/premium (scroll down) or nextgen.uvronline
inst. fullness 35.69, bleedless 37.59
*) Denoising for v1/2/1e recommended with: 1) ensemble noise/phase fix option for x-minus premium 1b) Becruily phase fixer | Colab (also since UVR beta patch #7) 2) Mel-Roformer de-noise non-agg. (might be better solution) 3) UVR-Denoise medium aggression (default for free users) 4) minimum aggression for premium/link (damages some instruments less) 5) UVR-Denoise-Lite [agg. 4, no TTA] in UVR - more aggressive method 6) UVR-Denoise [agg. 30/25, hi-end proc., 320 w.s., p.pr.] - even more muddy but preserves trumpets better

v1 might have more instruments missing vs v1e and less noise

0) inst_gabox2 (yaml) | Huggingface / 2

inst. fullness 36.03, bleedless: 38.02

-

0) Becruily inst model (again, because it fits here metrically)

Inst. fullness 33.98, bleedless 40.48, SDR 16.47
MK Colab | on MVSEP a.k.a. “high fullness” (the same model) | x-minus (w/ optional phase correction feature in premium) | UVR
For less vocal residues use phase fixer Colab (also in UVR>Tools) and “becruily's vocals as source and inst as target”

0) Inst_GaboxFv8 v2 model (yaml) | MK Colab

Inst. fullness: 33.21, bleedless: 40.73, SDR: 16.57

(again, just for metrics)

0) Gabox instV7plus bleedless model (experimental)

inst. fullness: 29.83, bleedless: 39.36, SDR 16.51
_


0) MVSEP SCNet XL (don’t confuse with undertrained weights on ZFTurbo’s GitHub)

inst. fullness 28.74, bleedless 39.42, SDR 17.27

“I've come across a lot of songs where high fullness [SCNet variant above] gives that annoying static noise. I'm starting to like basic SCNet XL more to the high fullness [model]. And also, less vocal residues.” - dca. There is crossbleeding of vocals in some songs. You can find the dca’s list for that model in further parts of this section.

0) MVSEP SCNet XL IHF

inst. fullness 28.87, bleedless 40.37, SDR 17.41

Some songs struggling with previous models might yield better results.

0) MVSEP SCNet Large

inst. fullness 27.10, bleedless 41.47, SDR 17.05

Higher bleedless (not so full)
(as above but more models and reversed order)

0) Gabox B/bleedless v3 (inst_gaboxBv3) | Huggingface / 2

Inst. fullness: 32.13, bleedless: 41.69, SDR 16.60

“can be muddy sometimes” but still fuller than the older one below.
Sometimes more buzzing than FV8B, although fuller, although it can change across the song when it will be exactly the opposite, and also buzzier than FV7Z and FV7B


0) Unwa Mel-Roformer inst v2 (similar but fewer vocal residues (not always), muddier, bigger, heavier model)

Inst. fullness 31.85, bleedless 41.73 (less bleeding than Gabox instfv5/6)

Model files | MK Colab | Huggingface / 2 | uvronline via special link for: free/premium (scroll down) | MSST-GUI (or OG repo) | UVR Download Center)

Might miss the flute. “Sounds very similar to v1 but has less noise, pretty good” “the aforementioned noise from the V1 is less noticeable to none at all, depending on the track”.  “V2 is more muddy than V1 (on some songs), but less muddy than the Kim model. (...) [As for V1,] sometimes it's better at high frequencies” Aufr33
Might miss some samples or adlibs while cleaning inverts. SDR got a bit bigger (16.845 vs 16.595).

“Significantly less noise than v1e, sounds full enough, despite the fullness inst score, and that it recognizes more instruments than v1 and v1e, added to the fact it has higher SDR so also slightly less vocal crossbleeding in instrumental.” - dca100fb8

0) Unwa BS-Roformer-Inst-FNO | MK Colab

Inst. fullness: 32.03, bleedless: 42.87, SDR: 17.60
(again, because it fits metrically, more info moved near the top to recent bleedless section)

0) Gabox Inst_GaboxFv7z Mel Roformer (yaml) | MK Colab |  uvronline/x-minus.pro

Becruily vocal used for phase fixer on x-minus.pro/uvronline (premium).

Fullness: 29.96, bleedless: 44.61, SDR: 16.62

(again, because it fits metrically, -||-)

0) Gabox inst_fv7b Mel Roformer (yaml) | MK Colab

Inst. fullness 27.07, bleedless 47.49, SDR 16.71

(duplicate from the above)

Fullness worse than even most vocal Mel-Roformers (incl. BS-RoFormer SW and Mel Kim OG model).

“on the fuller side, somewhere around inst v1e+, maybe a tiny bit below. The main thing I notice is it captures more instruments than v1e+, but isn't muddy like [HyperACE] (which also captures more instruments)

can be a little on the noisy side sometimes... but it at least isn't muddy and sounds natural (...) I'd still ensemble if you want the noise reduced - rainboomdash (src)

In some lo-fi beats, it can keep muffling hi-hats by being too muddy.

0) Rifforge by mesk model final | MK Colab (below)

Can be more destructive for drums than the two above, more muddy, but it has even less buzzing by straight up picking noise existing in the mixture as vocals. It’s slow and big. Generally dedicated to metal, achieving better results there, but can work for hip-hop occasionally.

_________________________

Last resort  - muddier but cleaner single vocal models with more bleedless tested for instrumentals (sorted by bleedless) here | descriptions

_________________________

Special purpose models

 

- (duplicated) Mesk’s Rifforge Mel-Roformer model final | MK Colab - focused on inst/voc separation for metal music

“The model can have some quirks (just like most models) but it's all-around clean for me to release.” “It kinda also fucks up in like cleans [not distorted/non-metal vocals]”

“A pretty insane instrumental (...) I can still hear the punch (...) drums punch is preserved!!!! result is 95% since it kept some deep madness vocals at the end and there is some subtle vocal reverb left but not noticed if you didn't focus” - mohammedmehditber
Training details:
“This is a dimension 512 depth 24 model (so fairly large file size at 1.9 GB!), with an SDR of 14.2436.

It's finetuned from an older Melband Roformer checkpoint with an SDR of 13.7.”

- “I think I found the (IMO) the best process for metal:

1. inferencing using the BS 07.2025 model on MVSEP

2. inferencing using my rifforge model

3. ensembling both with a min_fft ensemble” - mesk

it keeps the "fullness" of the rifforge model being an instrumental focused model but then also removes more stuff than my base model thanks to 07.2025”

- Max Spec ensemble of Deux and inst_gaboxFlowersV10 (below) - seems to yield good result for a metal album - mohammedmehditber

- Gabox inst_gaboxFlowersV10 Mel-Roformer (yaml) | MK Colab | uvronline via special link (for free/premium accounts) - duplicated

Inst. fullness 37.12, bleedless 36.68, SDR 16.95 (all of these metrics are better than Inst_FV8b).

The yaml is unique from the previous ones. Deux model fine-tune.
For less buzzing, use phase fix with becruily voc model - prodbyluke (phase fixer Colab).

For UVR RTX 5000 patch and torch._dynamo error, edit “use_torch_checkpoint: True” to False (or delete the line).

Chunk size for the flowers model in the Colab should be 352800 (yaml value, might be not evaluated).

It might have louder constant buzzing vs even inst_fv4 without phase fixer, but the FlowersV10 fixes some issues of crossbleeding existing in previous models, and also with some distorted vocals and screams (dtn, dca100fb8, Tobias51).

“a huge improvement over what Gabox has accustomed us to! I got amazing results. It truly achieves a stability that previous models lacked; the clarity is much greater in these.” - billieoconnell.

- (old) mesk’s “Rifforge” metal Mel-Roformer 14.01 fine-tune instrumental model (focused more on bleedless).

Inst. fullness: 28.49, bleedless: 42.38, SDR 16.67

“training is still in progress, that's why it's a beta test of the model; It should work fine for a lot of things, but it HAS quirks on some tracks + to me there's some vocal stuff still audible on some tracks, I'm mostly trying to get feedback on how I could improve it” known issues.

(dead link)

Be aware that MVSep’s BS-Roformer 2025.07 can be better for metal both for vocal and instrumentals than these mesk’s models, a lot of the time. It was also trained on mesk’s metal dataset.

- Custom model import Colab has currently some issues with it. Probably, using that old version will work (at least locally).

"My old MSST repo I'm using, but I removed all the training stuff

(dead)

pip install -r requirements.txt (u gotta have Python and PyTorch installed as well) for the script to work.

You just gotta put all the tracks you want to test on in the **"tracks"** folder then double-click on **"inference.bat"** to run the inference script

its like if you were to type in the command in cmd but its simpler, and I'm lazy" - mesk

- Older Mesk metal Mel-Roformer preview instrumental model

Inst. fullness: 28.81, bleedless: 42.16, SDR 16.66
Retrained from Mel Kim on metal dataset consisting of a few thousands of songs.

https://huggingface.co/meskvlla33/metal_roformer_preview/tree/main | Colab

(previous metrics were made on private dataset)

“Should work fine for all genres of metal, but doesn't work on:

- hard compressed screams

- some background vocals

- weird tracks (think Meshuggah's "The Ayahuasca Experience")”

P.S: Use the MK Colab “or training repo (MSST) if you want to [separate] with it. UVR will be abysmally slow (because of chunk_size [introduced since UVR Roformer beta #3])”

- Neo_InstVFX Mel-Roformer by neoculture | yaml | MK Colab | Huggingface / 2

Inst. fullness 39.88, bleedless: 32.56, SDR: 14.35
Focused on preserving vocal chops.

“Great model (at least for K-pop it achieved the clarity and quality that no other model managed to have) it should be noted that it has a bit of noise even in its latest update, its stability is impressive, how it captures vocal chops, in blank spaces it does not leave a vocal record, sometimes the voice on certain occasions tries to eliminate them confusing them with noise, but in general it was a model that impressed me. It captures the instruments very clearly” - billieoconnell.

“NOISY AF, this is probably the dumbest idea ever had for an instrumental model. Don’t use it as your main one, some vocals will leak because I added tracks with vocal chops to the dataset. Just use this model for songs that have vocal chops” - neoculture
New fine-tune is in works.

- Unwa BS-Roformer-Inst-EXP-Value-Residual (uses Mel v2 model type in UVR; If it wasn’t made compatible with MSST already, replace bs_roformer.py from this repo and
from bs_roformer.attend import attend

from models.bs_roformer.attend import attend

in bs_roformer.py file

generally not very good model, but sometimes capable: “successfully removed [vocals] and kept the digital choir atmosphere as well” vs deux, inst_gaboxFlowersV10 and HyperACE but it it’s considerably slower than deux - mohammedmehditber)

Older/other models (still less muddy than vocal models)

0) INSTV6 by Gabox | yaml | x-minus | MK Colab

Inst. fullness 37.62, bleedless 35.07, SDR 16.43

v1e still gives better fullness, but the noise in it is a problem

Opinions are divided whether v5 or v6 is better.

“Seems like a mix between brecuily and unwa's models”
“Slightly better than v5 (...) less muddy and also removes the vocals without adding that low EQ effect when the vocals would come in, so I feel it's better” zzz
Old viperx’ 12xx models have fewerproblems with sax.

- Inst_GaboxFVX | yaml | Huggingface / 2

Inst. fullness 38.25, bleedless 35.35, SDR 16.49

“instv7+3” - fuller than instv3

- Gabox instv10 (experimental) | yaml 

Less noise and vocal residues than V7, but muddier

0) Gabox Mel-Roformer instrumental model “inst_gabox.ckpt” (Kim/Unwa/Becruily fine-tuned)

Gabox’ models repo | MK Colab | Huggingface / 2

inst fullness 37.07 (better than unwa inst v1 and v2), bleedless 37.40 (better than v1e by 1.8, slightly worse than unwa’s v1)

“It’s like the v1 model with phase fixer, but it gets more instruments,

like, it prevents some instruments from getting into the vocals”, “sometimes both models don't get choirs”.

_____

List of fast inference models
(small model size/potentially workable on not ancient CPUs without GPU acceleration
or slow GPUs)

_____

Older fullness models

0)  Gabox F/fullness v2 | Huggingface / 2

inst fullness 37.46 | bleedless: 37.09

*) Gabox inst_Fv4 (F - fullness/v4) | (don’t confuse with vocal fv4) | yaml | MK Colab

inst fullness 39.40 | bleedless 33.49

Duplicate from the above

Others

0) intrumental_gabox | yaml | Huggingface / 2

__

0) Gabox B/bleedless v1 instrumental model | yaml | Huggingface / 2

Inst. fullness 35.03, bleedless 39.10, SDR 16.49

0) Gabox B/bleedless v2 instrumental model | yaml | Huggingface / 2
Inst. fullness 35.09, bleedless 38.38, SDR 16.49

(Gabox models repo)


0) Cut your song into fragments consisting from the best moments of e.g. v1e/v1/v2 into one (and optionally Mel-Roformer Bas Curtiz FT on MVSEP as it will give you even less vocal bleeding, but more muddiness if necessary)

0) Propositions of models for phase fixer to alleviate vocal residues (from the above)

a) Becruily voc with Becruily inst (muddy but very few residues if any)

b) FT3 with V1e+

c) Unwa Beta 6 with inst_gabox3 (although it might be less consistent than the top)

d) Unwa Revive model is also good with any instrumental model

e) Unwa Bigbeta5 used to be not bad either.

f) Or any of the vocal models above with e.g. V1e (it's pretty full, and it might be not enough for it nevertheless)

How to use the phase fixer in UVR?

Separate with vocal model, then with instrumental model. Go to Audio Tools>Phase swapper, and use vocal model result as reference, and instrumental as target

__

Older non-Roformers moved further with descriptions here

__


Ensembles

(for instrumentals; check out also DAW ensemble with the below)

If you find some phase fixer results (e.g. ensembled) unsatisfactory, use the Phase Fixer Colab - it's tweaked for better results than UVR and standalone scripts.

Phase fixing instructions | Ensemble instructions | Ensembles for vocals.

- Iterative Colab by IntroC - click (it uses V1e, Resurrection, Revive3e, SCNet, and MVSep BS-Rofomer 2025.07 via API)

- For now you can just use Deux [or maybe even paid BS becruily 124 on MVSEP) alone because if you ensemble it e.g. with Flowers or HyperACE or SCNet, these models are greatly inferior at preserving instruments, so if you mix a model that for example keeps some instrument like didgeridoo (such as deux) but the other model doesn't, you are going to get only 50% of the instrument in the ensemble result, and that's not good - dca100fb8

- These recommendations are for batch processing purposes, if the user just wants to do it for some songs only, then phase fixer could be applied to Flowers or HyperACE before ensembling with deux, but it's a pain to do in bulk.

- Max FFT and Max Spec below can be treated interchangeably/equal

Suggestions and quotations by dca100fb8 (if not said otherwise, further redacted):

Best ensembles (MVSep-exclusive currently):

- #1 - BS becruily 124 fullness + 2025.07 (Max FFT/Spec)

- #2 - BS becruily 124 fullness + SW (Max FFT)

--->These work well to fix missing instruments occurring with becruily 124 fullness, without adding vocal crossbleed.

- #3 - BS becruily 124 fullness + Mel Deux (Avg)

- #4 - BS becruily 124 fullness + SCNet XL IHF fullness (Avg)

Community models ensembles:

#1 - Mel Inst_GaboxFv9 + BS SW (Max Spec)

—->New best open ensemble for instrumental

(by nextgen.uvronline.app user, dca-approved),

can be used there when post processing options to show the ensemble menu for logged users is enabled.

- #2 - Mel v1e + BS SW (Max FFT)


- #3 - Mel Deux + BS SW (Max FFT)

- the inst. sum of the SW stems, or inverted inst stem from vocals (MVSep, or av. vocals-only variant)

—-> to fix missing instruments in Deux

#4 - Mel Gabox Flowers v10 + Resurrection voc (Max FFT)

-—-> permits to fix missing instruments in Flowers (without adding vocal bleed - SW adds more)

- #5 - BS Resurrection Inst + SW (Max FFT)

- #6 (old) - Mel Gabox Flowers v10 + BS SW (Max FFT)

-—->Permits to fix missing instruments issue, without adding vocal bleed (but more than Resurrection voc)

- #7 - BS HyperACE v2 inst + SW (Max FFT)

To sum up - if there are missing instruments with deux (very rare), you can add SW model with Max Spec. If there is too much fuzzy sound, you can add Flowers (Avg) to add more fullness. Is there is vocal bleed, you can replace parts of the song where there is vocal bleed with Resurrection Inst or Flowers, two models which are good at not creating vocal bleed

___

Older ensembles

- Mel Deux + MVSep SCNet XL IHF high fullness by becruily on MVSep (Average Spec/Wave)

—-> The former best, but ultimately decided to be better than some below.

Better than the Flowers ensemble - fullness instrumental Roformers have kind of the same artifacts so mixing these is not so ideal, it’s noisy.
Permit a cross arch ensemble, as diversity is better, and most of the artifacts are gone because with avg (50/50), if an artifact is present at one moment in one model, but not the other, in the ensemble the artifact will be heard twice time lower. Ensembling between Roformers create too much redundancy, even if the datasets between for example HyperACE and deux are different, with different loss etc but still.
I think it's a better ensemble than the others.

- Leap Xe Inst + deux (Avg)

>Only an experimental since this model has great issues with BV popping

- Mel FlowersV10 with MVSep BS 2025.07 (Max FFT),
with phase fix using the BS model as a reference (low cutoff 100 and high cutoff 1000), and scale factor set to 3 (maximum) and use the ensemble as the target.

Step by step:

1. Do a Max FFT ensemble of FlowersV10 and BS 2025.07

And then

2. Phase fix that ensemble using BS 2025.07 again

Warning. For the phase fix don't use UVR (it has the scale factor setting absent), but this notebook instead.

—->Possibly the older best ensemble.
It's a really aggressive phase fix but it doesn't seem to harm the fullness as much as Deux does as a single model. The issue with the previous ensemble [below] is that it's noisy among other things. This new ensemble offers full enough results (not as full as v1e tho), no noise, good instrument protection (due to the help of 2025.07 model), and rare vocal leak. So it's good at everything, while using deux only can lead to some vocal bleed sometimes and not full enough some other times.

- Mel deux + inst_gaboxFlowersV10 (Avg)

-—->The former new best ensemble and currently the best for local use (public models) and since HyperACE is not supported by UVR but only MSST.

Deux is great at everything and its only downside would be it's not full enough for some users (can sound "fuzzy"), so Flowers permit to fix that issue. This ensemble doesn't add bleed because Flowers has even less vocal bleed than deux.

- Becruily 124 fullness phase fixed + SCNet XL IHF fullness (Avg): https://mvsep.com/presets?hash=QQoUWhQKDTtPhKDB (still requires premium)

- Becruily 124 fullness phase fixed + Deux phase fixed (Avg): https://mvsep.com/presets?hash=i1x5OagBY3ajVa8I

---->Some ideas of more ensembles for the 124 model.

- Becruily 124 fullness phase fixed + SW phase fixed (Max FFT)

https://mvsep.com/presets?hash=T4CsAFVKZKZR5VPx

- Becruily 124 fullness phase fixed + 2025.07 (Max FFT): https://mvsep.com/presets?hash=817GC54ngy7asaa3


- Mel deux + BS 2025.07 from MVSep (Max FFT/Spec)

—->To fix missing instruments in deux

- Mel deux + BS HyperACE v2 instrum (Max FFT)

—->Fixes “muffled” or “fuzzy” results with deux

(HyperACE V2 has less crossbleeding than V1e)

- Mel deux + HyperACE v2 instrum results (Max FFT/Spec)

Note: For deux, it will be not the muddier stem named _other in the Colab, but _instrumental
b) then phase fix it with the ensembled inst result with inst stem of Mel becruily vocals (a.k.a. high fullness on MVSep) or older alternative: Mel Kim (if you don't mind taking out some instruments with it) as reference,

with 200 for Low Cutoff and 1000 for High Cutoff,
c) Ensemble the outputs with BS 2025.07 from MVSep or BS-Roformer SW (inst. result) (Max FFT)

—-> To fix noise in deux + HyperACE V2 regular ensemble

(I just added BS 2025.07 because it finds more missing instruments without adding bleed

And I chose not to phase fix the final result (which include BS 2025.07) because otherwise "becruily vocal" would remove the BS 2025.07 difficult instruments. So doing it before is good)

- Phase fix with inverted Deux acapella stem as a reference/source (so the muddy instrumental) and Deux instrumental stem (the non-muddy) as a target

Settings:

a) aggressive: 200 for Low Cutoff and 1000 for High Cutoff (jarredou's recommendation)

b) balanced: 500/5000 (default values)

c) the least aggressive: 500/9000 (santilli/MJ's recommendation

—->To alleviate noise in deux (if it ever occurs at all)

- a) Ensemble Mel Deux + BS HyperACE V2 (Max FFT)

b) Phase fix it with becruily vocals (inst stem) as reference

c) Ensemble above result with BS 2025.07 from MVSep or BS-Roformer SW (inst. result) (Max FFT)

—->Alleviates the noise in the ensemble in a)

(even if "becruily vocals" is good at "protecting" instruments, it's not as good as "deux", for example with the Kazoo instrument, or others even. Indeed, becruily vocals is a quite old model now)

- BS-Roformer Resurrection Inst (a.k.a. Inst high fullness on MVSep) with BS 2025.07 (Max FFT)

—->Fixing possible vocal crossbleeding in instrumentals

(Resurrection inst is good at it)

- a) Phase fix Mel V1e by unwa with BS 2025.07 (inst. stem) as reference

b) Ensemble the above with BS 2025.07 (fixes missing instruments)

—-->For the fullest result with most of the noise

- Average (avg_wave) ensemble of:
Mel-RoFormer by Gabox FV10 (aka Flowers V10) + Mel-RoFormer deux by becruily + BS-RoFormer HyperACE v2 instrum, with weights (in order): 0.35.; 0.25; 0.40

—-->Fullness oriented, seems to yield good results

Or without weighting

--> No HyperACE variant to avoid mainly BV leak/complex vocals in instrumental (second best ensemble)

- BS HyperACEv2 + Mel deux + Mel FV10 + SCNet XL IHF high fullness (Avg)

--> Combination of BS, Mel & SCNet archs for great quality but the drawbacks are it adds noise and vocal residues + it increases vocal crossbleeding (because of SCNet XL IHF high fullness mainly) at times and it's a lot of models if processed in batch so not ideal

- BS HyperACEv2 + Mel deux + Mel FV10 (Avg)

--> Current best ensemble, combination of the Mel & BS arch's current best models with AVG instead of MAX to avoid adding too much noise. If there is vocal crossbleeding, consider removing HyperACE from the ensemble

- BS HyperACEv2 + Mel deux (Avg)

--> Permits to enjoy deux with a higher fullness as deux can sound fuzzy or unclear at times

- BS HyperACEv2 + Mel FV10 (Avg)

--> Combination of two very high fullness inst models to obtain results close to v1e

- BS HyperACEv2 + BS 2025.07 (Max)

--> HyperACE is not so good at keeping instruments compared to deux, so if you want to use only HyperACE as a main model, you can add 2025.07 to get the instruments back in the instrumental

- Mel FV10 + BS 2025.07 (Max)

--> Same reason as before, there is often instrument bleed in vocals with FV10

- BS Large Inst v2 + Mel FV10 (Max)

--> Ensemble proposed by neoculture, permits to retain some instrumental details masked by the noise in fullness inst models

- BS Large Inst v2 + BS HyperACEv2 (Max)

--> Just like before but HyperACE instead of FV10, as FV10 sounds pretty much as full as HyperACE so can be used too

Notes:

- Most of the time deux won't need any other model for ensemble or phase fixing unless rare crossbleeding, or noise occurs, or if it's not full enough, then you can use the above solutions

- FlowersV10 has lower SDR than "deux" and "hyperace v2 inst" and it has problems with didgeridoo instrument (vide "Supersonic" by Jamiroquai), but is mostly good at protecting them (worse than deux, better than hyperace).

- HyperACE v2 inst has big problems with the harmonica instrument (e.g. it almost picked up the entire harmonica part in "Take The Long Way Home" by Supertramp)

- "deux" is better at managing reverb vocal removal compared to other models.

- "deux" sometimes can leave complete lead vocal parts, as if the removal didn't work (for example "WOTW/POTP" by Coldplay or "Strangers" by Portishead).

- HyperACE is better than deux to "detach" some instruments that are "attached" to the vocals, like in "Mon imagination" by Pierpoljak (flute) or "I Missed Again" by Phil Collins (elec guitar) or even "La Force de Melodie" by Thievery Corporation (flute).

_________

Other/older ensembles

0) Mel deux + BS HyperACE v2 instrum (Max FFT) with phase fix using becruily vocals (inst. stem) as reference and 100/100 for the low cut and high cut values

—->Former best

0) BS-Roformer HyperACE Inst v2 + 2025.07 (Max FFT) using 100/100 as values for phase fixing and 2025.07 as reference

—->Former best

0) BS HyperACE + BS 2025.07 (Max FFT with BS 2025.07 as phase fix reference and 3000/5000 for the values)

—->My [former] favorite ensemble right now. Though, I notice it produces vocal bleed sometimes and it can be noisy at some parts of the songs while the noise might be totally absent in other parts of the song.

0) BS-Roformer Resurrection Inst (phase fixed with BS 2025.07 using Low Cutoff 3000 and High Cutoff 5000) + BS 2025.07 (Max Spec)

—->Former best

0) Unwa Mel inst v1e+ + FNO + Becruilly inst (Max Spec)

—-> The best result back then

- sakkuhantano

0) Unwa Mel inst V1e + MVSEP BS 2025.07 (Max Spec) using BS 2025.07 as phase fix reference with 100/100 (it's better) as Low Cutoff and High Cutoff values

—->It's a very aggressive value because V1e is noisy, and it works quite well.

The older best for then, “BS Roformer Resur Inst [ensemble right below] is muddy compared to v1e, and I think fullness is the way. After phase fix the noise is barely noticeable”

0) Phase fix of Unwa BS Roformer Resurrection Inst with BS 2025.07 as a reference, ensembled with MVSEP BS Roformer 2025.07 (Max Spec/FFT)

 —->The least vocal crossbleeding (step-by-step process explained here)

Alternatively, you can use becruily vocal model instead of 2025.07 for the ensemble -
“Becruily vocal correctly recognize instruments far better than the instrumental one”

(Note: BS 2025.07, BS 2024.04, BS 2024.08 and SW were worse for Resurrection model as a phase fix source, viperx BS-Roformer 1297 better, but not so good for instruments preservation as BS-2025.07)

0) unwa v1e + Mel becruily vocal (Max Spec) + phase fix (using becruily vocal again as a source)

 —->The best instruments preservation (with more possible crossbleed)

0) Mel Gabox Fv7z + BS 2025.07 (Max Spec)

—->The least amount of noise (with more possible crossbleed)

0) Mel Gabox Inst V8 + BS 2025.07 (Max Spec) + phase fix (becruily vocal as reference)

—->A good balance between presence of noise and level of fullness

(occasional vocal crossbleeding)

0) Mel Becruily Instrumental (with phase fix, becruily vocal as reference) + SCNet XL IHF (Max FFT)

—->Why SCNet? Because it's better than Mel Roformer at the low frequencies, so why not ensemble both arch. (...) SCNet is already noisy from the start so the fullness models are even noisier obviously

Even older ensembles

0a) Unwa v1e+ + BS-Roformer 12xx by viperx (Max Spec) - musicalman

0a) FNO inst by unwa + BS-Roformer 12xx by viperx (might be optional) + v1e+ (or becruily inst Mel-Roformer)

“beware of the song where there's a vocal at the beginning of the song, using v1e+ will leave vocal residue. So decide to change into becruily inst as well.” - Sakkuhantano

0*) Mesks’s metal min_fft ensemble of BS 07.2025 model on MVSEP + rifforge model
“it keeps the "fullness" of the rifforge model being an instrumental focused model but then also removes more stuff than my base model thanks to 07.2025”

0*) Chained separation method by fabio06844 for “very clean and full” instrumental.

1) Go to MVSep and separate your song with the latest Karaoke BS-Roformer by MVSep Team

2) On its instrumental stem result use DEBLEED-MelBand-Roformer (by unwa/97chris)

(model | yaml | Colab)

(despite the fact that “the MVSep Team Karaoke uses the MVSep BS model to extract/remove vocals, then applies [the] karaoke model to that”, it was told to be not enough to just use BS 2025.07 model instead, leaving a little more residues).

0b) v1e phase swapped from Becruily vocals + BS 2025.07 (Max Spec)
(Phase fixer Colab/or UVR’s phase swapper+MVSEP separation>UVR Manual Ensemble)

(“brings: max fullness without a lot of noise since phase fix, rarely missing instruments, no robotic voice problem, rare vocal crossbleeding in instrumental ”) - older favourite dca100fb8’s ensemble

0b) v1e + Becruily vocal (Max Spec) (“If you had to keep one ensemble right now. v1e+ unfortunately is muddier than v1e and has that robotic issue sometimes”) - dca100fb8

0b) v1e + Becruily inst + Becruily vocal (Max Spec) (Becruily inst turned out to be “useless” in this bag of models) - -||-

0b) Unwa v1e (with phase fix) + BS Large V1 (Max Spec) (doesn't need Becruily vocal here as the third) - -||-

0) v1e + INSTV7 (Max Spec) - neoculture

0) Use MVSEP’s SCNet XL high fullness below 1000 Hz, and unwa’s v1e above 1000 Hz, and join the two in e.g. Izotope RX - “You can use vertical select in RX with feathering set to 1.00” (heuhew) or you can use linear phase EQ like e.g. free “lkjb QRange” (ensure to not overlap frequencies in the output spectrogram in the crossover point)

0) v1e+ + becruily inst

—> fixes some missing instruments occasionally in v1e+

- Sakku

0b) INSTV7 + Inst_FV8 (to check)

0b) unwa’s instv1e+, instv1+, instv2 and inst gabox, instv8 and instv7 - max FFT

“Then, I upscaled using Apollo. Afterward, I applied [Mel-Roformer] de-noise to remove background noise as needed and performed mastering” (Sir Joseph)

0b) Max Spec manual ensemble of: v1e+ + MVSEP BS-Roformer 2025.07 model

- It’s a Senn’s method below, simplified (IntroC)

“I think (...) [it] would be good enough. The way v1e+ keeps the noise is usually fine, and the mvsep 2025.07 model should bring back the lost masked frequencies for v1e+. Otherwise, just adjust the weight of v1e+ for maxspec to reduce the noise”

0b) Max Spec manual ensemble of: v1e+ + MVSEP BS-Roformer 2025.07 model + Becruily inst

- Senn’s method simplified (IntroC/Sakku)

“sometimes becruily can catch a tiny instr while v1e+ can't” - Sakku

*) Senn’s OG method:

Use the highest SDR BS Roformer model on MVSEP and the best Fullness Melband Ro-Former model (unwa instrumental v1e plus) - mixed both with one of them phase-inverted, then use Soothe2 to filter out resonants, leaving mostly only noise, further filter and mix them, use a few plugins to do a spectral flattening, very minor.

In other words:

“Pass the music through BS RoFormer's best SDR and the best Fullness, which is Unwa Instrumental v1e plus

in iZotope RX, Invert the phase of any of the tracks, and copy n paste to the other track, the result should be some ghastly sounding reverb of the vocals

using iZotope RX's Deconstruct, you wanna filter out the tonal signal of the voice as to remove the more obvious "sinusoidal" signals. It has to be fairly subtle as to not damage the noisy residuals

Now with Soothe2, you wanna filter out any of the more aggressive noisy components, I use this setting, but it might not work 100% for everyone https://imgur.com/2A5yn3c (You can replace this with any plugin that acts similar to Soothe2, but Soothe2 is the best compromise)

If needed, you can use Deconstruct as well but reducing the noisy aspects just to wipe out that aggressive noisy artifact

Mess with the gain, and then add it back to the BS RoFormer track (don't forget to invert the phase again), Ideally it should be fairly subtle

For post-processing:

I use MSpectralDynamics to add a slight spectral flattening to the track, and Unchirp to denoise very slightly the higher frequencies to remove that digital hissy artifact and also tighten more of the sound

A very subtle Gullfoss can also brighten the track slightly as well to compensate

here is the result” (thx senn)

Older ensembles (from before becruily inst/voc Mel models release)

UVR>Audio Tools>Manual ensemble (for models from outside UVR)

0) Unwa’s v1e inst + phase fixer/swapper (from Mel-Kim or sometimes Mel 2024.10 on MVSEP for less noise) + BS 2024.08 on MVSEP (Max Spec)
(fullness with less noise + retrieved missing wind instruments from v1e) - dca100fb8
Becruily vocal is even better at recognizing instruments compared to Mel 2024.10 or BS 2024.08 (vocal Roformer models have fewer problems with recognizing instruments than inst Roformers)

0) Other dca’s Max Spec ensembles of v1e with other Roformer models


II) v1e + phase fixer/swapper + Mel 1143 (because of fullness with less noise + retrieved missing wind instruments from v1e)

III) v1e + phase fixer/swapper + BS 1296 (because of fullness with less noise + retrieved missing wind instruments from v1e)

IV) v1e + phase fixer/swapper + BS 1297 (because of fullness with less noise + retrieved missing wind instruments from v1e)
V) v1e + phase fixer/swapper + BS Large V1 (because of fullness with less noise + retrieved missing wind instruments from v1e)
Recommended (or with v2/v1 instead) esp. when using a phase fixer due to bleed of instruments in the vocal track in Kim or its fine-tunes used for the tool.

VI) (extra) for slow CPU/GPU: Voc FT + HQ5 (Max Spec)

VII) HQ_5 with UVR’s Phase Rotate (which now can replace the above)

VIII) v1e + BS 2025.06 (Max Spec) - manual ensemble - the latter on MVSEP (because it keeps instruments in instrumental correctly [though less than becruily vocal] and it has less vocal crossbleeding in instrumental compared to becruily vocal)

0) Unwa’s Inst V2 and Inst Gabox (Avg) (1120 segments [6GB GPU]/4 overlaps) - cypha_sarin

0) Max ensemble of: instv1, instv2 and inst v1e  - erdzo125
(better fullness than inst v1e itself, but more noise)

0b) Models ensembled - available only for premium users on mvsep.com

Now also added “instrumental high fullness” variant for inst, voc ensemble.

For example, some lower inst, voc SDR ensembles available might be less muddy than 11.50 (e.g. 10.44), but the 11.50 one has the fewer amounts of vocal residues according to bleedless metric, but it can also sound very filtered. Newer ensembles added since then (track the leaderboard and click on entries to see also bleedless/fullness metrics).
(ensembles on MVSEP provide currently the best of SDR scores for 2 and 4 stem separators, higher SDR than free v.2.4/2.5 Colabs below; 2025.06.28 has currently the biggest SDR metric, and surpassed ByteDance private model)

There are shorter queues for single model separation for registered users with at least one point.

Possibly shorter queues between 10:00 PM - 1:00 AM UTC.

The ensemble option fixes some issues with general muddiness of older vocal Roformer models (but 11.50 is muddier than v. 2.4 Colab).


0) KaraFan (e.g. preset 5; fork of original ZFTurbo's MDX23 fork with new features by Captain FLAM with jarredou's help on some tweaks), offline version, org. Colab and Kubinka Colab (older version, less vocal residues vs. v.3.1, although v.3.2-4.2/+ were released with fewer residues).

Used to be one of the best free solutions for instrumentals (before some newer Roformers like unwa’s inst v1 were released), with not big amounts of vocal residues (sometimes more than below), and clear outputs. But no 4 stems unlike below:

0a) MDX23 by ZFTurbo (weighted UVR/ZF/VPR models) -

free modified Colab fork v. 2.1 - 2.4, 2.5 with fixes and enhancements by jarredou

(one of the best SDR scores for publicly available 2-4 stem separator,
v2.2.2 Colab with fullband MDX23C model might have more residues in instrumentals vs v. 2.1, but better SDR, 2.7b (with SCNet XL, not SDR evaluated - weight set by ear), 2.3, 2.4 with also 12xx BS-Roformer, v. 2.5 with also Kim’s Mel-Roformer (default settings can be already good and balanced, and weights further adjusted, read for more settings.
The caveat - it haven’t been updated by newer Roformer instrumental models at the top).

0a) dango.ai (tuanziai.com/en-US) - 8$ every 10 songs, currently one of (if not) the best instrumental separator so far; at least till unwa inst models came to the level of fullness of Dango now (you might find the latter even too muddy), but in cost of more noise, although “it can hande complex vocals/songs well so it's more reliable, and no vocal bleed in background of instrumental”.
Dango’s 10 Conservative mode give more fullness to instrumentals in cost of whispering artefacts (experimental for the time being - along with the Aggressive mode), and it doesn’t fix vocal popping using Smart Mode (default). Now Dango 11 is available. More crossbleeding than Unwa Bs-Roformer inst. What changed for better is less noise and better instrument detection


Ensembles (from before becruily models release; less noisy single models later below)

0b) Avg Spec ensemble of unwa inst v1 and v2

0b) Min Spec manual ensemble of vocals stems from these models>inversion with the original song (fuller, more noise)  
(UVR>Audio Tools>Manual Ensemble)

0b) Max ensemble of: unwa’s v1e + Mel 2024.10 + BS 2024.08 - dca100fb8
(older bag of models; 2024.08 on MVSEP; fixes the flute and trumpet issues)

0b) Separate the song twice - first with v1e, then with Unwa’s BS-Roformer Large and do a Manual Max Spec Ensemble via UVR - dca100fb8
(old; BS-Roformer is here to retrieve the missing instruments from v1e result, though BS 2024.08 & Mel 2024.10 on MVSEP work better for this task already)

(Ensembles from before unwa inst. models release - you might also try out replacing all the 12xx BS-Roformers below with newer unwa’s/becruily/Gabox models)

0b) Models ensembled - available only for premium users on x–minus.pro

- max_mag ensemble (with viperx 1297 Roformer)

- demudder (on Mel-Roformer)

- Mel-Roformer + MDX23C

UVR 5 ensembles (although beta 4 and inst v1 on their own might be better already)

For Roformers, min. RTX 3050 8GB or faster AMD/Intel ARC/Apple M1-3 recommended

(OpenCL is not as fast as CUDA in UVR; 6GB VRAM on CUDA should be enough too, min. 2K+ CUDA cores recommended)

0b) 1296 + 1143 (BS-Roformer in beta UVR) + MDX-Net HQ_4 (dopfunk)

[potentially try out Mel Kim instead of 1143 above already]

0b) Manual ensemble (in UVR’s Audio Tools) of:

BS-Roformer 1296 + file copy of the result + MDX23C HQ (jarredou; src)

or just 1296 + 1297 + MDX23C HQ for slower separation and similar result

0b) Manual ensemble of:

- BS-Roformer 1296 + drums stem from demucs_ft or

- Bs-Roformer 1143 result passed through demucs_ft for drums to ensemble with 1296 (max/max)

0b) MDXv2 HQ_4 + BS-Roformer 1296 + BS-Roformer 1297 + Melband RoFormer 1143 (Max Spec) “Godsend ensemble for demuddiness” (dca100fb8)

0b) Manual ensemble of HQ_5 (paid users) and Kim's Mel-Roformer (max_spec)

0b)  (for metal) “1 – pass through Kim's Vocal Melband Roformer (link in Single models below)

2 - Multi-stem Ensemble (Average algo):

1_HP-UVR

MGM_HIGHEND_v4

MGM_LOWEND_A_v4

(VR advanced settings: 320 window size, 5 aggression setting, batch size default, TTA enabled | Post Process and High-End Process CHECKED OFF)”

3 – Manual Ensemble both your Melband output and the Multi-stem instrumental output (with Average algorithm)

“best settings/models for metal” (~mesk)

0b) (older version of the above) 1_HP_UVR + UVR_MDX-NET-Inst HQ 4 + UVR_MDX-NET-Inst_Main 438 (VIP model)

(Min Spec / Average, WS 512, TTA Enabled, Post-Process and High-End Process off)

0b) 9_HP2-UVR and Kim Mel-Roformer (newer one for metal; mesk)

but not in multi-stem cos you need 3 or more models

VR: 320 window size, 1 aggression setting, Default batch size, TTA enabled (post process and high-end process isn't enabled)

0b) Mateus Contini's method e.g. #2 or #4

0b) MDX-Net Kim Vocals 1

MDX-Net MDX23C-InstVoc D1581

MDX-Net MDX23C-InstVocHQ 2

MDX-NET UVR-MDX-NET-Voc_FT

Demucs v4 | htdemucs_ft

avg/avg>"Vocal Splitter Options" and choose "VR Arc: 5_HP-Karaoke-UVR"

0b) 9_HP2-UVR + BS-Roformer 1297

0b) BS-Roformer ver. 2024.08 + MelBand Roforrmer (Bas Curtiz edition) + MDX-Net HQ4 + SCNet Large, Max Spec Ensemble (dca100fb8)

0b) BS-Roformer ver. 2024.08 + MelBand Roforrmer (Bas Curtiz edition) (Max Spec Ensemble) --> result. Result + MDX-Net HQ4 + SCNet Large (Average Ensemble) - -||-

0b) 1297 (ev. 1296) + MDX23C HQ2 (CZ-84)

[or potentially unwa’s BS-Roformer instead of 12xx]

See also DAW ensemble (older ensembles later below)

(more about) unwa’s instrumental Mel-Roformer v1 model | MVSEP | x-minus.pro

https://huggingface.co/pcunwa/Mel-Band-Roformer-Inst/tree/main | Colab | UVR instructions

"much less muddy (..) but carries the exact same UVR noise from the [MDX-Net v2] models"

But it's a different type of noise, so aufr33 denoiser won't work on it.

“you can "remove" [the] noise with uvr denoise aggr -10 or 0” although with -10 it will make it sound more muddy like Kim model and synths and bass are sometimes removed with the denoiser (~becruily)

“ if there is any voice left [or also background noise], use the Mel-Roformer de-noise with minimal aggression.

This inst model “doesn't eliminate vocoder voices well from an instrumental”.

For the noise in the model, vs the ensemble trick on x-minus using Mel-Roformer de-noise might be better alternative:

“removes more noise from the song keeping overall instrument quality more than the new button [on x-minus]” koseidon72. But the more aggressive variant of the Mel model sometimes deletes parts of the mix, like snares. UVR-Denoise-Lite doesn’t seem to damage instruments like non-lite UVR-Denoise in UVR, but still more than Mel denoise (recommended aggr. - 4, with 272 vs 512 windows size it’s less muddy, TTA can stress the noise more, somewhere above 10 aggr. it gets too muddy). UVR-Denoise on x-minus is even less aggressive (it’s medium aggression model for free users who don’t have aggression pick), but it might catch ends of some instruments like bass occasionally. Premium minimum aggression model is somehow more muddy, but doesn’t damage instruments.

For more muddy Roformers consider using Aufr’s demudder (it’s used for premium on x-minus for Kim Mel model) although it might increase vocal residues, and UVR demudder (explained there later below).

____

Muddier but cleaner single models (Roformer vocal models with fewer instrumental residues vs instrumental models without necessity of using phase fixer)

The list lacks some newer vocal models tested, so free to test some newer ones here as well

0c) MVSep BS-Roformer (2025.07.20) vocal model

Inst. fullness 27.83, bleedless 49.12, inst SDR 18.20

Probably a retrain of the SW model on a bigger dataset.

0c) BS-RoFormer SW 6 stem (MVSEP/Colab/undef13 splifft) / vocals only

Inst. fullness 27.45, bleedless 47.41, inst SDR 17.67

(use inversion from vocals and not mixed stems for better instrumental metrics)

Known for being good on some songs previously giving bad results.

0c) MVSep Mel-Roformer 10.2024 vocal model
Inst. fullness 27.84, bleedless 47.37, inst SDR 17.59

The cleanest, but muddy compared to models trained for instrumentals

Capable of detecting sax and trumpet, but still muddier than instrumental models above.
Bas Curtiz vocal model fine-tuned by ZFTurbo.

0c) Gabox voc_fv4 | yaml | Colab

Good for anime and RVC purposes (codename)

And also for instrumentals, if you need less vocal residues than typical instrumental Roformers (even less than Mel Kim, FT2 Bleedless, or Beta 6X - makidanyee).

0c) Unwa Beta 6X

Some people might even favored it over Resurrection inst - cyclorana

“the best universal vocal/instrumental model I’ve ever seen TBH” - NateTheGrate

0c) Unwa’s Kim Mel-Band Roformer FT2 | download | Colab

Inst. fullness: 28.36, bleedless: 45.58

A decent all-rounder too, sometimes less bleeding in instrumentals than 5e, although a bit worse transients.

0c) Unwa’s beta 5e model originally dedicated for vocals | Colab | MSST-GUI | UVR instr
Model files | yaml: big_beta5e.yaml or fixed for AttributeError in UVR
Inst. fullness: 27.63 (bigger than Mel-Kim) | bleedless 45.90 (bigger than Kim FT by unwa, worse than Mel-Kim) | Inst. SDR 16.89

Mainly for vocals, but can be still a decent all-rounder deprived of noise present in unwa’s inst v1e/v1/v2 models, also with fewer residues than in Kim FT by unwa, and also more consistent model than Kim Mel model in not muffling instrumental a bit in sudden moments.
The third highest bleedless instrumental metric after Mel-Kim model (after unwa ft2 bleedless in vocals).

It seems to fix some issues with trumpets in vocal stem (maxi74x1).
It handles reverb tails much better (jarredou/Rage123).

Noisier/grainier than beta 4 (a bit similarly to Apollo lew's vocal enhancer), but less muddy.
“The noise is terrible when that model is used for very intense songs” - unwa

Phase fixer for v1 inst model doesn’t help with the noise here (becruily).
“It's a miracle LMAO, slow instrumentation like violin, piano, not too many drums...
it's perfect... but unfortunately it can't process Pop or Rock correctly” gilliaan
“the vocal stem of beta5e may have fullness and noise level like duality v1, but it may also suffer kind of robotic phase distortion, yet may also remove some kind of bleed present in other melrofo's.” Alisa/makidanyee
“particularly helpful when you invert an instrumental and then process the track with it.” gilliaan

0c) Unwa Kim Mel-Band Roformer Bleedless FT2 | download | Colab

0c) Bas Curtiz' edition Mel-Roformer vocal model on MVSEP

(it was trained also on ZFTurbo dataset)

“Music sounds fuller than original Kim's one & the finetuned version from ZFTurbo [iirc below]. Even [though] the SDR is smaller than BS Roformer finetuned last version, but almost song has the best result in instrumental.” Henri

It can struggle with trumpets more than the other Mel-Roformer on MVSEP [whether 08.2024 or Mel-Kim, can’t remember].

0c) BS-Roformer 2024.08.07 vocal model on MVSEP

Inst. fullness 26.56 (less than Mel-Kim), Inst. bleedless 47.48 (the only single model with that better metric than Mel-Kim)

Inst SDR 17.62

vs 2024.04 model +0.1 SDR and “it seems to be much better at taking the vocals when there are a lot of vocal harmonies” also good for Dolby channels.

Capable of detecting flute correctly

0c) Mel-Roformer vocal model by KimberleyJSN - model | config | Colab

Inst. fullness 27.44 (worse than beta 5e and duality, but better than current BS-Roformers)

Inst. bleedless 46.56 (the best metric from public models)

Inst SDR 17.32

It became a base for many Mel-Roformer fine-tunes here.

(works in UVR beta Roformer/Colab/CML inference/x-minus/MDX23 2.5 (when weight is set only for Mel model)/simple model Colab (might have problems with mp3 files)

It’s less muddy than older viperx’ Roformer model, but can have more vocal residues e.g. in silent parts of instrumentals, plus, it can be more problematic with wind instruments putting them in vocals, and it might leave more instrumental residues in vocals. SDR is higher than viperx model (UVR/MVSEP) but lower than fine-tuned 2024.04 model on MVSEP.

0c) Unwa Revive 2 BS-Roformer (“my first impression is it may have less low end noise than fv4 but not the best in the overall quality and amount of residues in vocal” - makidanyee)


0c) BS-Roformer Large vocal model by unwa (viperx 1297 model fine-tune) download
Older BS model. It picks more instruments than 12xx models. More muddy than Kim’s Roformer, a bit less of vocal residues, a bit more artificial sound. Also tends to be more muddy than viperx 1297, sometimes muffling instrumental at times, but a bit less of vocal residues, a bit more artificial sound/a bit less musical. Sometimes it has more vocal residues than beta 5e.

Compared to BS, Mel-Roformers can be a good balance between muddiness and clarity for some instrumentals.

Compared to ZFTurbo (MVSEP) and viperx models, Kim’s trained on Aufr33’s and Anjok’s dataset.

UVR manual model installation (Model install option added in newer patches):
Place the model file to Ultimate Vocal Remover\models\MDX_Net_Models and the config to model_data\mdx_c_configs subfolder and “when it will ask you for the unrecognised model when you run it for the first time, you'll get some box that you'll need to tick "Roformer model" and choose its yaml” some models here are available in Download Center too.

Other unwa fine-tunes (originally vocal models)

0c) Mel-Roformer Kim | FT (by unwa) | Colab

https://huggingface.co/pcunwa/Kim-Mel-Band-Roformer-FT/tree/main

Inst. fullness 29.18 (lower than only unwa inst models)

Inst. bleedless 45.36 (lower than Beta 5e)

Inst. SDR 17.32

Has more vocal residues than Beta 5e

- Aname Mel-Roformer duality model .
It’s focused more on bleedless than fullness metric contrary to the unwa’s duality v2 model, but with bigger SDR.

Inst. fullness 24.36, bleedless 46.52, SDR: 17.15

- Mel-Roformer unwa’s inst-voc model called “duality v1/2” (focused on both instrumental and vocal stem during training; two independent and not inversible stems inside one weight file).

https://huggingface.co/pcunwa/Mel-Band-Roformer-InstVoc-Duality | Colab | MVSEP
V1: Inst fullness 28.03, bleedless 44.16, SDR 16.69.

V2: Inst SDR 16.67

Outperformed in both metrics by the unwa’s Kim FT.

Vocals sound similar to beta 4 model, instrumentals are deprived of the noise present in inst v1/e models, but in result, they don't sound similarly muddy to previous Roformers.

Compared to beta 4 and BS-Roformer Large or other archs’ models, it has fewer problems with reverb residues, and vs v1e, with vocal residues in e.g. Suno AI songs.
"other" is output from model, "Instrumental" is inverted vocals against input audio.

The latter has lower SDR and more holes in the spectrum, using MSST-GUI, leave the checkbox “extract instrumental” disabled for duality models (now it’s also in the Colab with “extract_instrumental” option) and probably for inst vx models.
You can use it in the Bas Curtiz’ GUI for ZFTurbo script or with the OG ZF’s repo code.

- unwa’s Mel-Roformer fine-tuned beta 3 (based on Kim’s model)

https://huggingface.co/pcunwa/Mel-Band-Roformer-big/tree/main | Colab
Inst SDR: 17.30

Since beta 3 there’s no ringing issues in higher frequencies like in previous betas.
Sometimes better for instrumentals than beta 4 - but tends to be too muddy at times, but with fewer vocal residues than beta 5.

- unwa’s Mel-Roformer beta 4 (Kim’s model fine-tuned)

https://huggingface.co/pcunwa/Mel-Band-Roformer-big/tree/main | Colab

Outperformed in both metrics by beta 5e.

Be aware that the yaml config is different in this model.

“Metrics on my test dataset have improved over beta3, but are probably not accurate due to the small test dataset. (...) The high frequencies of vocals are now extracted more aggressively. However, leakage may have increased.” - unwa

“one of the best at isolating most vocals with very little vocal bleed and still doesn't sound muddy” Can be a better choice on its own than some ensembles.

0c) SCNet XL (vocals, instum)

Inst SDR: 17.2785

Vocals have similar SDR to viperx 1297 model,

and instrumental has a tiny bit worse score vs Mel-Kim model.

0c) Older SCNet Large vocal model on MVSEP

“just like the new BS-Roformer ft model, but with more bleed. [BS] catches vocals with more harmonies/bgv” - isling. “it's like improved HQ4” - dca100fb8
Issues with horizontal lines on spectrogram.

0d) Aname Mel model trained from scratch a.k.a. Full Scratch

Inst. fullness: 25.10, bleedless: 37.13

Models for older archs (outdated models)

0c) MDX23C 1666 model exclusively on mvsep.com

(vocal Roformers are much more muddy than MDX23C/MDX-Net in general, but can be cleaner)

0c) MDX23C 1648 model in UVR 5 GUI (a.k.a. MDX23C-InstVoc HQ / 8K FFT) and mvsep.com, also on x-minus.pro/uvronline.app

Both sometimes have more bleeding vs MDX-Net HQ_3, but also less muddiness.

Possible horizontal lines/resonances in the output - fix DC offset and or use overlap “starting from 7 and going multiples up - 14 and so on.” Artim Lusis

0c) MDX23C-InstVoc HQ 2 - VIP model for UVR 5. It's a slightly fine-tuned version of MDX23C-InstVoc HQ. “The SDR is a tiny bit lower, but I found that it leaves less vocal bleeding.” ~Anjok

It’s not always the case, sometimes it can be even the opposite, but as always, all may depend on a specific song.

0d) MDX-Net HQ_4/3/2 (UVR/MVSEP/x-minus/Colab/alt) - small amounts of vocal residues at times, while not muffling the sound too much like in old BS-Roformer v2 (2024.02) on MVSEP, although it still can be muddy at times (esp. vs MDX23C HQ models), HQ_4 tends to be the least muddy out of all HQ_X models (although not always), and is faster than HQ_3 and below, it tends to have less vocal residues vs MDX23C.

Final MDX-Net HQ_5 seems to be muddier for instrumentals, although slightly less noisy, but better for vocals than HQ_4.

0d) MDX HQ_5 final model in UVR (available in its Download center and Colab)
Versus HQ_4, less vocal residues, but also muddier at times and a bit lower, 21,5kHz cutoff.
Sometimes even more muddy than narrowband inst 3 to the point it can spoil some hi hats occasionally.
Versus unwa’s v1e “HQ5 has less bleed but is prone to dips in certain situations. (...) Unwa has more stability, but the faint bleed is more audible. So I'd say it's situational. Use both. (...) Splice the two into one track depending on which part works better in whichever part of the song is what I'd do.” CC Karaoke

Model | config: "compensate": 1.010, "mdx_dim_f_set": 2560, "mdx_dim_t_set": 8, "mdx_n_fft_scale_set": 5120

0d) MDX HQ5 beta model on  uvronline via special link for: free/premium (scroll down)

Go to "music and vocals" and there you will see it

It's not a final model yet, the model was in training since April.

It seems to be muddier than HQ_4 (and more than Kim’s and MVSEP’s Mel-Roformer), it has less vocal bleeding than before, but more than Kim Mel-Roformer.

"Almost perfectly placed all the guitar in the vocal stem" it might get potentially fixed in the final version of the model.

0e) Other single MDX23C full band models on mvsep.com (queues for free unregistered users can be long)

(SDR is better when three or more of these models are ensembled on MVSEP; alternatively in UVR 5 GUI’s via “manual ensemble” of single models (worse SDR) or at best, weighted manually e.g. in DAW, but the MVSEP “ensemble” option is specific method - not all fullband MDX23C models on MVSEP, that’s including 04.24 BS-Roformer model are available in UVR)

- BS-Roformer model ver. 2024.04.04 on MVSEP (further trained from viperx’ checkpoint on a different dataset). SDR vocals: 11.24, instrumental: 17.55 (vs 17.17 in the base viperx model). Bad on sax. Less muddy than the three below.

Though, all might share same advantages and problems (filtered results, muddiness, but the least of residues)

- Mel-Roformer model ver. 2024.08.15 on MVSEP (fine-tuned on prob. Kim’s model)

- BS-Roformer 12xx models by viperx model in UVR beta/MVSEP and x-minus (struggles with saxophone too, but less (also vs Gabox inst v6), also struggles with some Arabic guitars, bad on vocoders)

“does NOT pick up on large screams that much (example being Shed by Meshuggah in my tests), well at least [vs] [kim’s] x-minus mel-rofo”

1297 variant is being used on x-minus. It tends to be better for instrumentals than the 1296 model.

- Older BS-Roformer v2 model on MVSEP (2024.02) (a bit lower SDR)

All vocal Roformer models may sound clean, but filtered at the same time - a bit artificial [it tends to be characteristic of the arch], but great for instrumentals with heavy compressed vocals and no bass and drums - the least amount of residues and noise - very aggressive.

- old MelBand Roformer model on MVSEP (don’t confuse with the Kim’s one x-minus - they’re different)

- GSEP (now paid) -

Inst fullness: 28.83, bleedless: 31.18, SDR: 12.59

Check out also 4-6 stem separation option and perform mixdown for instrumental manually, as it can contain less noise/residues vs 2 stem in light mix without bass and drums too (although more than first vocal fine-tunes of like MVSEP’s BS-Roformer v2 back then). Regular 2 stem option can be good for e.g. hip-hop, and 4/+ stems a bit too filtered for instrumentals with busy mix. GSEP tends to preserve flute or similar instruments better than some Roformers and HQ_X above (for this use cases, check out also kim inst and inst 3 models in UVR) and is not so aggressive in taking out vocal chops and loops from hip-hop beats. Sometimes might be good or even the best for instrumentals of more lo-fi hip-hop of the pre 2000s era, e.g. where vocals are not so bright but even still compressed/heavily processed/loud or when instrumental sound more specific to that era. For newer stuff from ~2014 onward, it produces vocal bleeding in instrumentals much sooner than the above models. "gsep loves to show off with loud synths and orchestra elements, every other mdx v2/demucs model fail with those types of things".

Older ensembles (among others from the leaderboard)

Q: How to ensemble BS-Roformer 1296 with Kim Mel-Roformer using UVR GUI?

I choose max/max vocal/instrumental, but on the list there is only 1296, and no Kim Mel-Roformer like in MDX-Net option [might have been fixed already]

A: “You have to set the stem pair to multi-stem ensemble, it can generate both vocal and instrumental from both models at the same time. Be sure to set the algorithm to max/max. Once that's done, find the ensemble folder and put the two instrumental files/two vocal files onto the input, provided that you have to go to audio tools first. Then set the algorithm to average and click on the start processing button” - imogen

0f. #4626:

MDX23C_D1581 + Voc FT

0g) #4595:

MDX23C_D1581 + HQ_3 (or HQ_4 now)

0h) Kim Vocal 2 + Kim Inst (a.k.a. Kim FT/other) + Inst Main + 406 + 427 + htdemucs_ft (avg/avg)

0i) Voc FT, inst HQ3, and Kim Inst

0j) Kim Inst + Kim Vocal 1 + Kim Vocal 2 + HQ 3 + voc_ft + htdemucs ft (avg/avg).

0k) MDX23C InstVoc HQ + MDX23C InstVoc HQ 2 + MDX23C InstVoc D1581 + UVR-MDX-NET-Inst HQ 3 (or HQ 4)

“A lot of that guitar/bass/drum/etc reverb ends up being preserved with Max Spec [in this ensemble]. The drawback is possible vocal bleed.” ~Anjok

0l) MDX23C InstVoc HQ + MDX23C InstVoc HQ 2 + UVR-MDX-Net Inst Main (496) + UVR-MDX-Net HQ 1

"This ensemble with Avg/Avg seems good to keep the instruments which are counted as vocals by other MDXv2/Demucs/VR models in the instrumental (like saxophone, harmonica) [but not flute in every case]" ~dca100fb8

0m) MDX23C InstVoc HQ + HQ4

0n) Ripple (no longer works) / Capcut.cn (uses SAMI-ByteDance a.k.a. BS-Roformer arch) - Ripple is for iOS 14.1 and US region set only - despite high SDR, it's better for vocals than instrumentals which are not so good due to noise in other stem (can be alleviated by decreasing volume by -3dB).

0n) Capcut (for Windows) allows separation only for the Chinese version above (and returns stems in worse quality). See more for a workaround. Sadly, it normalizes input already, so -3dB trick won’t work in Capcut. Also, it has worse quality than Ripple

The best single MDX-UVR non-Roformer models for instrumentals explained in more detail

(UVR 5 GUI/Colabs/MVSEP/x-minus):

0. full band MDX-Net HQ_4 - faster, and an improvement over HQ_3 (it was trained for epoch 1149). In rare cases there’s more vocal bleeding vs HQ_3 (sometimes “at points where only the vocal part starts without music then you can hear vocal residue, when the music starts then the voice disappears altogether”). Also, it can leave some vocal residues in fadeouts. More often instrumental bleeding in vocals, but the model is made mainly for instrumentals (like HQ_3 in general)

0b) full band MDX-Net HQ_5 - similarly fast, might be less noisy, but more muddy, although better for vocals, but “it seems it's the best workaround when there is vocal bleed caused by Roformers”

1. full band MDX-Net HQ_3 - like above, might be sometimes simply the best, pretty aggressive as for instrumental model, but still leaving small amounts of vocal residues at times - but not like BS-Roformer v2/viperx, so results are not so filtered like in these.

HQ_3 filters out flute into vocals. Can be still useful to this day for specific use cases “the only model that kept some gated FX vocals I wanted to keep”.

It all depends on a song, what’s the best - e.g. the one below might give better clarity:

2. full band MDX23C-InstVoc HQ (since UVR 5.60; 22kHz/fullband as well) - tends to have more vocal residues in instrumentals, but can give the best results for a lot of songs.

Added also in MDX23 2.2.2 Colab, possibly when weights include only that model, but UVR's implementation might be more correct for only that single model. Available also in KaraFan so it can be used there only as a solo model.

2b. MDX23C-InstVoc HQ 2 - worse SDR, sometimes less vocal residues

Older MDX models

2c. narrowband MDX23C_D1581 (model_2_stem_061321, 14.7kHz) - better SDR vs HQ_3 and voc_ft (single model file download [just for archiving purposes])

"really good, but (...) it filters some string and electric guitar sounds into the vocals output" also has more vocal residues vs HQ_3.

*. narrowband Kim inst (a.k.a. “ft other”, 17.7kHz) - for the least vocal residues than both above in some cases, and sometimes even vs HQ_3

*. narrowband inst 3 - similar results, a bit more muddy results, but also a bit more balanced in some cases

- Gabox “small” inst Mel Roformer model for faster inference than most Roformers | yaml
Be aware that it can have some audible faint constant residues.

*. narrowband inst 1 (418) - might preserve hihats a bit better than in inst 3.

3. narrowband voc_ft - sometimes can give better results with more clarity than even HQ_3 and kim inst for instrumentals, but it can produce more vocal residues, as it’s typically a vocal model and that’s how these models behave in MDX-Net v2 arch (you can use it e.g. as input for Matchering for cleaner, but more muddy model result)

*. less often - inst main (496) [less aggressive vs inst3, but gives more vocal residues]

*. or eventually also try out HQ_1 - (epoch 450)/HQ_2 (epoch 498) or earlier 403, 338 epochs, or even 292 is also used frequently from time to time) when VIP code is used.

Recommended MDX and Demucs parameters in UVR

- Ensemble of only models without bleeding in single models results for specific song

- DAW ensemble of various separation models - import the results of the best models into DAW session set custom weights by changing their volume proportions

- Captain Curvy method:

"I just usually get the instrumentals [with MDX23C] to phase invert with the original song, and later [I] clean up [the result using] with voc ft"

How to check whether a model in UVR5 GUI is vocal or instrumental?

(although in MDX23C there is no clear boundary in that regard)


> for vocals
(voc. ensembles, or click here for Karaoke, here for instrumentals)

MVSEP models without download links can be used only on MVSEP

(removing/isolating vocals from AI music can give muddy results and capture other unrelated instruments easily; also, Roformers tend to stress plosives which weren’t in the original vocals at time - cristouk)

Tips:

1. “I get less distortion on vocals if I delimited first on a loud mix” - 5b

2. Use the output file to separate with a Phantom Center model.

3. Then use the two output files to separate with a vocal model below one by one, then mix them together. You can get better results - gilliaan/straturkoise
4*. Or ensemble every of the two phantom center results with the vocal model or ensemble (e.g. big beta 7 and bs_roformer_mag [Max Spec/FFT]), and mix the results afterwards e.g. in Audacity - neoculture


There’s no one, the best model. It depends on the song.
Most commonly used/or the best models for the doc’s date (categorized below with links and metrics):

- MVSep BS-Roformer 124 bands (premium only, the best SDR)

MVSep BS-PolarFormer 124 bands (the 2nd best)
MVSep Becruily BS-Roformer 124 bands (fine-tune of the first one, but less bleedless)

- MVSep BS-Roformer 2025.07 (predominantly bleedless and accurate, muddy)

BS-Roformer SW Vocals-Only (bleedless, a bit similar, FT-ed above)

- Becruily Deux (fullness, but still a bit similar to the PolarFormer),

Anvuew BS_RoFormer_mag (fullness with more bleedless; very noisy at quiet parts),

BS_Roformer_mag_v2 (worse SDR, bigger fullness, you might like it more),

- unwa BS-Roformer-Leap Xe (best metrics for public model, e.g. the non Xe more highs, less mids vs Deux)

- HyperACE voc v2 (bleedless, with more fullness),

Resurrection voc (less bleedless, with even more fullness),

- Big Beta 6X (bleedless, with more fullness), 

Big Beta 7 (more bleedless, vs 12.55, SW),

Revive3e (more fullness, with less bleedless, struggles with harmonies, noisy)

- Outperformed metrically, but good at times:

Gabox voc_fv7 (bleedless, much less harmony issues vs 6X),

vocfv7 beta 2 (middle-ground fullness, less crossbleeding, fuller, noisier),

vocfv7 beta 3 (fuller; possible drum bleed in all betas)
vocfv7 beta 1 (better bleedless than voc_fv4, better BVs than 5e)

- Others/olders: BS-Roformer 12.45, voc_fv4 (might struggle with BVs vs 6X), fv5 (some might like it more vs 6X at times)
Big Beta 5e, Mel FT2 Bleedless (even more bleedless), voc_fv6, Becruily voc, Revive 2, FT3 Preview.

Full list of models


Bleedless models #1 (MVSEP exclusive; models available for download later below)

- MVSep BS Roformer 124 bands (currently only for premium/paid members of MVSep.com)

Vocals bleedless: 39.18, fullness: 18.11, SDR: 12.33' 

Metrics with 3 fullness levels | Model option link

“Because it is several times slower than the original BS Roformer (which is one of the most used models on the site), we have temporarily placed it behind a paywall to ensure our servers can handle the load. We plan to make it freely available later once we add more servers. The model can provide several levels of fullness for vocals and instrumentals out of the box (you need to choose it in the form).” - ZFTurbo

“+0.4 SDR is not a small jump, I mostly check the spectrograms and frequencies appear quite more defined.

So more accurate in preserving the correct instruments overall, sound wise mostly similar to everything else non-fullness (...)

There might be something wrong with the fullness inst output, it appears to be muddier than [the] baseline.” - becruily

“So looking at this, Roformer level 1 fullness has about the same bleedless as PolarFormer (non fullness), but more fullness

and the PolarFormer fullness model is basically irrelevant now because of the extremely low bleedless score is compared to the fullness tradeoff (...) Metrics wise it looks like there is no good reason to use level 1 for instrumentals we gain 2 points of fullness and lose like 8 bleedless - dynamic64

Q: Can the 124 BS-Roformer model make insts that have the fullness of deux/hyperace? I'm trying to avoid most of the muddiness. I don't have premium atm or I would just test myself lol

A: “No, not even close metrics wise tbh” - dynamic64

- MVSep BS PolarFormer model (124 bands) - only on MVSep (longer queue w/o premium/account)

Vocals bleedless: 39.26, fullness: 17.18, SDR: 12.02

Model option link | Metrics / fullness level 1 | Description
“Can be lacking some fullness, it's tuned for low noise.” - rainboomdash

“Very good for lead vocals. Unfortunately, it doesn’t sound good on backing vocals or choir vocals” - myself044. Sometimes deux, sometimes this picks up BVs better - pezz23

“On the contrary for me it works best to remove all vocals” - duhh_its.chris

So it might be song-dependent.
“there's not a huge dramatic difference between it and [deux] (...) There is definitely a difference, but you probably wouldn't notice it based on only one listen if you were directly comparing them (...) results sound realistic” - pezz23


“It currently achieves the highest SDR on the MultiSong dataset. It also introduces a new feature: generating two additional stems with higher fullness.” - ZFTurbo

“This model can extract Daft Punk vocoder cleanly” - RME

“the inst will always remain muddy no matter which fullness level it will be set to, since it's not the target instrument.

It sounds like phase remix demudder output, also what makes me think it’s a vocal model, is just how the inst sounds” - dca100fb8 (might be the case since the public PolarFormer is vocal target, indeed - squiddagod).

- MVSep BS-Roformer 2025.07 (only on MVSep)

Vocals bleedless: 38.25, fullness: 17.23, SDR: 11.89
The biggest bleedless metric for a single model so far. Compared to previous models, picks up backing vocals and vocal chops greatly where 6X struggles, and fixes crossbleeding and reverbs where in some songs previous models struggled before.
Sometimes you might still get better results with Beta 6X or voc_fv4 (depending on a song).

“07.2025 is by FAR the most accurate model i've tried, for instrumentals and vocals, but the mudiness kills me”

“Very similar to SCNet, very high fullness without the crazy noise” - dynamic64,

“handles speech very well. Most models get confused by stuff like birds chirping (they put it in the vocal stem), but this model keeps them out of the vocal stem way more than most. I love it!”

Works the best for orchestral choirs out of the long list of other models - .elgiano.

It can be better for metal both for vocal and instrumentals than the mesk’s models, a lot of the time (and sometimes the best).

The “only one i've found rivals that one in recognizing hard vocals is SW” - cyclorana

It’s the SW fine-tune: “it has the same params as SW, outputs are similar to [the] SW vocal-wise, but better obviously” - mesk

The first iteration of the model (2025.06: 37.83/17.30/11.82) received two small updates and was replaced by 2025.07.

It’s also good as a source model for phase fixer/swapper.

- Mel-Roformer Bas Curtiz edition (/w Marekkon5) (trained on also ZFTurbo dataset) only on MVSEP (older version of 2024.10 model)

Vocals bleedless: 39.20, fullness: 16.24, SDR 11.18.

Bleedless models #2


- unwa BS-Roformer-Leap vocal
Vocals bleedless: 38.90, fullness: 17.07, SDR: 11.72
MK Colab fork #2 | metrics (the highest vocal SDR for public model, lower than MVSep PolarFormer 62 bands)
“feels a bit fuller than [beta] resurrection v2, and it does a really good job picking up background vocals” - neoculture

“it sounds like there's a ton of fullness in high freq but more filtering in mid freqs. Imo the signature is kinda the inverse of deux.” - Musicalman [sonically, not like you were to invert something on your own]

"I found the vocals to be a bit muddy, but the instrumental is alright!" - gilliaan

"I think it['s] better on Vocal than Instrumental"

- Mel-Roformer 2024.10 (Bas Curtiz model fine-tuned by ZFTurbo) only on MVSEP
Vocals bleedless: 37.80, fullness: 17.07, SDR 11.28
Small amounts of bleeding from instrumentals (inst. bleedless 39.20), might struggle with flute occasionally, good enough for creating RVC datasets.

- BS-Roformer 2024.08 (viperx model fine-tuned v2 by ZFTurbo) on MVSEP
Vocals bleedless: 37.61, fullness: 15.89, SDR: 11.32
Good for inverts, Dolby, lots of harmonies, BGVs. Good or even the best vocal fullness for some genres ~Isling, decent all-rounder, but might be muddier than Mel models here, although it gives less vocal residues than all the Mel Kim fine-tune models here, can be also used for RVC). “I've found it very useful for extremely quiet vocals that Mel couldn't extract” - Dry Paint Dealer. It’s a second MVSEP’s fine-tune of viperx model.
Iirc, it’s used as a preprocessor model for "Extract from vocals part" feature on MVSEP.

#3

- MVSep Becruily/ZFTurbo BS-Roformer high instrum fullness 124 bands

Only for premium MVSep users (demanding architecture).

Vocals bleedless: 36.30, fullness: 21.38, SDR: 12.25

It’s a paywalled model due to increased separation time from the model architecture. Iirc 1 credit per minute.

“It’s a fine-tune of ZFTurbo's model [BS-Roformer 124 bands], so unfortunately private”

“It’s dual model so vocals are also fullness, but whether they’re better or not than the original level 1-3 I don’t know” - becruily

for vocals it has much better bleedless metric than Deux, but less fullness.

Perfect mix of fullness and bleedless - dynamic64

Q: How does it compare to deux and HyperACE, any insight?

A: “I like it more”, “I feel like it functionally replaces the BS Rofo 124 fullness model” - -||-
“In my opinion brecruily high inst fullness vocal result seems to give the same result of BS Roformer fullness level 1” - mrmason347

- MVSep Ensemble 11.93 (vocals, instrum) (2025.06.28) - only for premium users

Vocals bleedless: 36.30, fullness: 17.73, SDR: 11.93

Surpassed sami-bytedance-v.1.1 on the multisong dataset SDR-wise.

Community bleedless models

- BS-Roformer SW 6 stem / vocals-only variant (MVSEP | Colab | nextgen.uvronline) 

Vocals bleedless: 36.06, fullness: 16.95, SDR 11.36

Good for e.g. some deep voices. Might resemble MVSEP BS-Roformer 2025.07, as the SW was used as a base for it.

- unwa Big Beta 7 Mel-Roformer vocal model | MK Colab

Bleedless 38.77, fullness 16.20, SDR 11.20.

“SDR is lower than the BS model, but personally I prefer this one. (...) Although not reflected in the metrics, noise has been reduced in sections without vocals.” - unwa

More bleedless model than 6X, although people tend to prefer the latter.

It's free of distortion anvuew 12.55 model has.

- Unwa Kim Mel-Band Roformer Bleedless FT2 | download | Colab | Huggingface / 2 | Kaggle | UVR instruction

Vocals bleedless 39.30, fullness 15.77, SDR 11.05
(voc. fullness is worse than the Mel Kim - 16.26, bleedless better)
inst. bleedless is still lower than base Mel-Kim model: 46.30 vs 46.56)

“I usually use big beta 6x, big beta 5e if that fails, and FT2 bleedless if I want very low noise or instruments are quiet (it gets muddy quick)” - Rainboom Dash

- Unwa’s BS-Roformer Resurrection voc (voc. variant) | yaml | Colab

Vocals bleedless: 39.99, fullness: 15.14, SDR: 11.34

Shares some similarities with the SW model, including small size (might be a retrain). The default chunk_size is pretty big, so if you run out of memory, decrease it to e.g. 523776.

Make a back of your UVR folder before installing it. We had a report that this model might break your UVR installation probably permanently for some reason.

- Unwa’s Revive 2 BS-Roformer fine-tune of viperx 1297 model | config | Colab

Vocals bleedless: 40.07, fullness: 15.13, SDR: 10.97

“has a bleedless score that surpasses the FT2 Bleedless”

“can keep the string well”

It’s depth 12 and dim 512, so the inference is much slower than some newer Mel-Roformers.

Worse bleedless than above, higher fullness

- BS PolarFormer 124 bands fullness level 1 variant - only on MVSep

voc. bleedless 32.37, fullness 20.95, SDR 11.96

- Anvuew BS-Roformer 12.55 FT1 model | yaml | nextgen.uvronline

Vocal bleedless 35.17, fullness 19.88, SDR 11.54.

The metrics are better than the Unwa BS-Roformer HyperAce v2 and Gabox voc_fv7 Mel-Roformer.

“I compared the model with big beta 7. Both have pretty similar outputs. But for some reason, 12.55 has a bit of distortion, while big beta 7 sounds fine [and the 12.45] .

for the lead vocals, they’re quite similar, but for the backing vocals, 12.55 sounds slightly buried (you probably wouldn’t notice unless you listen really closely)” - neoculture

- Anvuew BS-Roformer 12.45 vocal model | download

For UVR RTX 5000 patch and torch._dynamo error, edit “use_torch_checkpoint: True” to False (or delete the line).

Can be muddy. Not so balanced like Beta6X or vocfv7beta1, but “it properly doesn't capture the instrument here. Even FT2 bleedless gets tricked by this part, but this does just fine.” - rainboomdash

- Unwa BS-Roformer HyperAce v2 vocal model | separate Colab #2 | MVSEP

voc. bleedless 34.08, fullness 19.10, SDR 11.40
JFYI - there was no v1 of the voc variant.

“Compared to the Resurrection vocal model, this new model achieves higher scores across the board except for the bleedless score.”

“I feel like most songs hyperace gets a good bleedless result but ill probably stick with voc_fv7 beta3 or revive e3” - 5b

“it seems to be like... trying really hard to pull certain stuff out that is causing noise, I think

it doesn't sound that bleedless to me, due to that reason (..) there's some noise during some quieter parts but I'll def be adding this to the list of models I use. I was testing against fv7beta2, during louder parts there seemed to be less noise, but during quieter parts there seemed to be more noise (...)

I'll prob use fv7beta2 for the most part still, but I'll try adding a hyperace vocal model to the mix of models I use. lol, soon I'll be using 10 models throughout one song” - Rainboomdash

Note: It also uses its own inference script (bs_roformer.py) and is different from the previous one, and it’s also incompatible with UVR. “You can use this model by replacing the MSST repository's models/bs_roformer.py with the repository's bs_roformer.py.”

To not affect functionality of other BS-Roformer models by that file, so older BS-Roformers will still work, you can add it as new model_type by editing utils/settings.py and models/bs_roformer/init.py here (thx anvuew).

For error while installing the py file for HyperACE model in Sucial’s WebUI:

from models.bs_roformer.attend import Attend

ModuleNotFoundError: No module named 'models'"

The fix: “SUC-DriverOld/MSST-WebUI use the name "modules" and ZFTurbo/Music-Source-Separation-Training use the name "models". And Unwa's  bs_roformer.py that you replace with, also use "models" - fjordfish
In order to make it work in Sucial MSST Web UI you have to edit line (...) in the bsroformer.py file that is included with (...) model and change the word "models" to "modules". - rage313_

- Gabox voc_fv7 Mel-Roformer | yaml | makidanyee’s Colab

voc. bleedless: 33.85, fullness: 17.99, SDR: 11.16

For UVR RTX 5000 patch and torch._dynamo error, edit “use_torch_checkpoint: True” to False (or delete the line)..

Sometimes results sound not as noisy as BigBeta5e (but it depends on a asong) and voc HyperACE (which might sound less muddy, but noisier), but it catches harmonies much better than BigBeta6X:

voc_fv7 is "a little noisy, quiet, but at least it's capturing [the harmonies] (...) As with all models, it varies drastically song to song on how full it is compared to other models

fv7 was also cutting the reverb really aggressively on one song compared to big beta 6x" - rainboomdash

- Unwa Mel-Roformer Big Beta 6X model | yaml | Colab | AI Hub Colab | Huggingface | uvronline for premium (not on MVSep)
voc. bleedless: 35.16, fullness: 17.77, SDR: 11.12

“it is probably the highest SDR or log wmse score in my model to date.”

“There's some noise audible; it doesn't sound as clean when you compare to a more bleedless model (...) but it's certainly not fullness... (...) I think calling it bleedless wouldn't be crazy... makes more sense than "middle of the road" - rainboomdash
“Significantly better” than 5e for some people, although slower. Some leaks into vocal might occur, plus “The biggest problem with the model is the remaining background noise. If it were cleaner, it would already be an almost perfect result.” - musictrack

“6X has a lot less noise on vocals, but it's pretty muddy. I would prefer something in between [5e and 6X]. I tried to apply the phase [fixer/swapper] to the vocals and the noise was reduced, but only slightly.” - Aufr33

Some people might prefer fv5 instead [at least on some songs] ~5b

“6x is picking up BV just fine, where voc fv4 is failing” - Rainboom Dash

Training details:

“dim 512, depth 12. It is the largest Mel-Band Roformer model I have ever uploaded.” - “the same as [the] Bas Curtiz Edition” model. It has a bigger SDR vs smaller depth 6 Big Beta 6 model. “I've added dozens of samples and songs that use a lot of them to the dataset”

Fullness models

- becruily dual Mel-Roformer model “deux” (its vocal stem):
voc. bleedless: 28.30, fullness: 23.25, SDR: 11.37

Compatible with UVR Roformer patch (including the RTX 5000 one) | Colab | uvronline | MVSEP
“Unfortunately [both stems] won't null to the mixture perfectly, this applies to any multi stem model” - becruily
“vocal model picks up some things super well, like certain backing vocals or there was yelling in the background that it picked up that fv7beta didn't (...) generally slightly less full than fv7beta3 from the songs I tested, but has less instrumental bleed and slightly less noise.”, “generally does pretty well with harmonies”

“it’s good, but lot of noise during quieter parts that other vocal models don't have” - rainboomdash

“deux vocal stem has more backing vocals than Gabox fv7 beta 1-3 and other models I've ever tried, but it may be rather noisy on silent parts or fadeouts” - makidanyee

“I find deux does the best at keeping reverb trails intact” - Pipedream
“sounds great! I don't hear the dips in the vocals from some songs that are overly compressed” - Rage313

It doesn't work with speech and dialogues from movies - dca100fb8

“it leaves spectrums looking strange and wavey in my experience, like every time” - cyclorana, dynamic64

Noisier at silent sections than mag v2 - dca

“I exported it using FP16 just for the smaller size [only 432 MB], the quality is the same.” - becruily

- Colab auto-cleaning Deux result acapella by billieoconnell using the invert-clean model by becruily

- anvuew BS_RoFormer_mag | MK Colab

voc. bleedless: 32.17, fullness 22.15, SDR 11.09'

“specialized for magnitude spectrum accuracy” - anvuew, “isn't made to be used as a standard vocal model. It's only where magnitude is the most important (like RVC [SVC, TTS]). [The] mag is also perceptually pretty noisy”, “very noisy at quiet parts” - rainboomdash
“pure magnitude spectrum loss, so its phase is less accurate.”

“managed to pull out harmonies I knew were missing from my test track as well as not extracting some vocal sampled perc that had been in other models previously” - cristouk

“I found it to be less noisy than deux (for vocals obviously) and almost same fullness”, “like deux but without the noisy problem” - dca

- anvuew BS_Roformer_mag_v2 | DL | Colab | (you'll find it below deux and 1297 models) | nextgen.uvronline.app | Metrics

Vocals bleedless: 27.50, fullness: 26.22, SDR: 10.74

“too noisy and SDR drops a lot, so I just shared it on GDrive instead of HF” - anvuew

The older mag had 22.15 fullness, and it's the only metric which increased in the V2.

Still, it turned out to be a favourite model of one of our members (5b), so it probably needs some recognition anyway.

- Unwa bs_roformer_revive3e | config | Colab | nextgen.uvronline.app

voc. bleedless: 30.51, fullness: 21.43, SDR: 10.98

“A vocal model specialized in fullness.

(...) the opposite of version 2 — it pushes fullness to the extreme.

Also, the training dataset was provided by Aufr33. Many thanks for that.” - Unwa

“seems to sound better than beta5e, it sounds fuller, but this also means it sounds noisier” - gilliaan. For some people, it’s even the best, but “has issues with harmonies

and noisy” - rainboomdash

Even more fullness, less bleedless

- Gabox experimental Mel-Roformer voc_fv6 model | yaml | Colab

voc. bleedless: 26.61, fullness: 24.93, SDR: 10.64
“Definitely not bleedless” - rainboomdash, “Sounds like b5e with vocal enhancer. Needs more training, some instruments are confused as vocals” - Gabox. “fv6 = fv4 but with better background vocal capture” - neoculture

“very indecisive about whether to put vocal chops in the vocal stem or instrumental stem.

sometimes it plays in vocals and fades out into instrumental stem and sometimes it just splits it in half kinda and plays in both at the same time lol” - Isling

“I think [it] is the fullest vocal model I've heard, aside from maybe the SCNet high fullness ones lol. Oh and revive 3e and b5e are full too but yeah.” - Musicalman

- SCNet XL very high fullness on MVSEP

voc. bleedless: 25.30, fullness: 23.50, SDR: 10.40

- SCNet XL IHF (high instrum fullness by bercuily)

voc. bleedless: 25.48, fullness: 22.70, SDR: 10.87

(it was made mainly for instrumentals, but “It can also be an insane vocal model too”

____

- MVSEP SCNet XL IHF

voc. bleedless 28.31, fullness 17.98, SDR: 11.11

“It has a better SDR than previous versions. Very close to Roformers now.” also, vocal bleedless is the best among all SCNet variants on MVSEP. Metrics. IHF - “Improved high frequencies”.

“Certainly sounds better than classic SCNet XL (...) less crossbleeding of vocals in instrumental (...), and handles complex vocals better” - dca

Middle of the road #1 (lower fullness)

“Beta 1 and 3 are higher fullness, beta 2 is more middle-ground (still a fullness model)

I think 3 might be more consistent with not having some songs be overly noisy” - rainboomdash

“The fv7 betas sound great but always too much drum bleed in vocal stem for me” - cyclorana

- Gabox vocfv7 beta 2 Mel-Roformer model | yaml | Colab 

voc. bleedless: 31.55, fullness: 20.44, SDR: 10.87

“fullness went down a little bit” vs beta 1 (...) Definitely an improvement over fv4 (...) still quite a bit fuller than big beta 6x, but has less noise than even fv4 (...) at least when the instruments are loud, fv7beta2 is usually quite a bit less noisy than fv4, while still maintaining a decent amount of fullness... it is a bit less, but not too much (...) both are pretty noisy with fv4 (...) sometimes the noise can be pretty significant with fv7beta1, and fv7beta2 may have the fullness you desire. (...) “I'm really liking the balance of fullness and noise for most songs. fv4 and fv6/fv7beta1 are usually pretty noisy... this is less noisy, but still has a good amount of fullness.” still gonna have an issue with backing vocals compared to fv7beta1 sometimes… (...) “Fv7beta2 has still been significantly better with BV than fv4, despite quite a bit less noise” but “significant issues on one song, while fv6/fv7beta1 didn't” - rainboomdash

- Gabox vocfv7 beta3 Mel-Roformer | yaml | Colab | nextgen.uvronline.app

voc. bleedless 30.83, fullness 21.82, SDR 10.80

“beta 1 and 2... eh, pretty close to same instrumental bleed,

but beta 3 [is] definitely a step up from the two songs I compared (...)

most songs so far, fv7beta3 is fuller than fv7beta1,

def less robotic sounding at times (when a voice gets quiet/hard to capture, and it just fails).

Just had another song where fv7beta1 was fuller than fv7beta3, but it was also a lot noisier.

A large majority of the songs I tested, fv7beta3 was fuller... I think fv7beta3 is usually a bit noisier than fv7beta1? But also sounds fuller in those cases, I'd say it's generally worth [the]

instrumental bleed. Usually worse with fv7beta3 versus fv7beta1, but it depends.

Fv7beta2 is always less full/less noise, but [has] only slightly less instrumental bleed than fv7beta1” - rainboomdash

- Gabox Mel-Roformer vocfv7 beta 1 (a.k.a. vocfv7beta1) model | yaml | Colab

voc. bleedless: 30.81, fullness: 21.21, SDR: 10.96

“one step below the extreme fullness models (...) fv6 on average is more full” - rainboomdash. "Just a better fv4 it seems, better bleedless" (bleedless: 29.07, fullness: 21.33, SDR 10.58)

vs fv4 "It is noisier... Kinda closer to beta 5e?” “It's slightly less noise and fullness than beta 5e but picking up the backing vocals REALLY well, significantly better than beta 5e”

But it's pulling the backing vocals out even better than 5e” “the backing vocals are so good!

“it does have significant synth bleed, too... it at least wasn't coming through at full volume

when I say fullness, I specifically mean how muddy it sounds”,
“seems to capture partial instruments a lot.” - rainboomdash

_______

- Gabox Mel-Roformer voc_fv4 | yaml | Colab | Huggingface / 2

voc. bleedless 29.07, fullness 21.33, SDR 10.58

“Very clean, non-muddy vocals. Loving this model so far” (mrmason347)

Good for anime and RVC purposes, currently the best public model for it (codename)

“The important thing for an RVC dataset is to get lead vocals so fv4 is good for that

The newer karaoke models are also helpful” - Ryan

Some might prefer voc_gabox2 instead, occasionally - chroniclaugh.

The opposite of Beta6x which has “lower noise but [is] less full/muddier (...) noise/muddiness seems between 6x and 5e, but even 6x is picking up BV just fine, where voc fv4 is failing”

Some people might want to test it with even overlap 32, and then:

“It's close to perfect, the only thing is it kinda struggled with picking up the adlibs and the delay, but the lead vocal is almost perfect I think. (...) on another song (...) 5e is just too noisy and 6x is muddy, fv4 is best of both worlds (...) has segments with constant significant vocal bleed (for the most part, it's not audible at all) (...) I was trying to get an acapella and every model failed except this one. It's not perfect, but I guess some songs are just too hard for the AI.” - Rainboom Dash

Good also for instrumentals, if you need less vocal residues than typical instrumental Roformers (even less than Mel Kim, FT2 Bleedless, or Beta 6X - makidanyee.

“even beta 6x is a lot better at pulling that background vocal out than voc fv4...

and that's a less full model. hmm, fv6 is noisier and also not picking up the backing vocals as full as the last mel band roformer” - rainboomdash

- Unwa Mel big beta 5e vocal model | Colab | Huggingface / 2 | MVSEP | MSST-GUI | UVR
yaml: big_beta5e.yaml or fixed yaml for AttributeError in UVR
voc. bleedless: 32.07, fullness: 20.77 (the biggest for now), vocals SDR: 10.66

“feel so full AF, but it has noticeable noise similar to lew's vocal enhancer”
You can alleviate some of this noise/residues by using phase fixer/swapper and using becruily vocals model as reference (imogen).
It seems to fix some issues with trumpets in vocal stem - maxi74x1.
“It's noisy and, IDK, grainy? When the accompaniment gets too loud. (...) Definitely not muddy though, which is a welcome change IMHO. I think I prefer beta 4 overall” - Musicalman “ending of the words also have a robotic noise” - John UVR
“Perhaps a phase problem is occurring” - unwa. Phase swapper doesn’t fix the issue (it works for inst unwa’s models).
If you try big beta 5e on a song that has lots of vocal chops, the vocal chops will be phasing in and out and sound muddy (Isling).

“Excellent for ASMR, for separating Whispers and noise, the quality is super good

That's good when your mic/pc makes a lot of noise. All the denoise models are a bit too harsh for ASMR (gilliaan)”

Worse for RVC than Beta 4 model below (codename, NotEddy)

It can handle reverb tails better than 6X.

- Mel-Roformer vocal by becruily model | config for ensemble in UVR | MVSEP | Colab
voc. bleedless: 31.26, fullness: 20.72 (on pair with 5e), SDR: 10.55 | Huggingface / 2

Lower bleedless than 5e,“pulling almost studio quality metal screams effortlessly, wOw ive NEVER heard that scream so cleanly”

(on older UVR beta patches) If you use lower dim_t like 256 at the bottom of config for slower GPU these are the first models to have muddy results with it.
Consider setting 485100 chunk_size in the yaml for the highest SDR.
Currently used on x-minus/uvronline as a model for phase fixer.

- Gabox Mel-Roformer voc_fv5 | yaml | Colab

voc. bleedless: 29.50, fullness: 20.67, SDR: 10.56

“fv5 sounds a bit fuller than fv4, but the vocal chops end up in the vocal stem. In my opinion, fv4 is better for removing vocal chops from the vocal stem” - neoculture. Examples

Other models

- BS-Polarformer (PoPE) 62 bands vocal model on MVSep

Bleedless 35.90, fullness 17.68, SDR 11.75.

Looks like further trained from the below.

- BS-Polarformer a.k.a. PoPE vocal by ZFTurbo model vocals | HF inference / #2
SDR: 11

Trained on a new arch enhancement - you need updated MSST to run it.

Decent results, just needs some polish, maybe fine-tuning. Only a bit worse metrics than the Mel Kim.

- BS-Polarformer Aname instvoc

mirror: model | yaml | HF inference / #2

Fine-tune of the ZFTurbo’s, also decent.

- MVSep BS-Roformer vocals from “MVSep Mega 53 Stems” single model | yaml | splifft | MVSepless: HF / HF CPU / Colab / newer (you can pick which stems you want there, but it’s in Russian, and HF doesn’t support auto-translate, but Colab might)

The model can be muddy, it's small. It was further trained on MVSep in the all-in-one model and returns only present stems there.

Might be a retrain from some of the MVSep BS-Roformer vocal models on the site, not sure, but it's only 77MB, so probably very fast, but its performance might be limited.

Hint: On HF, if you don't check any models from the list after choosing the OG model, all will be processed.

Older or outperformed models

- Gabox Mel-Roformer voc_gabox2 model | yaml | Colab

voc. bleedless: 33.13, fullness: 18.98, SDR: 10.98

- Gabox Mel-Roformer Vocal F (fullness) v3 model | Colab | Huggingface / 2

voc bleedless: 32.15, fullness 19.97

- Gabox Mel-Roformer Vocal F (fullness) v2 model | Colab | Huggingface / 2

voc bleedless: 33.40, fullness: 19.31

- Aname Mel FullnessVocalModel (yaml) model | Colab | Huggingface / 2

voc. bleedless: 32.98 (less than beta 4), fullness: 18.83 (less than big beta 5e/voc_fv4/becruily, more than beta 4)

- Gabox Mel-Roformer voc_gabox (Kim/Unwa/Becruily FT) model | Colab Huggingface / 2

voc. bleedless: 34.66 (better than 5e, beta 4 and becruily voc), fullness 18.10 (on pair with beta 4, worse than 5e and becruily)

- Mel-Roformer unwa’s beta 4 (Kim’s model fine-tuned) download | Colab | Huggingface / 2

voc. bleedless: 33.76, fullness: 18.09
“Clarity and fullness” - even compared to newer models above.

Beta 1/2 were more muddy than Kim’s Roformer, potentially a bit less of residues, a bit more artificial sound. Ringing issues in higher frequencies fixed in beta 3 and later. It’s good for RVC (and favourite codename’s public model for RVC before voc_fv4 was released). Fuller vocals than Bas Curtiz FT on MVSEP (but can bleed more synths) ~becruily
Unwa’s vocal models are capable of handling sidechain in songs - John UVR

Bleedless models #3

- BS-Roformer Revive unwa’s vocal model experimental | yaml 

voc. bleedless: 38.80, fullness: 15.48, SDR: 11.03

viperx 1297 model fine-tuned. “Less instrument bleed in vocal track compared to BS 1296/1297” but it still has many issues, “has fewer problems with instruments bleeding it seems compared to Mel. (...) 1297 had very few instrument bleeding in vocal, and that Revive model is even better at this.

Works great as a phase fixer reference to remove Mel Roformer inst models noise” (dca)

- SYHFT V5 Beta - only on x-minus/uvronline (still available only with this link for premium users, and for free)

Vocal bleedless: 37.27, fullness, 16.18, SDR: 10.82

Other models #2

- Unwa’s Kim Mel-Band Roformer FT2 | model | Colab
voc. bleedless: 37.06, fullness: 16.61 (fullness worse vs the previous FT, but both metrics are better than Kim’s)
It tends to muddy instrumental outputs at times, similarly like the OG Kim’s model was doing, which didn’t happen in the previous FT below. Metrics

- Unwa Kim Mel-Band Roformer FT3 Preview | model | yaml | Colab | uvronline via special link for: free/premium (scroll down) or nextgen.uvronline.app

voc. bleedless: 36.11, fullness: 16.80, SDR: 11.05

“primarily aimed at reducing leakage of wind instruments to vocals.”

For now, FT2 has less leakage for some songs (maybe till the next FT will be released)

- Unwa’s Mel Big Beta 6 vocal model | yaml | Colab | Huggingface / 2 | AI Hub Colab
Similar to FT series. “Although it belongs to the Big series, the characteristics of the model are similar to those of the FT series. (...) this model is based on FT2 bleedless with the dim increased to 512”.

Muddier than Big Beta 5[e], might be better than FT2 at times.
“If you liked the output of the Big Beta 5e model, you may not like 6 as much; it does not have the output noise problem of 5e, but instead sacrifices Fullness. (...) Simply put, it is a more conservative model” (unwa)

For anime and RVC “isn't as audibly and spectrally full as fv4 + can at times have flat-line artifact at the very top, but then, fv4 can sometimes have "crunchy" noise present at some places, so an ensemble of those 2 is probs a good idea (or might be fv4 flash more on less aggressive scenes).” codename

- Unwa’s Kim Mel-Band Roformer FT vocal model | Colab
Voc. bleedless 36.95 (vs 36.75), fullness 16.40 (vs 16.26) metrics for vocals are enhanced vs the original Mel Kim model. SDR-wise it’s a tad lower (10.97 vs 11.02).

Older models list continues later below

Tips for separating vocals

- Separate with becruily Mel Vocal model and its instrumental model variant, then get vocals from the vocal model, and instrumental from instrumental model, import both stems for the DAW of your choice (can be Audacity) so you’ll get a file sounding like original file, then export - so perform a mixdown of both stems, then separate it with vocal model (mrmason347 /Havoc)

- “In my testing, I've found that SCNet very high fullness (on MVSEP) put through Mel-Roformer denoise (average) and UVR denoise (minimum) has the best acapella result” dynamic

Depending on a model, some Roformers might be muddy. Then consider using ensembles or Apollo enhancer model by Lew v2 , although it might be noisy. Might work the best on BS+Mel ensembles (max spec, though avg might work better in some cases), v2 (model | config | Colab | Inference) (this Colab now probably works instead), v1 (model | config), and it also can be used in the latest UVR Roformer beta.

- Sometimes using EQ stressing vocals properly might be beneficial for separation too

- You might potentially also try to experiment with demudder added with the beta patch #14 linked above. Normally demudder works only for instrumentals, but when you switch in the config editor to vocal stem being instrumental and in reverse, then demudder will work vocals. If your model have “other” stem instead of “instrumental” or “vocal”, you’ll need to rename it. Demudder requires stem labelled as instrumental to work with.


Ensembles

(for vocals)

Note: Max FFT and Spec can be used interchangeably

Ensemble instructions


MVSep-exclusive (and paid):

- #1 - BS becruily 124 fullness + PolarFormer fullness lv 1 (Avg)

- #2 - BS MVSep 124 fullness lv 1 + PolarFormer fullness lv 1 (Avg)

--->The first one should be better (dca100fb8)

Open/free ensembles:
It’s still better to use Deux on its own, but in cost of more bleed, you can get more BVs

#1 BS anvuew Mag v2 + SW (Max Spec)

#2 BS Mag + SW (Max Spec)

#5 BS anvuew ft1 + SW (Max Spec) - noisy at silent passages

#6 Mel deux + BS SW (Max Spec)

#5 BS HyperACE v2 Voc + BS SW (Max Spec)
#2 Mel voc fv6 + BS SW (Max Spec)

#4 Mel vocfv7beta1 + BS SW (Max Spec)
- (dca100fb8)

E.g Mag v2 and voc fv6 ensembled with the SW can be also used on nextgen.uvronline.app.

“Both free users and paid users can use it.” - squiddagod (more)

- MVSep presets editor now can be used to create a node for the ensemble of your choice (“If 2+ algorithm nodes, it is premium only”), avg used by default. Some presets are already shared by community

- SESA Colab

- MVSepless Colab

(older ensembles)

- Mel vocfv7beta3 + BS SW (Max Spec)

- Mel deux + BS 2025.07 (Max FFT/Spec) (former “best Ensemble for vocals”) (-||-)

- MVSep Polarformer, Deux and MVSep 2025.07 (Max FFT) (“The result is great.

Has all the vocals from Polarformer and 2025.07 but the fullness of Deux, along with a bit of instrumental bleed. But the result is better than any of the separate models”) (Kashi/omaroka)

- deux + big beta 7 (Max FFT) (“another good ensemble”) (neoculture)

- BS Roformer HyperACE Voc v2 + 2025.07 (Max FFT) (formerly best) (dca)

- BS Revive 3e + BS 2025.07 (Max FFT) (former “best vocal ensemble”) (-||-)

- Mel Becruily Vocal + MVSEP’s BS 2025.07 (Max FFT) (-||-) (-||-)

- unwa’s bigbeta5 + becruily vocal (Max FFT) (midol)

- voc_gaboxFv2 + becruily vocal (heauxdontlast/gilliaan)

- big beta 7 + bs_roformer_mag (Max FFT) (good for RVC) (neoculture)


- Unwa “Big Beta 4 + Big Beta 5e (Average Spec) (“really good to reduce the noise while keeping the fullness”) (gilliaan)

- unwa beta6 + voc_fv4 (good for anime and RVC)

- unwa beta6x + voc_fv4 (“some songs I can use big beta 6x, and it's enough, others I need to ensemble it with voc_fv4”) (Rainboom Dash)

- unwa beta6x + voc_fv6 (“would make a good ensemble, but the amount of noise is horrific

and I heard that phase swapper would fix it”)

- unwa beta 6x + beta5e (avg/avg) (good for RVC) (gilliaan)

- BSRoformer-Viperx1297, BSRoformer-LargeV1 by Unwa, unwa_ft2_bleedless, mel_band_roformer_vocals_becruily, Gabox voc_fv4 - Average/Average Spec (good for cleaning inverts) (AG89)

- Models ensembled (inst, voc) available for premium users on mvsep.com

(SDR 10.44-11.93 and “High Vocal Fullness” variants)

RVC models choice with AI Hub advice (subject to change;
read their current docs too)

If you can separate with these models downloaded from above locally, see also here for the list of all cloud sites and Colabs.

“If you need to remove multiple noises, follow this pipeline for the best results:

Remove instrumental -> Remove reverb [probably on vocals] -> Extract main vocals -> Remove noise

Or also Isling’s approach “gives insanely clean results”:

Vocals>De-reverb>Karaoke

“That’s how I get completely raw lead vocals if I don’t have the multitracks to a song:

Big Beta 6X (Unwa) -> Karaoke (Frazer & Becruily) -> De-Reverb V2 (Anvuew) -> De-Reverb Room -> (Anvuew) -> Less Aggressive Denoise (Aufr33)” - natethegratevhs

Note: The room model outputs mono natively, you might need to process every channel separately if you don't use MVSEP.

“I’m using deux to extract vocals. For lead vocals, I ensemble frazer becruily karaoke and anvuew karaoke [max spec], but I target the backing vocals instead of the lead. Once I get the backing vocals from the ensemble result, I open Audacity and invert the ensemble backing vocal track against the deux vocal track (that’s how I get a cleaner lead vocal, at least for my tracks). After I’ve collected the lead vocals, i run anvuew dereverb mono to remove the reverb” - neoculture

Recommended all-vocals models/ensembles for RVC:

- MelBand Roformer | Vocals FV4 (a.k.a. voc_fv4) by Gabox
- Gabox vocfv7beta1 (“seems to give better results than fv4”)

- Mel 2024.10 (mentioned in the MVSEP section of AI Hub, but -)

- BS-Roformer 2025.07 (now has all the metrics better)

- big beta 7 + bs_roformer_mag (max spec) - neoculture (the mag was trained with “pure magnitude spectrum loss (...) for tasks that only need magnitude, like SVC and TTS”)

- unwa beta6/x + voc_fv4 (also good for RVC)

- deux vocal stem (“has more backing vocals than gabox fv7 beta 1-3 and other models I've ever tried, but it may be rather noisy on silent parts or fadeouts” - makidanyee)

- unwa beta 6x + beta5e (avg/avg) - gilliaan

- unwa beta 4 (was better than big beta v5e (NotEddy/codename); research also voc_gabox2)

Potentially:
anvuew mag or mag_v2 (metallic noise, Leap xe “eliminates that noise but filters out too many high frequencies above 16kHz” - adrielmz_

To experiment more, visit ensembles for general vocal sepration

Instrumentals

- MelBand Roformer | INSTV7 by Gabox
(unwa instrumental v1e+ OR Mel 2024.10 are also mentioned in their MVSEP section and Gabox Fv7z is mentioned in the x-minus)

De-reverb

- MelBand Roformer | De-Reverb by anvuew
(it’s probably v2 variant [also mentioned there], or also Sucial V2 (MelRoformer) mentioned in their MVSEP section [“if I'm unhappy with the results I go for Sucial

- isling”] - it probably follows the model naming scheme of UVR UI on HF, also the new mono-dereverb model is being used occasionally)

Backing Vocals

- Mel-Roformer-Karaoke-Aufr33-Viperx (surpassed by Becruily and Frazer Karaoke, but the first can be more consistent; anvuew's Karaoke model have fuller lead vocals; also older Model fuzed gabox & aufr33/viperx (SDR: 9.85) is mentioned in their MVSEP section)

De-noise

- Mel-Roformer-Denoise-Aufr33-Aggr (they mention also “Mel denoiser v2” in UVR section)

Restoration

- For lossy mp3/mixtures: Apollo Universal by Lew (sometimes AudioSR can be better)

- For voice: AP-BWE or ClearerVoice-Studio's Clear Voice “my favorite is the 2nd one” - codename0)

Fast inference models for general use

Above an hour on i3-7100u, small - the lightest Roformers, while most used to be 870 MB):

For vocals

- Unwa Resurrection BS-Roformer (yaml | Colab, 195 MB)

- BS-Roformer SW vocals only (mask_estimators.0 on the regular 6 stem model, 195 MB)

Models like BS_RoFormer_mag or Anvuew BS-Roformer 12.45 also have similar size, although not all small size models have to be similarly fast like above, but feel free to test (e.g. hyperacev2 is much slower than the Resurrection).

- Vocals from MVSep 53 stem model (single model | yaml, 77 MB)
(it’s even smaller, but its performance might be mediocre compared to the above)

Older models

- Aname Mel-Roformer small (203MB)

- Unwa Mel-Roformer small (203MB)

Older arch (faster; 25-60 minutes+ on weak i3u/C2Q respectively)

- voc_ft (probably the fastest, but uses outperformed MDX-Net v2 arch, also it’s narrowband)

- Kim Vocal 2 (or ev. 1, -||-, older model)

For instrumentals

- Unwa BS-Roformer Resurrection inst (yaml) | a.k.a. “unwa high fullness inst" on MVSEP | uvronline free/premium | Colab | UVR (don’t confuse with Resurrection vocals variant, 204 MB)

- Gabox BS_ResurrectioN (model | yaml, 204 MB)

For both
(dual stem model, it don’t invert - you might save time instead of using two models)

- Becruily Mel Deux (decent vocals and instrumentals, although sometimes bleedy)

Older models

- Unwa BS-Roformer-Inst-FNO (works only in MSST after modifying py file like in the model card, similar to decently performing Resurrection inst model, 332 MB)

- Gabox Mel-Roformer small_inst | yaml (experimental, 203 MB)

- Unwa BS-Roformer-Inst-EXP-Value-Residual (uses Mel v2 model type in UVR; If it wasn’t made compatible with MSST already, replace bs_roformer.py from this repo and
from bs_roformer.attend import attend

from models.bs_roformer.attend import attend

in bs_roformer.py file

generally not very good model, but sometimes capable: “successfully removed [vocals] and kept the digital choir atmosphere as well” vs deux, inst_gaboxFlowersV10 and HyperACE but it it’s considerably slower than deux - mohammedmehditber)

—-

4 stems

- Faster FP16 version of BS-Roformer 6 stems called splifft (by undef13; a tad lower SDR; only 334MB vs 700 MB in the OG weight, CPU/NVIDIA compatible, and potentially AMD ROCm, only bigger variant works in UVR; the OG “Conversion done after 2 hours for a 2 minute 49 second file” on 2/4 i3 7100u) - on CPU it might be slower than the OG, as it might not support FP16 natively due to even possible emulation. But probably Turing GPUs with tensors (e.g. RTX or T4) and newer, probably have FP16 acceleration, while non-RTX 16XX sometimes not.

Faster, lower quality:

- KUIELab-MDXNET23C (4 stems) - its first scores were probably from ensemble of its five models, and in that configuration it had better SDR than demucs_ft on its own, and drums had better SDR than “SCNet-large_starrytong” (so single models’ score of any of these MDX23C models is probably lower than in demucs_ft).
> Lighter “model1” drums sound surprisingly better than htdemucs non_ft v4 on previously separated instrumental. It handles trap really well and preserves hi-hats correctly, but at the cost of other stem bleeding. v4 model can be used to clean it a bit further,

- htdemucs v4 non-ft (UVR default) - it can clean up other stem bleeding of the above

The fastest models, usually lower quality, CPU-friendly

Instrumentals

- MDX-Net HQ_3, 4, 5 (the last is the fastest, 56 MB)

- MDX-Net inst3, Kim inst (older, narrowband models, but can be useful too in some cases, 63 MB)

Vocals

- MDX-Net voc_ft

- MDX-Net Kim vocal 2/1

(all narrowband models)

4 stems

- htdemucs_mmi - iirc the fastest demucs (v3) model, but worse quality than the others

- kuielab_b - lighting-fast, but quality is mediocre (but rather still better than Spleeter which might be even faster, but not necessarily)

________________

Older vocal models for general use (moved here for archiving purposes)

- Mel-Roformer unwa’s inst-voc model called “duality v1/2” (focused on both instrumental and vocal stem during training, but you can now test newer V1e+ single stem for this purpose too).

https://huggingface.co/pcunwa/Mel-Band-Roformer-InstVoc-Duality | Colab | MVSEP

Vocals sound similar to beta 4 model, but with more noise,
instrumentals are deprived of the noise present in inst v1 and later inst models, but as a downside, they’re more muddy for instrumentals.
v2 have slightly a bit better SDR and fewer residues

Because duality is a two stems target model.
"other" is output from model

"Instrumental" is inverted vocals against input audio.

The latter has lower SDR and more holes in the spectrum.

So, using MSST-GUI, leave the checkbox “extract instrumental” disabled for duality models.
You can use it in the Bas Curtiz’ GUI for ZFTurbo script (already added) or with the OG ZF’s repo, or in the Colab.

- Aname duality Mel model

- Aname Full Scratch Mel-Band Roformer model

bleedless 30.75 fullness 13.24, SDR: 8.01


- SYHFT (a.k.a. SYH99999/yukunelatyh) MelBandRoformer V3 | model

VS previous SYH’s models “this version is more consistent with separation. It's not what I'd call a clean model; It sometimes lets background noise bleed into the vocal stem. But only somewhat, and depending on how you look at it, it can be a good thing since it makes the vocals sound less muddy.” Musicalman


- MelBandRoformerBigSYHFTV1Fast | model - more vocal fullness metric, but more bleeding (although less than duality models and even Kim’s purely metric-wise). “same parameters size with Kim's. Other models are 2x scale parameter size to compare my model”

- Mel-Roformer model by Kim | model | config

Vocals bleedless: 36.75, fullness: 16.26, SDR: 11.07

(Colab/Huggingface/2/MVSEP/uvronline via special link for: free/premium (scroll down)/UVR beta Roformer (available in Download Center)/MSST-GUI/simple Colab/CML inference)

Usual base for lots of Mel fine-tunes on that list.

Sometimes might leave instrumental residues in vocals, but can be less muddy than other BS-Roformers - the same goes to any fine-tunes of this model vs BS 2024.08, so effectively all the Mel models above)

“godsend for voice modulated in synth/electronic songs” vs 1296 can be more problematic with wind instruments putting them in vocals.

- unwa’s instrumental Mel-Roformer v1e+

- unwa’s instrumental Mel-Roformer v2 model (similar to v1, but less noise, muddier, bigger, heavier model)

Model files | Colab |  uvronline via special link for: free/premium (scroll down) | MSST-GUI (It's now included in ZFTurbo's repo, it's the "gui-wx.py" file)

Might miss some samples or adlibs while cleaning inverts. SDR got a bit bigger (16.845 vs 16.595) “Sounds very similar to v1 but has less noise, pretty good” “the aforementioned noise from the V1 is less noticeable to none at all, depending on the track”.  “V2 is more muddy than V1 (on some songs), but less muddy than the Kim model. (...) [As for V1,] sometimes it's better at high frequencies” Aufr33

- older BS-Roformer 2024.02 on MVSEP (generally BS-Roformer models “can be slappy with choir-like vocals and background vocals” but “hot on pre-2000 rock”)

These older Roformers “kinda does poorly on large screams” in metal music, but not always. Sometimes even HQ_4 can catch them better than, e.g. viperx models.

- Mel-Roformer fine-tuned 17.48 model on MVSEP (works e.g. for live shows that have crowd)

(it’s different from the one on x-minus)

- Gabox BS-Roformer instrumental, which doesn’t struggle so much with choirs like most Mel-Roformers, although it may not help in all cases (link)

- “ver. 2024.04” SDR 17.55 on MVSEP - fine-tuned viperx model v1 (can pick in adlibs better, occasionally picks some SFX’, sometimes one, sometimes the other is “slightly worse at pulling out difficult vocals”)

- BS-Roformer Large v1 unwa’s vocal model (viperx 1297 model fine-tuned) download | mirror | yaml | Colab

More muddy than Kim’s Roformer, potentially a bit less of residues, a bit more artificial sound. Better than viperx model - “captures more nuances, subtle elements and details” ~A5
It can be better for some older music like The Beatles than above models.

- BS-Roformer viperx 1297 model (UVR/MVSEP a.k.a. SDR 17.17 model | yaml 

- BS-Roformer viperx a.k.a playdasegunda 1296 variant model | yaml
previously called just “BS-Roformer” on uvronline via special link for: free/premium and in nextgen.uvronline.app (legacy)

- Mel-Roformer viperx 1143 vocal model (UVR>Download More Models)

(don't confuse with 1053 which separates drums and bass in one stem).

The first Mel-Roformer vocal model trained by viperx. It was before Kim model which introduced changes to the config, and fixed the problem of lower SDR vs models trained on BS-Roformer.

Most people back then preferred Kim Mel-Roformer instead, but Mel viperx’ “does background voices correctly unlike the Kim's (it does not recognise background 'breee's)” “Iirc Viperx Mel Rofo doesn't struggle with instruments counted as vocals”.

Also, both Mel and BS variants of viperx model struggle with saxophone and e.g. some Arabic guitars. It can still depend on a song whether these are better than even the second oldest Roformer than on MVSEP (from before viperx model got fine-tuned version). Beside problems with recognizing instruments, they're very good for vocals (although Mel-Roformer by Kim on x-minus tends to be better).

Muddy instrumentals when not ensembled with other archs (but we didn’t have typically instrumental stem target models back then), maybe Mel variant less.

Be aware that names of these models on UVR refer to SDR measurements of vocals conducted on private viperx dataset, not even older Synthetic dataset, instead of on multisong dataset on MVSEP, hence the numbers are higher than in the multisong chart on MVSEP.

Older ensembles for vocals

- Models ensembled option on x-minus.pro (available only for premium users)

> Mel-Roformer + MDX23C (can be picked after you uploaded/processed a track [at least with Mel-Roformer model chosen]).

> Mel-Roformer + demudder

“I recommend mel-roformer + demudder to remove vocals from songs that contain only backing vocals that are so faint that our ears can barely hear them.”

- MDX23 by ZFTurbo (v. 2.5 jarredou Colab fork)

- Ensembles on MVSEP.com (for premium users)

- Ensembles in UVR 5:

a) 1296 + 1143 (BS-Roformer in beta UVR) + Inst HQ4 (dopfunk)

(there might be instrumental residues from HQ4 in some cases)

b) 1296 + 1297 + MDX23C HQ

c) Manual ensemble in UVR of models BS-Roformer 1296 + copy of the result + MDX23C HQ (jarredou) - for faster result and similar quality vs the one above

More ensembles beneath

- KaraFan (preset 4, but may give worse results than Mel-Roformer)

___

Older single models for vocals (available in UVR 5 | MDX-Net Colab for non-23C models | MVSEP)

- UVR-MDX-Net-Voc_FT (narrowband, further trained, fine-tuned version of the Kim vocal model; Roformers might be better now)

>If you still have instrumental bleeding, process the result with Kim vocal 2

>Alternatively use MDX23C narrowband (D1581) then Voc-FT, "great combination" (or MDX23C-InstVoc HQ instead of D1581)

(so separate with the D1581 or InstVoc model first, then use the separated result as input, and separate it further with voc_ft)

- Kim Vocal 1 (can bleed less than 2, but more than voc_ft, might depend on a song)

- Kim Vocal 2

>MDX-Net HQ_3/4/5 (HQ_4 can be sometimes not bad on vocals too, even less muddy than voc_ft, though more noisy, and e.g. HQ_3 had more vocal residues then Kim Vocal 2 in general, HQ_5 have stronger and fuller vocals than HQ_4)

>MDX23C-InstVoc HQ (can have some instruments residues at times, but it’s fullband - better clarity vs voc_ft and Kim Vocal 1/2 -

“This new model is [vs the narrowband vocal models], by far, the best in removing the most non-vocal information from an audio and recovering formants from buried passages... But in some cases, it also removes some airy parts from specific words, and some non-verbal sounds (breathing, moaning).”

- newer MDX23C epochs available on MVSEP like 16.66.

MDX23C models are go-to models for live recorded vocals

(available also in MDX23 Colab v2.3/2.4 when weight set only for InstVoc model)

Older UVR ensembles (from before Roformer models release)

>Voc FT + MDX23C_D1581 (avg/avg)

>292, 496, 406, 427, Kim Vocal 1, Kim Inst + Demucs ft (#1449)

>Kim Inst, Kim Vocal 1 (or/and voc_ft), Kim Vocal 2, UVR-MDX-NET Inst HQ 2 (or 3/4), UVR-MDX-NET_Main_427, htdemucs_ft (avg/avg IRC)

>Kim Vocal 1+2, MDX23C-InstVoc HQ, UVR-MDX-NET-Voc_FT

(jaredou)

> More ensembles

>You can also check some ensembles for instrumentals

Your choice of the best vocal models only (up to 4-5 max for the best SDR - more)

If your separation still bleeds, consider processing it further with models in Debleeding section further below.

___

Other services (multipurpose)

- Ripple (no longer works; since BS-Roformer models release it might be obsolete; it's very good at recognizing what is vocals and what's not and tends to not bleed instrumental into vocal stem; very good if not the best solutions for vocals)

- music.ai (paid; presumably in-house BS-Roformer models)

“almost the same as my cleaned up work (...) It seems to get the instrument bleed out quite well”)

“Beware, I've experienced some very weird phase issues with music.ai. I use it for bass, but vocals are too filtered/denoised IMO, and you can't choose to not filter it all so heavily. ” - Sam Hocking

- https://myxt.com/ (paid; uses Audioshake)

- moises.ai (paid; uses in-house BS-Roformer models, sometimes better results than the one on MVSEP)

- ZFTurbo’s VitLarge23 e.g. on MVSEP or 2.3/2.4 Colab (it's based on a new transformers arch. SDR-wise it's not better than MDX23C (9.78 vs 10.17), but works "great" for an ensemble consisting of two models with weights 2, 1. It's been added in 4 models ensembled on MVSEP (although the bag of current models is a subject to change any time)

- ZFTurbo’s Bandit Plus (MVSEP)

Other decent single UVR models

- Main (427) or 406, 340, MDXNET_2_9682 - all available in UVR5, some appear in download center after entering VIP code)

- or also instrumental models: Kim Inst and HQ_3 (via applied inversion automatically)

Other models

- ZFTurbo's Demucs v4 vocals 2023 (on MVSEP, unavailable in Colab, good when everything else fails)

- MDX23 Colab fork 2.1 / 2.2 (this might be slow) / 2.3 / 2.4 / 2.5 (it's generally better than UVR ensembles SDR-wise, but it's not available in UVR5) (MDX23 Colab is good also for instrumentals and 4 stems, very clean, sometimes more vocal residues in specific places vs single MDX-UVR inst3/Kim inst/HQ models, but it sounds better in overall, especially the Colab modification/fork with fixes made by jarredou)

- HQ_3 (inverted result giving vocals from instrumental in 2nd stem) - more instrumental residues than e.g. Kim Vocal 2, but no 17.7 cutoff)

- Narrowband MDX23C_D1581 “Leaves too much instrumental bleeding / non-vocal sounds behind the vocals. Formants are less refined than on any of the top vocal models (Voc FT, Kim 1, Kim 2 and MDX23C-InstVoc HQ).”

- Kavas' methods for HQ vocals:

Ensemble (Max/Max) - Low pass filter (brickwall) at 2k:

- MDX23C

- Voc FT

Voc FT - High Pass Filter (brickwall) at 2k

(“Sometimes it leaves some synth bleeding in the mids" then try out min/min)

Or:

Multiband EQ split at 2kHz with a low & high pass brickwall filter with:

-MDX23C-InstVoc from 0 to 2kHz and:

-Voc_FT from 2kHz onwards

(InstVoc gives fuller mids, but leaves transients from hats in the high end, whereas Voc ft lacks the mids, but gets rid of most transients. Combine the best of both for optimal results.)

- Any top ensemble or AI appearing on MVSEP leaderboard (but it depends, - sometimes it can be better for instrumental, sometimes vocals

Ensembles are resource consuming, no cutoff if one model is fullband and the other is narrowband. Random ensembles can result in more vocal or instrumental residues, as mentioned above.

Models not exclusive for MVSEP are all available in UVR5 GUI, or optionally you can separate MDX models in Colab and perform manual ensemble in UVR5 (no GPU or fast CPU required for this task) or use manual ensemble in Colab [may not work anymore]) or also in DAW by importing all the stems together and decreasing volume (you might want to turn on limiter on the sum).

Speech

“There isn't one specifically trained for anime, try your luck with the current available models”

The list by Musicalman (mostly from before the Gabox models release, check newer vocal models too)

“Any vocal model in the past few years should work for speech separation. My favorites at the moment are:

- MDX23C Inst-Voc HQ

- other similar MDX models for least aggressive, but bleedy, only really useful for denoising

- Unwa's Mel-Roformer big beta 4 or beta 5e vocal models - for less bleed. Atm, 5e is my go-to as it sounds less filtered.

~ I've heard people praise BS-Roformers a lot, haven't really tested those much, though.

- Becruily's vocal model [that old one] can also be better at SFX separation, but can overestimate reverb in the vocal stem sometimes.

- Mel-Roformer Karaoke by viperx and aufr33 - for more aggressive separation (removes a bit more SFX)

- And the most aggressive are Bandit models and the DNR v3 models on MVSEP, though they tend to be a bit too aggressive for my taste, so I only use them selectively.

This is just my own opinions though, subject to change at a moment's notice lol”

[you’ll find them in SFX section]

- clearvoice - it's a set of speech enhancement/separation models. My favorite model of the set is MossFormer2_SE_48K. Its dialog extraction seems to be similar to Bandit v2, though clearervoice sounds fuller to me, and separation is usually a bit better. Might be especially good in an ensemble with Bandit or vocal sep models eg. unwa, gabox etc.

- BS-Roformer 2025.06 on MVSEP - “handle speech very well. Most models get confused by stuff like birds chirping (they put it in the vocal stem), but this model keeps them out of the vocal stem way more than most. I love it!”

- iZotope Dialogue Isolate (in RX 10, and esp. 11, 12)

- iZotope RX 12 Scene Rebalance (paid; dialogue, music and effects stems)

- Acon Digital Extract:Dialogue 2

- Accentize VoiceGate and

- Acon Extract Dialogue

(to debleed dialogues from SFX models)

See also:

- Various speakers isolation 

- Harmonies

- Two singers isolation 

- Karaoke 

Can’t find model link?

- Results containing models in e.g. #946 (e.g. 406, 427, 438) or other ensembles mentioned above, still have public models available in UVR, but you can access them by entering the download/vip code in UVR, so more models will show up

You cannot use VIP code on older beta UVR Roformer patches (updates), then to use any other VIP model with Roformers (e.g. D1581), you need to install the stable 5.6 from official GH repo, download the model, and update the installation with the old Roformer patch afterwards if you need such version

- Be aware that MDX23C Inst Voc HQ2 is not accessible in beta Roformer patch when VIP code is inserted. You need to download the model file manually, and paste into models\MDX_Net_Models folder.

(Config is detected automatically, as it uses existing model_2_stem_full_band_8k config - the same as for Inst Voc HQ)

- UVR Denoise non-lite model disappeared from Download Center. Here it is: https://github.com/TRvlvr/model_repo/releases/download/all_public_uvr_models/UVR-DeNoise.pth

- You cannot use some models from x-minus/uvronline.app or used on MVSEP, e.g. used for Ensemble of 4 and 8 models in UVR, as they contain models not available in UVR, and not available for download. You can only perform manual ensemble of single models processed by MVSEP or x-minus, in UVR, but it will not give the same result as ensemble on MVSEP, as it uses code more similar to MDX23 Colab code, so sometimes weighted ensemble instead of. e.g. avg spec (don’t confuse with MDX23C arch models).

- E.g. for 16.10.23, “MVSep Ensemble of 4” consists of 1648 previous epoch (maybe later updated to 16.66), VitLarge, and Demucs 2023 Vocals and beside the first, none of these models work in UVR, even if downloaded manually (plus VitLarge arch is not supported in UVR at all). Currently, there are various ensembles to choose from on MVSEP.

As for 4/8 models ensemble on MVSEP - they’re all only for premium users, as many resources and models are being used to output these results

- 1648 on MVSEP is MDX23C HQ1 model (a.k.a. 8K FFT)

- SYHFT V4 and V5 beta by SYH99999 were never publicly released.

V5 Beta is only on x-minus.pro/uvronline and got deleted from the main models view, but might be still accessible via the following links:

https://uvronline.app/ai?hp&test (premium)

https://uvronline.app/ai?test (free)

Models repository

All arch models with configs in one place:

https://huggingface.co/noblebarkrr/mvsepless_resources/tree/main

https://huggingface.co/miercolesv/MVSEP-Central-Backup/tree/main

(it was 5 months older at the moment of writing, but contained a few more models)

https://huggingface.co/Politrees/UVR_resources/tree/main/models

(it’s also the same 5 months older at the moment)

All single models available in UVR 5’s download center
- repository backup as separate links (excluding VIP models, which contain offline links after decrypting):

https://github.com/TRvlvr/model_repo/releases/tag/all_public_uvr_models

configs (non-VR/Demucs ones)

All of publicly available MVSEP models (including checkpoints just for further training):
https://github.com/ZFTurbo/Music-Source-Separation-Training/releases

(refer to the list of models in this document for descriptions of the best models)

Alternative model links lists (some can be offline):

https://bascurtiz.x10.mx/models-checkpoint-config-urls.html 

https://github.com/SiftedSand/MusicSepGUI/blob/main/models.json

https://huggingface.co/spaces/TheStinger/UVR5_UI/blob/main/assets/models.json

Some of the older UVR5 GUI models described in this guide can be downloaded via expansion packs:

https://github.com/Anjok07/ultimatevocalremovergui/releases/download/v5.3.0/v5_model_expansion_pack.zip

https://github.com/Anjok07/ultimatevocalremovergui/releases/download/v5.3.0/models.zip

https://github.com/Anjok07/ultimatevocalremovergui/releases/download/v4.0.1/models.zip

Some of the models used by KaraFan:

https://github.com/Eddycrack864/KaraFan/releases/tag/karafan_models

MDX23C HQ 2

https://github.com/deton24/Colab-for-new-MDX_UVR_models/releases/download/v1.0.0/MDX23C-8KFFT-InstVoc_HQ_2.ckpt

427:

https://drive.google.com/drive/folders/16sEox9Z_rGTngFUtJceQ63O5S9hhjjDk?usp=drive_link (just in case)

Copy it to Ultimate Vocal Remover\models\MDX_Net_Models and rename the model name to: UVR-MDX-NET_Main_427

Some direct links


VOCALS-InstVocHQ

Config: https://raw.githubusercontent.com/ZFTurbo/Music-Source-Separation-Training/main/configs/config_vocals_mdx23c.yaml

Checkpoint: https://github.com/ZFTurbo/Music-Source-Separation-Training/releases/download/v1.0.0/model_vocals_mdx23c_sdr_10.17.ckpt

VOCALS-MelBand-Roformer (by KimberleyJSN)

Config: https://raw.githubusercontent.com/ZFTurbo/Music-Source-Separation-Training/main/configs/KimberleyJensen/config_vocals_mel_band_roformer_kj.yaml

Checkpoint: https://huggingface.co/KimberleyJSN/melbandroformer/resolve/main/MelBandRoformer.ckpt

VOCALS-BS-Roformer_1297 (by viperx)

Config: https://raw.githubusercontent.com/ZFTurbo/Music-Source-Separation-Training/main/configs/viperx/model_bs_roformer_ep_317_sdr_12.9755.yaml

Checkpoint: https://github.com/TRvlvr/model_repo/releases/download/all_public_uvr_models/model_bs_roformer_ep_317_sdr_12.9755.ckpt

VOCALS-BS-Roformer_1296 (by viperx)

Config: https://raw.githubusercontent.com/TRvlvr/application_data/main/mdx_model_data/mdx_c_configs/model_bs_roformer_ep_368_sdr_12.9628.yaml

Checkpoint: https://github.com/TRvlvr/model_repo/releases/download/all_public_uvr_models/model_bs_roformer_ep_368_sdr_12.9628.ckpt

KARAOKE-MelBand-Roformer (by aufr33 & viperx)

Config: https://huggingface.co/jarredou/aufr33-viperx-karaoke-melroformer-model/resolve/main/config_mel_band_roformer_karaoke.yaml (dead)

Checkpoint: https://huggingface.co/jarredou/aufr33-viperx-karaoke-melroformer-model/resolve/main/mel_band_roformer_karaoke_aufr33_viperx_sdr_10.1956.ckpt (dead)

OTHER-BS-Roformer_1053 (by viperx)

Config: https://raw.githubusercontent.com/TRvlvr/application_data/main/mdx_model_data/mdx_c_configs/model_bs_roformer_ep_937_sdr_10.5309.yaml

Checkpoint: https://github.com/TRvlvr/model_repo/releases/download/all_public_uvr_models/model_bs_roformer_ep_937_sdr_10.5309.ckpt

CROWD-REMOVAL-MelBand-Roformer (by aufr33)

Config: https://github.com/ZFTurbo/Music-Source-Separation-Training/releases/download/v.1.0.4/model_mel_band_roformer_crowd.yaml

Checkpoint: https://github.com/ZFTurbo/Music-Source-Separation-Training/releases/download/v.1.0.4/mel_band_roformer_crowd_aufr33_viperx_sdr_8.7144.ckpt

VOCALS-VitLarge23 (by ZFTurbo)

Config: https://raw.githubusercontent.com/ZFTurbo/Music-Source-Separation-Training/refs/heads/main/configs/config_vocals_segm_models.yaml

Checkpoint: https://github.com/ZFTurbo/Music-Source-Separation-Training/releases/download/v1.0.0/model_vocals_segm_models_sdr_9.77.ckpt

CINEMATIC-BandIt_Plus (by kwatcharasupat)

Config: https://github.com/ZFTurbo/Music-Source-Separation-Training/releases/download/v.1.0.3/config_dnr_bandit_bsrnn_multi_mus64.yaml

Checkpoint: https://github.com/ZFTurbo/Music-Source-Separation-Training/releases/download/v.1.0.3/model_bandit_plus_dnr_sdr_11.47.chpt

DRUMSEP-MDX23C_DrumSep_6stem (by aufr33 & jarredou)

Config: https://huggingface.co/noblebarkrr/mvsepless_resources/blob/main/mdx23c/mdx23c_drumsep_6stem_aufr33_jarredou_config.yaml 

Checkpoint: https://huggingface.co/noblebarkrr/mvsepless_resources/blob/main/mdx23c/mdx23c_drumsep_6stem_aufr33_jarredou.ckpt

4STEMS-SCNet_MUSDB18 (by starrytong)

Config: https://github.com/ZFTurbo/Music-Source-Separation-Training/releases/download/v.1.0.6/config_musdb18_scnet.yaml

Checkpoint: https://github.com/ZFTurbo/Music-Source-Separation-Training/releases/download/v.1.0.6/scnet_checkpoint_musdb18.ckpt

For a more recent list see this MK Colab and its cells containing all the links there too.

____

Why not to use more than 4-5 models for ensemble in UVR - click

____

Other models

- GSEP AI - new model called “Vocal Remover”, and old instrumental, vocal, 4-6 stem model (it applies additional denoiser for 4/6 stems) - piano and guitar (free). As for 2 stems, it gives very good instrumentals for songs with very loud and harsh vocals and a bit lo-fi hip-hop beats, as it can remove vocals very aggressively. Sometimes even more than HQ_3. The new model might be good at removing SFX (instrumental stem is the old model).

In specific cases (can have more vocal residues in instrumentals vs HQ_3 at times - less in jarredou's Colab):

- original MDX23 by ZFTurbo (only this OG version of MDX23 still works in the offline app, min. 8GB Nvidia card required [6GB with specific parameters]) - sounds very clean though, and not that muddy like inst MDX models, in this means, comparable with even VR arch or better (because of much less vocal residues).

- Demucs_ft model (both 3 stems to mix in e.g. Audacity for instrumental) / sometimes 6s model gives better results, or in very specific cases when vocals are easy to filter out - even the old 4 stem mdx_extra model (but SDR wise full band MDX 292 is already better than even ft model). The 6s model is worth checking with shifts 20.

Might be still usable in some specific cases, despite the fact that MDX23 uses demucs_ft and other models combined.

- VR models settings + VR-only ensemble settings (generally deprecated, but sometimes more clarity vs MDX v1, though frequently more vocal residues. Some people still uses it e.g. for some rock, when it can still can give better results than other models, and also for fun dubs, but for it if you have two language tracks of the same movie, you can test out Similarity Extractor instead, but Audacity center extraction works better than that linked Colab)

- Alternatively, you can consider using narrowband Kim other ft model with fullband model settings parameters in this or the new HV Colab instead. Useful in some specific parts of songs like chorus, where there are still no persistent vocal residues using this method (clearer results than even Max-Spec) or e.g. MDX23 still doesn't give you enough clarity in such places to maybe merge fragments manually of results from different models.

Paid

- Audioshake (non-copyrighted music only, can be more aggressive than above and pickup some lo-fi vocals where other fails [a bit in manner of HQ models])

How to bypass the non-copyright music restriction (1, 2).

"They also reserve themselves the right to keep your money and not let you download the song you split if they discover that you are using a commercially released song and that you don't have the rights to it." but generally we didn't have such a case with slowed down songs (otherwise they might not pass anyway)

4 stems might be better at times then Demucs ft model.

- Dango.AI (a.k.a. tuanziai.com) free 30 seconds samples; can be the most aggressive for instrumentals vs, e.g. inst 3, tested on Childish Gambino - Algorithm). Since then, models/arch were updated and instrumentals in 9.0 seem to be the cleanest or the closest to original instrumentals for 12.08.23 at least in some cases (despite low SDR).

> If you care only about specific snippet in a song, then since 30 second samples to separate are taken randomly from a whole song, to have specific fragment separated, you can copy the same fragment over and over to make a full-length track of it, and it will eventually pick up a whole snippet for separation.

X Uploading snippet shorter than or exactly 30 seconds will not result in the whole fragment being processed from the beginning to the ending.

>Sometimes using other devices or virtual machine in addition to incognito/VPN/new email might even be necessary to reset free credits. It's pretty persistent.

https://tuanziai.com/encouragement

Here you might get 30 free points (for 2 samples) and 60 paid points (for 1 full songs) "easily".

>>>

Everything else for 2 or 4 stems than above is worse for separation tasks:

Lalal, RipX (although now it uses some UVR models (?), Demix, RX Editor 8-11, Spleeter and its online derivatives.

Debleeding/cleaning vocals/instrumentals/inverts

1) For constant vocal shells/buzzing of Roformer instrumental models (don't confuse with residues)

- Phase fixer (for instrumentals - uses phase of inst vocal model result on inst model result)

- MVSep PolarFormer (for “pulling out some of the vocal noise leftover by the BS-Rofo high fullness models, without muddying the sound” - CC Karaoke).

- Gabox Mel denoise/debleed model (when used on the mixture) | yaml | Colab | SESA Colab v3 - for fullness models (tested on v5n) - it can't remove the vocal residues

- Lew Apollo Universal on hi-hat and snare from drumsep 6 stems model by jarredou from drums model from e.g. flowersv10 inst model (I've noticed that most of the residual "Fuzz" is usually trapped in the snare or cymbals. - CC Karaoke)

- Weighting manually Drumsep model results with SAM-Audio-prompted drumsep in L/R because SAM is mono (to get drumsep from drums too, but also with SAM) - to get cleaner results of snare and hi-hat (“ I lower the original snare/HH drum stem by maybe 5 or 6dB, and the SAM drum stem by 1 or 2dB... basically the idea is to make sure the volume of the drums in the mix is about the same as it was originally; but lowering out the sound of the noise while keeping as much of the original sound as possible.” more)

- Denoise models (e.g. Aufr33’s Mel is less aggressive and effective than phase fixer, UVR Denoise more aggressive and destructive, you can try out Mel on mixture first)

- DEBLEED-MelBand-Roformer (by unwa/97chris) model | yaml | Colab

(it can work with e.g. inst v1 and its noise, or with v1e, or even MVSEP Karaoke BS-Roformer instrumental stem for “very clean and full” result. Also, sometimes the debleed model can remove some bleed also after using phase fixer)

2) For crossbleeding/residues/leftovers/cleaning inverts

- Mel-Roformer invert-clean by becruily | Colab (for cleaning acapella inverts [so when you use e.g. official instrumental with mixture inverted in order to get vocals and it gives some residues]

“Most regular models don't do the job (...) It removes peaks/pops, clicks, instrumental residue (...)

I didn't focus on things like fullness” - becruily

“Cool model, it's a little muddy on loud/clipped vocals but otherwise it works great (...)

it does work well for leftovers” - gabiigh

“It works great (...) the result is fantastic!” - billieoconnell.

“this has to be [becruily's] best model ever, so useful, sounds so great too... I'm shocked it truly is so good”

Consider using it with overlap 4 - gilliaan)

- Colab auto-cleaning by billieoconnell - uses Deux result acapella using the invert-clean model above.
“Most of the acapellas come out fine, but in some, when I run them through the AI, there's a drop in the bass where the drums hit, making the vocals sound like they're on a bass drum. (...) when I run it through a model on other websites or similar platforms, the result is great.”

- MVSEP BS-Roformer 2025.07 (e.g. “after making the deux hyperace ensemble ensemble, at least in my cases, works good to remove any small residues that are left in quieter/airy parts of songs” - arxynr)

- VR’s HP2-4BAND-3090_4band_arch-500m_1 in HV Colab a.k.a. 9_HP2-UVR in UVR, aggressiveness 1.0 on MVSEP/Colab, 100 in UVR (works for vocal residues from inst Roformers, “can take some vocals residues that even those bleed suppressor models can't (...) still works somehow for cleaning that even BS-Roformer 2025.07 couldn't”, helps with some vocal reverb and some choir vocals and old anime songs and regular nu metal songs - mohammedmehditber)

- MVSEP Choir (same as with 9_HP2, but esp. for choir residues - mohammedmehditber)

- UVR Denoise Standard

- UVR Denoise model with 0.1-0.2 aggressiveness (for debleeding vocals from instruments)

- UVR-MDX-NET Crowd HQ 1 (“sometimes fixes screaming Roformers can't, but it has more muddiness and bleeding in general”)


- Roformers: Mel 2024.10 on MVSEP a.k.a. Bas Curtiz FT (for vocals debleeding)
- Or earlier Kim Mel-Roformer model (for vocals; if other model was prev. used)

- Unwa inst v1 (to clean-up vocals from Mel 2024.10 model)

- unwa’s inst v1/e/2 (for OG instrumentals with bleeding [better than Dango for it])
- unwa big beta 5 (“my go-to clean-up artifacts model after phase inverting master + official instrumental”)

- Unwa BS Revive 3e (although Revive 2 has bigger bleedless metric)

- voc_fv6 (for vocal inverts - ezequielcasas)

- Mel avuew’s v2 de-reverb or unwa’s BS Large, MDX-Net HQ_5
(“great for cleaning acapellas from bits of instrumentals”)
- syftbeta 5 on x-minus.pro (probably still available with this link for premium, and for free)

- Ensemble of BSRoformer-Viperx1297, Unwa BSRoformer-LargeV1, unwa_ft2_bleedless, mel_band_roformer_vocals_becruily, Gabox voc_fv4 on Average/Average
(good for cleaning inverts - AG89)

- Ensemble of big beta6x, revive2, unwa ft2 bleedless
(for cleaning instrumental inverts - AG89)

- Or just experiment with other models with the highest bleedless metric (instrumental | vocals)

- SFX models are more aggressive than vocal models (not tested for this purpose yet)

- yxlllc’s harmonic noise separation VR model (“Very good at further removing the noise from a dereverbed vocal, yet it is mono. (...) It did have two channels on export but both of them have audible information except one has only some noise that can be easily removed, then mix the other channel up to stereo” - mohammedmehditber. Maybe attached CLI code will have better mono model handling.
Can be used in UVR: rename “model” to some model name, and pt extension to pth, then use Install model option and set config settings to: VR 5.1, 32/128, 1band_sr44100_hl512)

- Sam Audio by Meta - Iirc, on mid side processed files, it got rid of some noise and knocking when separating SFX and instrumental with a prompt “music” and later with a prompt “knocking” on the result file “and selected the part at 0:35 and 0:36 for span prompting” - nicov_na

- RX10 De-bleed feature for instrumentals (video)

(older methods)

- Gabox Mel denoise/debleed model | yaml | Colab | SESA Colab v3 - for noise from fullness models (tested on v5n) - it can't remove the vocal residues
- Mel denoise (iirc Aufr33’s) - that model “removed some of the faint vocals that even the bleed suppressor didn't manage to filter out” before”. Try out denoising on a mixture first, then use the model.

- (for saxophone bleeding) ~“1. Take the original song in FLAC or WAV 2. Use MVSEP Saxophone 3. Take the other stem from it - there should be everything else and most of the sax should be gone (for me, there was a small part left) 4. Use Unwa Big Beta 5 on it (so Other-> uvronline/xminus/Colab Unwa Big Beta 5) - then vocals should be very clean no sax bleeding” cali_tay98

- (“In case there's any wind instruments that could potentially bleed into the vocals”)
MVSep Wind

- (when “models don't pick up the noise)

gently bring back a bit of the original music/instrumental on the inverted track and use AI again.

By gently, I mean no more than 6 dB“ - becruily

- (“If your result have "vocal chops" left in the instrumental separation and no models could remove them completely)

then it's likely MDX HQ_5 or VitLarge23 v2 will fix it” dca

- Acon Digital DeBleed:Drums “it's just an advanced gate. It doesn't remove bleed when it's overlapping the wanted audio (we can still hear hihat/cymbals on snare with the plugin enabled in their demo)” - jaredou.

- Audio-Bleeding-Removal - by its-rajesh

- Try out some L/R inverting, try out to separate multiple times to get rid of some vocal pop-ins like this (fix for ~"ah ha hah ah" vocal residues)

- Accentize VoiceGate and Acon Extract Dialogue

(to debleed dialogues from SFX models)

Older de-bleeding models

- Ripple (defunct) “AWESOME to use after inverting songs with the official instrumental”

Instrumentals can be also further cleaned with Ripple, and then with Bandlab Splitter

(Roformer models may potentially replace Ripple models in that matter now)

- Top ensemble in UVR5 (starting from point 0d)

- GSEP - very minor difference between both for cleaning vocals (maybe GSEP is better by a pinch).

You can try separating e.g. vocal result double using different settings (e.g. voc_ft>kim vocal 2)

- MVSEP 11.50 Ensemble (the least amount of bleeding in inst. separations at least)

- MDX23 jarredou's fork Colab (maybe this version at first)

- use voc_ft model on the result you got (so separate twice if you already used that model)

Cleaning inverts means - cleaning up residues - e.g. left by the instrumental after an imperfect phase cancellation, e.g. when audio is lossy, or maybe even not from the same mixing session

Aligning for bad inverts

- "Utagoe bruteforces alignment every few ms or so to make sure it's aligned in the case that you're trying to get the instrumental of a song that was on [e.g.] vinyl."

[The previous] UVR's align tool is handy just for digital recordings… [so those] which don't suffer from that [issue] at all."

Utagoe will not fix fluctuating speed issues, only the constant ones.

- Anjok already "cracked" how that specific Utagoe feature works, and introduced it to UVR.

“Updated "Align Tool" [is] to align inputs with timing variations, like Utagoe.”

Sometimes it can give even better results than utagoe e.g. when inverting “a full track and instrumental while automatically matching the waveform”

- “Some users had good results with Auto Align Post 2 plugin to resync tracks before inverting them.”

- For problematic inverts, you can also try out azimuth correction in e.g. iZotope RX.

Notes:

1) Make sure your stems align in DAW, check across the whole song

2) If stems become suddenly misaligned in certain parts of song, use utagoe or align inputs in UVR>Tools

3) Make sure you don't use separated stems from service like MVSep with volume normalized, so the output will be quieter than the OG mixture/input file (or use 32-bit float output for premium)

4) Make sure you don't use lossy separation output - it might start in a different place vs mixture

5) Make sure one of your stems doesn't have a different sample rate

6) Some instrumentals/acapellas derive from different mixing/mastering session or weren't properly exported (e.g. without stems sidechaining or dedicated option but instead e.g. with just stems soloed) so it won't invert with the OG song, or there might be some discrepancies in various places of the song

7) The sums of stems of multistem models like deux or SW won't invert with the OG mixture/input file (e.g. splifft by undef13 compensates fixes it at least for the SW)

Declicking vocals

- BS-Rofomer vocal model (iirc; “Can also fix hard clicks in vocals. It is even better than RX in this, but still there is a tiny wave fade residue in some cases”)

- Kim vocal first and then separate with instrumental model (e.g. HQ_3 or 4). You might want to perform additional separation steps to clean up the vocal from instrumental residues first, and invert it manually to get cleaner instrumental to separate with instrumental model to get rid of vocal residues

Removing metronome (e.g. from a mixture or vocals)

- Use a good instrumental model so you will be left with metronome + vocals in one stem, then use a drums model - “Then the drum trick worked better but still not very good, a regular extraction worked better this time though!” - brianghost

Removing bleeding of hi-hats in vocals

- Kim Mel-Roformer model

- Use MelBand RoFormer v2 on MVSEP (e.g. after using MDX23C Inst HQ)

Bleeding in other stems

- RipX Stem cleanup feature (possibly)

- SpectraLayers 10 (eliminates lots of bleeding and noise from MDX23 Colab ensembles)

"You debleed the layer to debleed from using the debleed source. Results vary. Usually it's better to debleed using Unmix and then moving the bleed to where it belongs" Sam Hocking

Video

Bleeding of claps in vocals

- Reverse polarity and/or remove DC offset of the input file

- KaraFan (for general drum artefacts, but it doesn’t work well for inverts, try out modded preset 5 here)

- Remove drums with e.g. demucs_ft first, then separate the drumless mixture from inversion, consider slowing down your input without stretching to 0.75.

- Ensemble of old VR models (search the names here for UVR names equivalents)

- Kim Vocal 2 (but it has a cutoff and creates a lot of noise in the output)

- Denoise model with 0.1-0.2 aggressiveness

- Sam Hocking method

- possibly de-crowd models

Bleeding of guitars/winds/synths in vocals

- BVE (Karaoke) models

Fixing overlapped/misrecognized stems

- Spectralayer's 9/+ Cast & Mold

Fixing low-end rumble

- Spectral Editing:

a) RX Editor’s brush (video by Bas)

b) Audacity (image) - “you can, just barely”

Potential alternatives for spectral painting:

Free: ISSE, Ampter, Filter-Artist, AudioPaint

Paid: RipX, SpectraLayers, Melodyne, prob. Revoice Pro 5

Cleaning the white noise/sizzle from vocals

(from e.g. Roformer models)

- big beta 5e (if you have vocals “really great” - gilliaan)

- MDX23C model (e.g. the latest on MVSEP or HQ in UVR)

- aufr33 denoise model (generally meant also for white noise; doesn’t always work, can pick vocals - gustownis, “when it works its one of the best ones” - cyclorana)

_______

Debleeding guide by Bas Curtis (other methods, e.g. Audacity)

Denoising and dereverberation/apps later below.

See also “Vinyl noise/white noise” from the end of the list.

_______

How to check whether a model in UVR5 GUI is vocal or instrumental?
  • Read carefully the models list above - they're categorized
  • If you want to experiment with other models:

The moment you see "Instrumental" on top (and "Vocal" below) in the list where GPU conversion is mentioned, you know it's an instrumental model.

When it flips the sequence, so Vocal on top, you know it's a vocal model.

Same happens for MDX and VR archs.

  • “Be aware that MDX23C/MDXv3 models can be multisource - it depends on the training, so it can be only vocals, or only instrumental, or vocals+instrumental, or vocals+drums+bass+other (like baseline models are), or whatever else.
  • You can know it looking at the config file of the model, for example InstVocHQ,

https://github.com/Anjok07/ultimatevocalremovergui/blob/master/models/MDX_Net_Models/model_data/mdx_c_configs/model_2_stem_full_band_8k.yaml

Seeing by the instruments line above, D1581 and InstVocHQ models are instrumental+vocal.

Config for the rest of the models:

https://github.com/Anjok07/ultimatevocalremovergui/blob/master/models/MDX_Net_Models/model_data/model_data.json 

(decoded hashes)

____________________________________________________________________


Keeping only backing vocals in a song (lead vocal extractor) a.k.a.: 
>Karaoke

1. You might want to use a good vocal model as a preprocessor to use with the models below (if MVSEP/x-minus don’t do it already), but sometimes it can degrade the quality, so experiment both ways.

2. Optionally you may also de-reverb vocals with a good model/plugin first before proceeding, but only if the specific model/method doesn’t filter out a lot of BGVs in your vocals.

3. As a preprocessor for all-vocals model, experiment with different panning settings using e.g. A1StereoControl plugin beforehand (might work as counterpart of the panning trick used on uvronline.app for better LV/BV recognizability).

Also check here for backing vocal extraction.

- Mel-Roformer small_karaoke_gaboxauf by Gabox and Aufr33 | yaml | Colab | Metrics

For RTX 5000 patch and torch._dynamo error, edit “use_torch_checkpoint: True” to False (or delete the line).

“it sounds much more fuller than other karaoke models I use [anvuew bs roformer, bs karaoke gabox, and becruily karaoke] and clears out bg vocal much more accurately” - baptizedinfear

“Great! got to hear vocals I could not hear before.” - makesomenoiseyuh

“it's not great with duets btw” - Gabox

Lowering chunk_size to 352800 makes it muddier, but less noise, “maybe in between would be nice” - rainboomdash

- Anvuew’s Karaoke BS-Roformer model | metrics | MSST | MVSEP | Colab

(Lead vocals/backing with instrumental)

For RTX 5000 patch and torch._dynamo error, edit “use_torch_checkpoint: True” to False (or delete the line).

“extracts lead vocals a bit better than karaoke becruily frazer, and in some parts, the lead vocals from karaoke anvuew still sound brighter compared to karaoke becruily frazer, which sounds a bit more compressed. oh, and for some reason, the becruily frazer model doesn’t detect vocals with radio effects, while anvuew’s model handles them just fine” - neoculture

“lead vocals leak into instrumental (...) Mel Becruily and Frazer’s BS don’t have this problem”

In that case, maybe “isolate the acapella first in almost all cases of using a karaoke model” or use the model below instead.

- BS-Roformer Karaoke model by becruily & frazer | metrics | MVSEP | uvronline

(Lead vocals/backing with instrumental)

Make sure you don’t have the option “Vocals only” checked in UVR.

8GB VRAM users of AMD and Intel ARC GPUs need to use 160000 chunk_size, or the separation will be very slow.

“After dozens of tests I can tell this (...) is the best (better harmony detection, better differentiation between LVs and BVs, sounds fuller, less background Roformer bleed, better uncommon panning handling etc) (...) I noticed "lead vocal panning" works really well” - dca

“It also can detect the double vocals” - black_as_night

It works the best for some previously difficult songs. Aufr33 and viperx model seems more consistent, but the new BS is still the best in overall - Musicalman

“My OG Mel also catches some of the FX/drums, I guess quite a difficult one due to how it's mixed” - Becruily

“It does do better on mono than previous, sometimes confuses which voice should be the lead, but all models do that on mono in the exact use-case I normally test” - Dry Paint Dealer Undr

“[In] no way inferior to the viperx (Play da Segunda)” - fabio5284

“The new karaoke model doesn't actually differentiate between LVs & BVs and there's some lead vocal bleeding in the instrumental stem” - scdxtherevolution

VS the newer BS-Roformer MVSEP team model: “sound isn't as clear, but it does an infinitely better job at telling lead/bgv apart”

Becruily:

“I want to remind something regarding my (and the frazer) models

they're made to separate true lead vocals, meaning either all of the main singer's vocals, or if it's multiple singers - theirs too

this means if the main singer has stuff like adlibs on top of the main vocals, these are considered lead vocals too - they go together

if there are multiple singers singing on top of each other, including harmonise each other, and if there are additional background vocals behind those - all the singers will be separated as one main lead vocal, leaving only the true background vocals”

“Noticed it a couple days ago. I've had fantastic results with it so far. Much MUCH better at holding the 'S' & 'T' sounds than the Rofo oke (for backing vox). Generally seems to provide fuller results... but also the typical 'ghost' residue from the main vox can end up in the backing vox sometimes, but it's usually not enough to be an issue. I won't go so far as so say that it's replacing the other backing vox models for me entirely... but it feels like the best of both worlds that Rofo and UVR2 provide.” - CC Karaoke

Tips for the model

- “I had success by setting the BS Rofo Karaoke model to 100% Right, and then taking the 'Other' result and reprocessing it at 100% Left to get the backing vocals cleanly out.

Curious note on something I've never tried or had to do before but it has worked wonderfully;

I'm isolating the backing vocals on Radiohead's Let Down.” - CC Karaoke

- If you use e.g. vocfv7beta1 as preprocessor for the model, you may get some quieter backing vocals better - Rainboom Dash

- If you have “state_dict” error in MSST, edit it to false in line 790 in separate.py as following: model.load_state_dict(checkpoint, strict=False

If it's still the same, update MSST and retry the step.

- MVSEP’s BS Roformer by MVSep Team (SDR: 10.41)
under option "MVSep MelBand Karaoke (lead/back vocals)", metrics. Might be a fine-tune. Use the option “extract vocals first”.
(“In contrast with other Karaoke models, it returns 3 stems: "lead", "back" and "instrumental".)

“If I had to compare it to any of the models, it is similar to the frazer and becruily models. Sometimes it does not detect the lead vocals especially if there's some heavy hard panning, but when it does, there is almost no bleed, and it works very well with heavy harmonies in mono from what I tested.” - smilewasfound

“becruily & frazer is better a little when the main voice is stereo” - daylightgay

Less clarity thean the models above - wancitte

“On tracks I tested, harmony preservation was better in becruily & frazer (...) the new model isn't worse, I ended up finding examples like Chan Chan by Buena Vista Social Club or The Way I Are by Timbaland where it is better than the previous kar model. The thing is, with the Kar models, it's just track per track. Difficult to find a model for batch processing as it's really different from one track to another” - dca100fb8

As for MVsep Team: “It’s the only model that combines the lead vocal doubles with the lead vocals stem. It’s far more useful for dissecting harmonies on songs with vocal doubles like Backstreet Boys” - heuheu

“I also found the new model to not keep some BGVs, mainly mono/low octave ones, despite higher SDR” - becruily

“I think I've found a solution for people who don't like the new model.

If you put an audio file through the karaoke model and then put the lead vocal result through that, it usually picks up doubles.

Which you can then put in your BGV stem if you'd like” - dynamic64

- MVSep Choir (works for e.g. choirs buried under main vocal in e.g. pop music)

Ensembles

(MVSep-exclusive models and paywalled)

- Becruily 124 fullness phase fixed + Karaoke MVSep Team (Max FFT): https://mvsep.com/presets?hash=eR69hwVtSsAdBtO8

- Becruily 124 fullness phase fixed + Karaoke Frazer/becruily (Max FFT): https://mvsep.com/presets?hash=s78ySK0F4fQ9GfuK

(free/open ensembles)

#1 BS anvuew kar + Mel Inst_GaboxFv9 (Max Spec)

#2 BS anvuew kar + Mel deux (Max Spec)

#3 BS frazer/becruily kar + Mel Inst_GaboxFv9 (Max Spec)

#4 BS frazer/becruily kar + Mel deux (Max Spec

(Older ensembles)

- Ensemble of frazer becruily karaoke and anvuew karaoke (max spec)

- Ensemble of Mel v1e + BS Karaoke MVSep Team with extract vocals first option, Max Spec (using BS 2025.07 as reference and 200/200 for the values)

“It's very aggressive values cuz v1e is noisy, and it works quite well”, the best ensemble for now (dca100fb8)

- Ensemble of BS Roformer Karaoke by anvuew + BS Resurrection Inst (aka "unwa high instrum fullness" on mvsep) + phase fix (using BS 2025.07 as reference), older, more crossbleeding - (dca100fb8)

- "Ensemble of 3 models "MVSep + Gabox + frazer/becruily" gives 10.6 SDR on the leaderboard. I didn't upload it yet, but I had local testing.” - ZFTurbo

- Ensemble of :

BS-Roformer Karaoke Frazer & Becruily + BS-Roformer Karaoke Anvuew (avg/avg)

(metrics)

- MVSEP fusion model of Gabox Karaoke and Aufr33/viperx
(tend to confuse BV/LV more than single models)

- Gonzaluigi Karaoke fusion models (standard | aggresive | yaml) -

also confuses BV/LV more

Other single models

- MVSep MelBand Karaoke (lead/back vocals) SCNet XL IHF by becruily (SDR: 9.53)

Worse SDR than the top performing Roformers, but works best in busy mix scenarios, and when Mel-Roformer models fail. Generally bleedier arch. To fix the bleed in the back-instrum stem, use “Extract vocals first”, but “I noticed a pattern that if you hear the lead vocals in the back-instrum track already (SCNet bleed), don't try to use Extract vocals first because there will be even more lead vocal bleed” - dca (iirc it uses the biggest SDR BS-Roformer vocal model as preprocessor).

“Separates lead vocals better than Mel-Roformer karaoke becruily. It's not perfectly clean, sometimes a bit of the backing vocals slips through, but for now, SCNet karaoke model [is] still the most reliable for lead vocals separation (imo)” - neoculture

(it was written probably before Frazer and anvuew models)

Is it good?

“If you ask me - yes. In some cases it might sound more natural for the BGVs because by design it will keep a small amount of lead leftover on top of the BGV which perceptually sounds much smoother/nicer than the muddy Roformer models. Also sometimes just understands BGVs better, but that model might be undertrained. Could be more useful for BVE than karaoke, but who knows.” - Becruily (March 2026)

- (x-minus/uvronline) Lead and Backing vocal separator (in “Extract Backing vocals>Mel-Roformer Lead/Back)

It uses big beta 5e model as preprocessor for becruily Mel Karaoke model  

“In fact, the big beta 5e Deux model is run after becruily Mel Karaoke” - Aufr33, Gabox (so you don’t need the additional step to use this separator), plus it also allows controlling option for lead vocal panning like for BVE v2 (It’s to “to "tell" the AI ​​where the main vocals are located (how they are mixed)”. “Doesn't even need Lead vocal panning a lot of the time, [the] ability to recognize what is LV and what is BV [is] impressive” - dca).

“The new separator is available in the free version, however, due to its resource intensity, only the first minute of the song will be processed.” if you don’t have a premium.

The difference from using single becruily Kar model is that, here, “you get the third track, backing vocals.”.

- Becruily: “Probably too resource-intensive, but you could try adding demudders to each step

1) Karaoke model + demudding

2) Separate vocals of BGV + demudidng

But not sure how much noise this will bring

(Or even a 50:50 ensemble with BVE OG)”

- Mel-Roformer Karaoke by becruily model file | Colab | MVSEP | x-minus

(back/main vox/instrumental - 3 stems)

Use voc_fv4 vocal model before running it (less bleeding than voc_fv6) or:

“first extract the vocals with a fullness model [it was Mel becruily vocal back then] and combine the results with a fullness instrumental model.” - becruily

“It's a dual model, trained for both vocals and instrumental. It sounds fuller + understands better what is lead and background vocal,
and to me, it is better than any other karaoke model.

Important note: This is not a duet or male/female model. If 2 singers are singing simultaneously + background vocals, it will count both singers as lead vocals. The model strictly keeps only actual background vocals. The same goes for "adlibs" such as high notes or other overlapping lead vocals.

The model is not foolproof. Some songs might not sound that much improved compared to others. It's very hard to find a dataset for this kind of task.

Compared to Aufr33’s Melband model below, it can achieve e.g. cleaner pronunciation in some songs (examples).

“Better than Mel Kar, UVR BVE v2, lalal.ai, Dango...”

“For sure better than [the] older Karaoke (aufr's model) for harmonies. Though I can say Dango can remain useful in certain situations” - dca

On x-minus there’s lead vocal panning setting added for Mel-RoFormer Kar by becruily model. It’s to “to "tell" the AI ​​where the main vocals are located (how they are mixed).”.

“doesn't even need Lead vocal panning a lot of the time, [the] ability to recognize what is LV and what is BV [is] impressive” - dca
“Sometimes struggles when the backing vocals are the same notes as the lead vocals” - isling. Seems like xminus panning can’t solve such issues either.

“Had a similar issue as you however with the Chase Atlantic vocal, MDX Kar V2 with stereo 80% then Chain or Max mag extracts the leftovers works very very well not perfect (at least works for most CA song) but It's enough for me to do an edit” - cali_tay98“

It seems demudder shouldn't be used when Lead vocal panning is set to something different than center, I noticed it brings back the lead vocals in the inst w/BVs as it was before changing LV panning” - dca

- Mel-Roformer Karaoke (by aufr33 & viperx) on x-minus.pro / uvronline.app / mvsep

model files (UVR instruction) mirror | yaml - online version above might work better, not sure about preprocessor model, maybe even old voc_ft, but not necessarily.

This Mel may extract more than BVE V2, if "extract directly from mixture" on MVSEP doesn’t detect the BVs (x-minus behavior for this model), the chances are "extract from vocals part" on MVSEP (which uses BS-Roformer 2024.08 for it) will detect more BVs (although with possible cross bleed between lead/back in inst+bv)

- If you ensemble the model above with Unwa v1e (Max) it removes all the muddiness of Mel Kar (dca100fb8)

"You can do it via:

Choose Process Method --> Audio Tools --> Manual Ensemble --> Max Spec

Q: How do you select both inputs?

A: Via Select Input, but if you want to batch process Manual Ensemble it's not possible yet" (dca100fb8)

Ensemble algorithms like ~“min_spec in the direct ensemble are only available if the selected type of stems in the yaml corresponds with the target in UVR.

Example:

If the target in the yaml is:

  -vocals

  -other

then you can't use it in the Vocal/Instrumental selection because it's not written that way in the yaml.” (mesk)

- Gabox Mel KaraokeGabox model (uses Aufr’s config) | Colab - finetune of becruily model

“The lead vocals are good and clean!
While the backing tracks are lossy for this model, [it still] provide[s] great convenient for those who need LdV”

“The model doesn't keep the backing vocals below the main vocals, sometimes the backing vocals will be lost even though there are backing vocals there.”

- BVE v2 model on x-minus.pro/uvronline.app for premium users | model (uses “4band_v4_ms_fullband” stock config) by Aufr33

Place the model file in Ultimate Vocal Remover\models\VR_Models and config file in lib_v5\vr_network\modelparams (if doesn’t exist already). Then pick “4band_v4_ms_fullband.json” and BV when asked to recognize the model (it has the same checksum as in modelparams folder if it’s there already). Seems like it works with “VR 5.1 model” checked (and probably without it too).

“Note that this model should be used with a rebalanced mix.

The recommended music level is no more than 25% or -12 dB.

If you use this model in your project, please credit me.”
(it's v1 version is also added in UVR’s “Download More Models”, but also without stereo width feature which can fix some issues when BVs are confused with other vocals).

On x-minus “When you select stereo, it applies a stereo narrower before AI processing.”

It used to be one of the best models for this purpose. On x-minus at certain point it used voc_ft for all vocals as a preprocessor already (not sure if it got changed).

"BVE sounds good for now but being an (u)vr model the vocals are soft (it doesn’t extract hard sounds like K, T, S etc. very well)"

"Seems to begin a phrase with a bit of confusion between lead and backing, but then kicks in with better separation later in the phrase."

“If something struggles to separate on bve v2 I change the lv panning option to either 50 or 80% [stereo or center], and it separates it amazingly.

It even allows me to separate backing backing vocals from backing vocals”
- For UVR BVE v2 LV bleed in Song without LV - download the BV track and add it to inst v1e model result - no vocal bleed/residues (introC).

“I tried it ensembled with Gabox's model [kar_gabox”] and they are amazing together. Yes you have to make the primary stem of aufr33's model "Instrumental" if you're ensembling” - AG89

- Newer Gabox experimental Karaoke model (June 2025). It’s one stem target so keep extract_instrumental enabled for the rest stem.

“really hard to tell the difference between this and becruily's karaoke model” minus the latter has more target stems.

- Chain ensemble mode for B.V. models (available on x-minus.pro for premium users, added in UVR beta 5.5/9.15 beta patch already):

It is possible to recreate this approach using non-BVE v2 models in UVR by processing the output of one Karaoke model by another (possibly VR model as the latter) with Settings>Additional Settings>Vocal Split Mode option (so it separates using the main model for all vocals, then it uses the result as input for the next model).

So you might experiment with using voc_ft or Kim Vocal 2 or 1296 as the main vocal model in the main UVR window, and in Vocal Split Mode use HP5 or HP6 or BVE model, so you won’t have to make the process in 2 steps manually, so separating the result with another model once the first separation is done. Although Vocal Split Mode was designed mainly for BVE models, so in case of any problems with HP5/6 or Karaoke, you can test out also Settings>Choose Advanced Menu>[model arch]>Secondary model instead.

Don't forget reading vocals to find the best separation method for your song to use it for separation with Karaoke/BVE models.

Recommended ensemble settings for Karaoke in UVR 5 GUI (instrumentals with backing vocals):

- 5_HP-Karaoke-UVR, 6_HP-Karaoke-UVR, UVR-MDX-NET Karaoke 2 (Max Spec)

(in e.g. “min/max” the latter is for instrumental)

- Alternatively, use Manual Ensemble with UVR with Max Spec using x-minus’ UVR BVE v2 result and the UVR ensemble result from the above.

Or single model:

- HP_KAROKEE-MSB2-3BAND-3090 (a.k.a. VR's 6HP-Karaoke-UVR)

- UVR BV v2 on x-minus (and download "Song without L.V.". Better solution, newer, different model)

- 5HP can be sometimes better than 6HP

(UVR5 GUI / x-minus.pro / Colab) - you might want to use Kim Vocal 2 or voc_ft or 1296 or MDX23C first for better results.

- UVR-BVE-4B_SN-44100-1

Q: What are the differences between Mel-Roformer Karaoke and the last model?

A: If the vocals don't contain harmonies, this model (Mel) is better. In other cases, it is better to use the MDX+UVR Chain ensemble for now.

- Gabox denoise/debleed Mel-Roformer | model | yaml | Colab

“better results than kar v2”

- Gabox kar v2 | model | yaml 

- De-echo VR model in UVR5 GUI set to maximum aggression

- MedleyVox with our trained model (more coherent results than current BV models)

Or ensemble in UVR:

"The karaoke ensemble works best with isolated vocals rather than the full track itself"

- VR Arc: 6HP-Karaoke-UVR

- MDX-Net: UVR-MDX-NET Karaoke 2

- Demucs: v4 | htdemucs_ft

Or:

- VR Arc: 5HP-Karaoke-UVR

- VR Arc: 6HP-Karaoke-UVR

- MDX-Net: UVR-MDX-NET Karaoke 2

(Max Spec, aggression 0, high-end process)

Or:

- VR arc: 5_HP-Karaoke

- MDX-Net: UVR-MDX Karaoke 1

- MDX-Net: UVR-MDX Karaoke 2

(you might want to turn off high-end process and post process)

Or:

- VR Arc: 5HP-Karaoke-UVR

- VR Arc: 6HP-Karaoke-UVR

- MDX-Net: UVR-MDX-NET Karaoke 1

- MDX-Net: UVR-MDX-NET Karaoke 2

(Min/Min Spec, Window Size 512, Aggression 100, TTA On)

If your main vocals are confused with backing vocals, use X-Minus and set "Lead vocal placement" to center (not in UVR5 at the moment).

Or Mateus Contini's method.

How to extract backing vocals X-Minus Guide (can be executed in UVR5 as well)

Vinctekan Q&A

Q: Which BVE aggression settings (for VR model, e.g. uvr-bve-4b-sn-44100-1) is good for backing removal?

A: “I recommend starting from exactly from 0 and working from there to either - or +.

0 is the baseline for BVE that are almost perfectly center.

If it's off to the left or right a little bit, I would start from 50”

Q: How do I tell what side BVs are panned or if they are Stereo 50 % or 80 % without extracting them?

A:  “It's more about listening to the track. The way I used to it is to invert the left channel with the right channel. In most cases this should only leave the reverb of the vocals in place, but if there is backing vocals that is panned either left or right, then it should be a bit louder than the reverb. Audacity's [Vocal Reduction and Isolation>Analyze] feature usually can give a rough estimates as to how related the two channels are, but that does not tell where the backing vocal actually is. I would only recommend doing the above with a vocal output, though.”

Q: Does anyone know how to tell what side BV's (backing Vocals) are panned similar to this? Like, is there a way to tell using RipX? Or another tool. In my case I think mine might be Stereo 20 30 percent or lower

A: “Your ears [probably the least effective]

If you have Audacity, select your entire track, and select [Vocal reduction and Isolation] and select the [Analyze] but it won't tell you which direction the panning is in.

Or use it to isolate the sides, and just take a look at the output levels of each channel.

Spectralayers's [Unmix>Multichannel Content] tab can measure the output of frequencies in the spectrogram and can tell you when certain elements are not equal in loudness, which you can restore.”

- Dango.ai has also a good BVE model (expensive) - at least sometimes it gives better results than uvr bve v2 to get songs without lead vocals. “meant for separating melody from harmony, not separating singer from singer, so you'll hear both [singers in one] stem” if present

Later a new Advanced repair tool was added to “fix any backing-vocals errors”

- AudiosourceRE Demix Pro has BVE/lead vocals model

- lalal.ai has a new decent lead and backing vocals model

If bve or mel karaoke model CAN do it, then they'll do it better, but if they CAN'T do it, then lalal will do it better. "I have seen lalal work better on mono audio than bve model." ~Isling

dca100fb8: “For Instrumentals with Backing Vocals:

If Mel Kar doesn't work, it's likely Dango Backing Vocal Keeper will not too, although it’s not always the case, and still worth trying out - separating left and right channel with Dango Backing Vocal Keeper once fixed the issue.

For back vocals/lead vocals model

If neither Mel Kar or UVR BVE v2 work, it's likely lalal.ai Lead & Back Vocal Splitter will work instead. Its deep Extraction seems to provide better results than Clear Cut”

- Advanced chain processing chart (image)

It’s a method utilizing old models, and e.g. Kim Vocals 2 can be potentially replaced by unwa’s BS/Mel-Roformer models in beta UVR (or other good method for vocals) or ensembles mentioned in this document. Check the best current methods for vocals in one stem to find what works the best for your song to get all vocals before splitting to other stems using this diagram.

htdemucs v4 above can be replaced by htdemucs_ft, as it's the fine-tuned version of the model (or MDX23 Colab). Even better, you can use some of the methods for 4 stems in this GDoc (like drums on x-minus).

De-echo and reverb models can be potentially replaced by some better paid plugins like:

DeVerberate by Acon Digital, Accentize DeRoom Pro (more in the de-reverb section).

UVR Denoise can be potentially replaced by less aggressive Aufr33 model on x-minus.pro (used when aggressiveness is set to minimum), and there’s also newer Mel-Roformer (read de-reverb section).

As for Karaoke models, there's e.g. a Mel-Roformer model on x-minus.pro for premium users or MVSEP/jarredeou inference Colab.

"If the vocals don't contain harmonies, this model (Mel) is better. In other cases, it is better to use the MDX+UVR Chain ensemble for now.". It is possible to recreate to some extent this approach while not using BVE v2 models, by processing the output of main vocal model by one of Karaoke/BVE models in UVR (possibly VR model as the latter) using Settings>Additional Settings>Vocal Splitter Options, so it separates using one model, then it uses the result as input for the next model (see the Karaoke section).

MedleyVox (not available in UVR) will be useful in the end in cases when everything else fails after you obtain all vocals in one stem, as it's very narrowband. But you can use AudioSR on it afterwards.

>Keeping only lead vocals in a song (backing vocals extractor)

Sometimes the same model might work for lead, sometimes for back vocals depending on a song.

Sometimes extracting vocals with a good model can degrade the quality, experiment both ways but “if they are really quiet in the mix, sometimes they first need to be extracted to come out clear, but lead vocals usually aren't that way… and you can easily confuse the model by extracting them first” - rainboomdash

- Anvuew’s Karaoke BS-Roformer model | metrics | MSST | MVSEP | Colab

(stems: Lead vocals/backing with instrumental)
“this is the best for extracting lead vocals, currently” - rainboomdash

“extracts lead vocals a bit better than karaoke becruily frazer, and in some parts, the lead vocals from karaoke anvuew still sound brighter compared to karaoke becruily frazer, which sounds a bit more compressed. oh, and for some reason, the becruily frazer model doesn’t detect vocals with radio effects, while anvuew’s model handles them just fine” - neoculture

“lead vocals leak into instrumental (...) Mel Becruily and Frazer’s BS don’t have this problem”

In that case, maybe “isolate the acapella first in almost all cases of using a karaoke model” or use the model below instead.

- Ensemble of becruily frazer karaoke + anvuew karaoke (Max Spec) -

to get the backing vocals, use instrumental only in UVR to get only them - neoculture

If one model in the ensemble does a bad job, it will spoil the result, so be aware.

- (x-minus/uvronline) Lead and Backing vocal separator (in “Extract Backing vocals>Mel-Roformer Lead/Back)

It uses big beta 5e model as preprocessor for becruily Mel Karaoke model  

“In fact, the big beta 5e Deux model is run after becruily Mel Karaoke” - Aufr33, Gabox (so you don’t need the additional step to use this separator), plus it also allows controlling option for lead vocal panning like for BVE v2 (It’s to “to "tell" the AI ​​where the main vocals are located (how they are mixed).”. “Doesn't even need Lead vocal panning a lot of the time, [the] ability to recognize what is LV and what is BV [is] impressive” - dca).

“The new separator is available in the free version, however, due to its resource intensity, only the first minute of the song will be processed.” if you don’t have a premium.

The difference from using single becruily Kar model is that, here, “you get the third track, backing vocals.”.

- Mel-Roformer small_karaoke_gaboxauf by Gabox and Aufr33 | yaml | Colab | Metrics.

For RTX 5000 patch and torch._dynamo error, edit “use_torch_checkpoint: True” to False (or delete the line).

“it sounds much more fuller than other karaoke models I use [anvuew bs roformer,  bs karaoke gabox, and becruily karaoke] and clears out bg vocal much more accurately” - baptizedinfear

“great! got to hear vocals I could not hear before.” - makesomenoiseyuh

“it's not great with duets btw” - Gabox

Lowering chunk_size to 352800 makes it muddier, but less noise, “maybe in between would be nice” - rainboomdash

- Mel-Roformer Karaoke by becruily model file | Colab | MVSEP | x-minus

“It's a dual model, trained for both vocals and instrumental. It sounds fuller + understands better what is lead and background vocal

- Gabox Mel KaraokeGabox model (uses Aufr’s config) | Colab

“The lead vocals are good and clean!
While the backing tracks are lossy for this model, [it still] provide[s] great convenient for those who need LdV”

“The model doesn't keep the backing vocals below the main vocals, sometimes the backing vocals will be lost even though there are backing vocals there.”

- Newer Gabox experimental Karaoke model (June 2025). It’s one stem target so keep extract_instrumental enabled for the rest stem.

“really hard to tell the difference between this and becruily's karaoke model” minus the latter has more target stems.

- BVE v2 model on x-minus.pro/uvronline.app for premium users | model (“4band_v4_ms_fullband” stock config) by Aufr33

Place the model file in Ultimate Vocal Remover\models\VR_Models and config file in lib_v5\vr_network\modelparams. Then pick “4band_v4_ms_fullband.json” and BV when asked to recognize the model (it has the same checksum as in modelparams folder if it’s there already). Also, I think it's not a VR 5.1 model.

“Note that this model should be used with a rebalanced mix.

The recommended music level is no more than 25% or -12 dB.

If you use this model in your project, please credit me.”
(it's v1 version is also added in UVR’s “Download More Models”, but also without the stereo width feature which can fix some issues when BVs are confused with other vocals).

On x-minus “When you select stereo, it applies a stereo narrower before AI processing.”

(sometimes it might work for lead, sometimes for back vocals)

It used to be one of the best models so far. On x-minus it used voc_ft for all vocals as a preprocessor already.

"Seems to begin a phrase with a bit of confusion between lead and backing, but then kicks in with better separation later in the phrase."

“If something struggles to separate on bve v2 I change the lv panning option to either 50 or 80% [stereo or center], and it separates it amazingly.

It even allows me to separate backing backing vocals from backing vocals”

"BVE sounds good for now but being an (U)VR model the vocals are soft (it doesn’t extract hard sounds like K, T, S etc. very well)"
- For UVR BVE v2 LV bleed in Song without LV - download the BV track and add it to inst v1e model result - no vocal bleed/residues (introC).

- MVSep MelBand Karaoke (lead/back vocals) SCNet XL IHF by becruily (SDR: 9.53)

“Could be more useful for BVE than karaoke [models].” - becruily

- MVSep BS-Roformer Lead Vocal model from “MVSep Mega 53 Stems” all-in-one model or single models (less VRAM-hungry) - 1.27GB | splifft | MVSepless HF / HF CPU / Colab / newer (you can pick which stems you want there, but it’s in Russian, and HF doesn’t support auto-translate, but Colab might)

The model can be muddy, it's small. It was further trained on MVSep in the all-in-one model and returns only present stems.

- Mel-Roformer Karaoke (by aufr33 & viperx) on x-minus.pro / uvronline.app / mvsep

model file (UVR instruction) - online version above might work better

(may extract more than BVE V2), if "extract directly from mixture" on MVSEP doesn’t detect the BVs (x-minus behavior for this model), the chances are "extract from vocals part" on MVSEP (which uses BS-Roformer 2024.08 for it) will detect more BVs (although with possible cross bleed between lead/back in inst+bv)

- Gabox Mel KaraokeGabox model (uses Aufr’s config) | Colab

“The lead vocals are good and clean!
While the backing tracks are lossy for this model, [it still] provide[s] great convenient for those who need LdV”

“The model doesn't keep the backing vocals below the main vocals, sometimes the backing vocals will be lost even though there are backing vocals there.”

- Gabox Mel KaraokeGabox model (uses Aufr’s config) | Colab

“The lead vocals are good and clean!
While the backing tracks are lossy for this model, [it still] provide[s] great convenient for those who need LdV”

“The model doesn't keep the backing vocals below the main vocals, sometimes the backing vocals will be lost even though there are backing vocals there.”

- anvuew dereverb Mel-Roformer model v2 (“on some songs I tried it worked better than karaoke models”) not for duets, it debleeds too, “cleaner backing vocals [than the below], can sometimes mistake main vocals/delay for backing vocals more often than [the below])”
- karokee_4band_v2_sn a.k.a. UVR-MDX NET Karaoke 2 (MVSEP [MDX B (Karaoke)] / Colab / UVR5 GUI / x-minus.pro) - “best for keeping lead vocal detail” on its own, “cleaner main vocals, has significantly less bleeding than the MVSep counterpart”, removes backing vocals from a track, but when we use min_mag_k it can return similar results to:

- Demix Pro (paid, “keeps more backing vocals [than Karaoke 2] (and somehow the lead vocals are also better most of the time, with fuller sound”

“Demix is better for keeping background vocals yes, but for the lead ones they tend to sound weaker (the spectrum isn’t as full and has more holes than karaoke 2, but this isn’t always a bad thing because the lead vocals themselves are cleaner, the mdx karaoke 2 might produce fuller lead vocals, but you will most certainly have some background vocals left too”)

- Fusion model of Gabox Karaoke and Aufr33/viperx model on MVSEP
(tend to confuse BV/LV more than single models)

- Gonzaluigi Karaoke fusion models (standard | aggresive | yaml) -

also confuses BV/LV more

- MDX B Karaoke on mvsep.com (exclusive) - good, but as an alternative you could use MDX Karaoke 2 in UVR 5 (they are different)

“I personally wouldn't recommend 5/6_hp karaoke, except for using 5_hp karaoke as a last resort, you could also use the x minus bve model in uvr which sometimes is good with lead vocals”

- UVR-BVE-4B_SN-44100-1
- Center extraction model

- Melodyne guide

- RipX

(doesn’t work for everyone)

- MDX-UVR Inst HQ_3 - new, best in removing background vocals from a song (e.g. from Kim Vocal 2)

Or consecutive models processing:

- Vocals (good vocal stem from e.g. voc_ft or 1296 or MDX23C single models or ensembles of MDX23C / MDX23 2.2 / UVR top/near top SDR / Ensemble of only vocal models: Kim 1, 2, voc_ft, MDX23C_D1581, eventually with demucs_ft

>The vocal result separated with->

Karaoke model -> Lead_Voc & Backing_Voc

Tutorial 

 (+ experimentally split stereo channels and separate them on their own, then join channels back)

- arigato78 method

“Karaoke 2 really won't pick up any chorus lead vocals EXCEPT for ad-libs

6-HP will pick up the melody, although it's usually muffled as hell”

“Q: is mdx karaoke 2 still the best for lead and back vocals' separation?

A: I'm finding it's the best for "fullness" but 6-HP picks up chorus melody while K2 only usually picks up ad-libs

I personally like mixing K2, 6-HP (and sometimes 5-HP if 6-HP sounds very thin) together

also, let's say a verse has back vocals that are just the melody behind the lead vocal (instead of harmonies) for a doubling effect, sometimes K2 will still pick up both the lead and double.”

AG89 avg ensemble:

UVR-5-1_4band_v4_ms_fullband_BVE_V2,

Karaoke_GaboxV2,

mel-band_karaoke_fusion_standard

Harmonies

E.g. for more layers from the above.

On MVSEP (or on your own) “'Use As Is' [so the mixture - your input file will be used] and 'Extract Vocals First' [so using vocal model] might be the difference between splitting a vocalist's 'double vocal' and not.” or on your own. - cypha_sarin.

Also, as a preprocessor for all vocals, experiment with panning using e.g. A1StereoControl VST2 32 bit plugin beforehand - dementia2009 (might work as a counterpart of the panning trick in UVRonline).

- becruily & frazer BS-Roformer Karaoke model (in most cases not using vocal model on the top of it is better for doubles, sometimes “it's necessary if the vocal was very, very quiet, otherwise it's extremely muddy” - rainboomdash; or check more Karaoke models)

- Stereo width feature for uvr bve v2 by setting it to 80% (it might use voc_ft as preprocessor already) - on x-minus.pro/uvronline.app

- MVSep SATB Choir (soprano, alt, tenor, bass; much better metrics than the old SATB below, and better BS-Roformer model architecture; use “a karaoke model first and then using the SATB model on the backing vocals, it tends to have less confusion and more opportunities for manual cleanup” - dynamic64).

- MVSep Choir (works for e.g. choirs buried under main vocal in e.g. pop music)

- Older public SATB choir models by Dry Paint Dealer Undr (soprano, alto, tenor, and bass - vocal harmonies) “I think it’s a great model for anyone looking to study harmonies more closely. Better than the Medley Vox model. So far, it’s the best result I’ve found" - Roberto89.

Nevertheless, it can have bleeding.

“scnet_masked model was the best in the end” - Dry Paint Dealer Undr

“I've tried the SCNet one, it's really noisy, and it has a lot of bleed, it kinda works. I can see the potential of this kind of model ngl.” - smilewasfound

“You can’t install [the VR ones] into UVR since that only supports VR v5 [and 5.1] not VR v6

For VR6 bass and soprano model -X flag is needed.

Demucs and SCNet_masked are not compatible with UVR for now either (use MSST instead), so only regular SCNet variant can be used in UVR.

- Medley Vox (free, 24kHz SR model trained by Cyrus, Colab, local installation tutorial, more info, use e.g. AudioSR afterwards; “help[ed] me separate a few layers harmonies on one song, though I had to keep mixing certain ones back together and run them through the same model to get a cleaner result”)

- Sucial Mel-Roformer dereverb/echo model #3 called “fused”: model | yaml

- Melodyne (paid, 30 days trial) - “the best way to ensure it’s the correct voice”, or

- Hit 'n' Mix RipX DAW Pro 7 (paid/trial)

In Melodyne it is "harder to do but can be cleaner since you can more easily deal with the incorrect harmonics than RipX sometimes choses"

"every time I’d run a song through RipX I was only able to separate 4-5" harmonies

(or also)

- (prob.) Revoice Pro 5

- Mel-Roformer Karaoke (by becruily) model file (better than aufr’s model)

- Dango.ai (paid, might be still useful in certain situations)

- Choral Quartets F0 Extractor - “midi outputs, but it works”

For research

https://c4dm.eecs.qmul.ac.uk/ChoralSep/

https://c4dm.eecs.qmul.ac.uk/EnsembleSet/ (similar results to MedleyVox)

For two singers in a duet from one song

(use on already separated vocals)

- MVSep SATB Choir (in case of crossbleeding “when i put them in the DAW and inverted it gave really satisfactory results”, worked for two main male vocals where the below failed - lekt0rs)

- becruily & frazer BS Karaoke (sometimes can separate even 3 singers if the vocals aren't completely glued into one, especially if it's male and female - maxerv19,

"It's getting better than MedleyVox" - ryanz48)

- Becruily Mel Karaoke
- Dry Paint Dealer Undr MelBand Roformer Duet model (Singer 1 and Singer_2) | MK Colab fork

For RTX 5000 patch and torch._dynamo error, edit “use_torch_checkpoint: True” to False (or delete the line).

“only works on isolated vocals”, “It's not perfect but it can work, although how well it works really varies from song to song. I was originally going to hold off on releasing this to see if I could get it any better but I saw people wanted a model like this and thought I should probably just release it now.” - dpdu
“Surprisingly good. The model has slight bleed from the backing vocals but less than the karaoke models. Well, about the same. Still has quite a bit which is what i was expecting” - isling

- MedleyVox (trained by Cyrus, 24kHz SR - use e.g. AudioSR afterwards, MVSEP, Colab, local installation tutorial [use vocals 238 model], more info)

“best at it, but not guaranteed to work always. 5% chances of perfectly separating duets on MedleyVox, or else it always false detects and switches back and forth” “also does a pretty good job on other solo instruments”

- MVSEP Multispeaker model (Experimental section at the bottom)
Works well for rap overlapped with singing in one already separated vocal stem.

“Seems very picky with audio, most of the songs/files I tried didn't work

MedleyVox works on the other side (regardless that it's of lower quality)” - becruily

- MVSEP Male/Female:

a) Mel-Roformer Male/Female separation model 13.03 SDR

a) SCNet XL Male/Female separation model on MVSEP (same model base)

SDR on the same dataset: 11.8346 vs 6.5259 (Sucial)

Sometimes the old Sucial model might still do a better job at times, so feel free to experiment.

- Aufr33 BS-Roformer Male/Female beta model | config | Colab | x-minus (uses Kim-Mel-Band-Roformer-FT2 as preprocessor) | MVSEP

(based on BS-RoFormer Chorus Male Female by Sucial) SDR 8.18

- Male/female BS-Roformer model by Sucial | config for UVR | tensor match error fix
If they sing at intervals [one by one - not together] they cannot be separated. | MVSEP

- MVSep Karaoke BS-Roformer by MVSep team (works for double-tracked vocals)

- Mel-Roformer Karaoke on x-minus.pro (model files in Karaoke)

- MDX-UVR Karaoke models

- VR's 5_HP or ev. 6_HP in UVR

- BVE v2 on x-minus (already uses voc_ft as preprocessor for separating vocals)

It might be still not enough, then continue and/or look for Dolby Atmos rip and retry; works for several backing vocals when lead vocal panning is set to center, “then running the bgv through bve v2 again, but this time set lead vocal panning to 80% but be aware the lead vocal quality will not be that good with this model” - Isling)

- Dry Paint Dealer Undr’s Melband Roformer and Demucs Lead and Rhythm guitar models

- duet-svs-diffusion (“mono/16kHz, 24kHz sample rate, and quality is lower than MedleyVox models”)

- RipX (paid)

- Melodyne (paid, “with polyphonic mode, with a lot of manual finetuning in the detection tab, and this can only work if the voices are not on same pitch”).

- SpectraLayers 11 (but it’s mainly dedicated for voice, not singing)

Spectral painting:

- ISSE (free, you can figure out which voice is who's just by frequencies alone; use on e.g. separated vocals too)

- RX Editor’s brush (video by Bas)

- Audacity (image) less effective

- Ampter 

Filter-Artist

- AudioPaint

Notes

If artists sing the same notes, Karaoke models will rather not work in this case.

If BVs are heard in the center, don't use the MDX karaoke model but the VR karaoke model instead.

Use the chain algorithm with mdx (kar) v2 on x-minus which will use uvr (kar) v2 to solve the issue. It will be available after you process the song with MDX. (Aufr33/dca)

“The MDX models seem to have a cleaner separation between lead and backing/background vocals, but they often don't do any actual separation, meanwhile the VR models are less clean, but they seem to be better at detecting lead and background”

“MDX models basically require the lead to be completely center and the BV to be stereo

whereas VR ones don't really care as much about stereo placement”

You could also ask playdasegunda/play da primeira/viperx for separation, as he has some decent private method/models for double vocals better than becruily model (although the latter can be still close), although newer frazer & becruily model is “no way inferior to the ViperX (Play da Segunda). The rumour says, viperx model on Play da Segunda was trained on 40GB dataset allegedly (it would be small), and the model can be bought for 500$ when you contact via email. It's actually a set of models used for the final inference.
For the record, the open-sourced model by the duo costed 600$ on compute and probably used a bigger dataset, achieved a bit smaller SDR, but it’s a single model.

For research:

“These archs are [...] really promising for multiple speakers separation, and should be working for multiple singers separation if trained on singing voice:

https://github.com/dmlguq456/SepReformer (current SOTA)

https://github.com/JusperLee/TDANet

https://github.com/alibabasglab/MossFormer2

> Separating two main vocals

E.g. one panned about 30% left and the other right

- “use bve v2 and click the “lead vocal panning” button” on x-minus premium

For vocals with vocoder 

- voc_ft

Alternatively, you can use:

- 5HP Karaoke (e.g. with aggression settings raised up) or

- Karaoke 2 model (UVR5 or Colabs). Try out separating the result obtained with voc_ft as well.

- BS-Roformer model ver. 2024.04 on MVSEP (better on vocoder than the viperx’ model).

"If you have a track with 3 different vocal layers at different parts, it's better to only isolate the parts with 'two voices at once' so to speak"

Various speakers' isolation (from e.g. podcast or movie)

- MVSEP Male/Female SCNet model

- MVSEP Male/Female MelRoformer model

- Aufr33 BS-Roformer Male/Female beta model | config | Colab (based on the model below)
- Male/female BS-Roformer model by Sucial | config for UVR | tensor match error fix
(if they sing at intervals [one by one] they cannot be separated)

- Multispeaker model on MVSEP

- BS-Roformer becruily & frazer Karaoke model

Guide and script for WhisperX (separating people in recording)

- https://github.com/alexlnkp/Easy-Audio-Diarisation

- Spectralayers 11’s unmix multiple voices option

(for further research) - some of these tools might get useful:

https://github.com/dmlguq456/SepReformer (SOTA for 2 speakers)

https://paperswithcode.com/task/speaker-separation/latest

https://arxiv.org/abs/2301.13341

https://paperswithcode.com/task/multi-speaker-source-separation/latest

____________________________________________________________________


> 4-6 stems (drums, bass, others, vocals + opt. guitar, piano):
- Currently when used on AI-generated music, usually hihats will be left behind.

- All the metrics provided for Multisong dataset on MVSEP

- You might want to use the already well-sounding instrumental (or even ensemble), possessed with 2 stem model in the section above first, and then separate using the following models.
- Furthermore, you can slow down your song by x0.75 speed - the result can be - more elements in other stem and better snaps and human claps using 4 stems.

Read the Tips to enhance separation #4 for more.

- MVSep Ensemble (vocals, instrum, bass, drums, other) (2025.06.30) - premium only

SDR bass 14.85, drums 14.33, other 9, vocals 11.93

All the metrics are better than the SW, but consider passing the other stem through the SW below to get piano and guitar stems too

- Logic Pro (May 2025 update) / BS-Roformer SW 6 stem | MVSEP | uvronline

SDR bass 14.57, drums 14.05, piano 7.79, guitar 9.00, other 8.66, vocals 11.27

Currently, the best single model bag SDR for all stems but vocals, drums have lower fullness than MVSep SCNet XL drums 14.26 vs 21.21). Excellent guitar and piano.

“guitar model sounds better than demucs, mvsep, and moises” - Sausum

“it's not a fullness emphasis or anything, but it's shockingly good at understanding different types of instruments and keeping them consistent sounding” - becruily

vocals doesn’t have the biggest metric, but are good for deep voices.
Drums lacks some fullness but “I got better drums/bass separation with that model than with any others when input is some live/rehearsal recordings with shitty sound” - jarredou

Although, compared with Mel-Roformer drums on uvronline:

“separates far far better when it’s programmed instruments compared to actual recorded ones” - isling

“Roformer SW is putting finger snaps and foot taps as vocals and in the vocal stems.” - GodzFire

“gets not just drums but anything percussive/non-melodic. Which I personally don't mind, but yeah it does cause problems with drumsep models because they're only expecting standard drums.”

Bass can be occasionally worse vs demucs_ft as bass stem Demucs “considers not only spectrograms but also waveforms”.

“can't differ an electric bass with pedal effect from an electric guitar” - qraiqu

"better than Lalal.ai by a long shot too" - nowarrantywarren

As for MVSEP “just as good a job on vocals as the paid version's ensemble preset.”

In UVR it takes two hours for overlap 11 for a 4 minute file to separate on 6700 XT. Keep it at 2 (the fastest) - it will also take 2 hours, but on only i3-7100u (GPU Conversion disabled).
Highest SDR with 882000 chunk_size. 352800 on 4GB NVIDIA GPUs works faster with this or lower than 485100 setting. Some people prefer using it with overlap 8.

SESA Colab can apply phase remix demudder for this model.

- MDX23 v.2.5 by ZFTurbo, fork by jarredou (Colab; 4 stems when they're enabled)

Multisong dataset SDR bass: 12.58, drums: 11.97, other: 7.28, vocals: 11.10 (v2.4)

It’s weighted ensemble of various older 4 stem Demucs models with weighted ensemble of now outdated 2 stem models for 4 stem input, so the metrics for RAW 4 stems output (without getting instrumental from ensemble first) will be a bit lower, and more for other stem - even by 1.38+, and 0.24+ for bass, and 0.02+ for drums (read more).
~"compared to this, demucs_ft drums sound like compressed".

- Other Ensembles 2/4/8 stems (MVSEP in premium) - similar or better results with newer single stem models combined, various ensembles to choose from, freedom to experiment.

Read all the ensembles metrics sorted by instrumental bleedless.

- Demucs_ft (Colab | UVR | MVSEP |  MVSepless HF)

- the best variant of Demucs 4 models, frequently better pick for drumsep than the SW

- xlancelab BS-Roformer (inference | model | paper | HF inference / #2) - trained on the SW model with additional percussion, synth and orchestra stems

- Huge-SCNet v1.2 model (4 stems) by Aname-Tommy

- SCNet XL IHF model (4 stems) by ZFTurbo

- BS-Roformer “MVSep Mega 53 Stems” all-in-one model | or single models (less VRAM-hungry) - 1.27GB | splifft | MVSepless HF / HF CPU / Colab / newer (you can pick which stems you want there, but it’s in Russian, and HF doesn’t support auto-translate, but Colab might)

Now the all-in one model is also available on MVSep (it’s been further trained, and skips undetected stems, FLAC is used to save space).

The model has vocal and lead vocal models, drumsep (hh, kick, toms, snare), drums, electric, acoustic and guitar, and lots of orchestral instruments (full list).

Lacks chello, backpipes, braam, xylophone, Choir (choir/other), SATB Choir (soprano, alto, tenor, and bass) which are available on MVSep separately.

They might be muddier than the single models on MVSep (each public is only 77MB, so even less than the SW).

The resulting stems won't invert with the mixture, and work independently, so e.g. despite the BV model, vocals will also have BVs.

“Definitely sounds rough in many ways, and I wouldn't use it for anything critical. However, it has its strong spots and imo has a lot of potential” - musicalman

“most [stems] came out really muddy just bc of the song I used I would test more if it didn't take so long to run the model fully” - 5b

At least free Colab currently gives ^C error (from what e.g. jesse wrote), unless you use a 30-second song fragment with chunk_size 352800. Use the HF above instead, or:

“Try with forcing batch_size=1.

The Colab is forcing batch_size=2, it was a workaround for a click issue, but not needed anymore since the issue was fixed at source a while back”

In the separation cell, edit:

data['inference']['batch_size'] = 2

to 1

 - jarredou

Consider using 88200 or 112455, so, a much lower than in the OG config from the config (441000), but the results can be unpredictable.

Works on default settings with 4070 Super and 4090, and with decreased chunk_size to at least 112455 on 4060 8GB.

You can use this memory optimized .py to process stems one by one, instead of all at the same time. It might increase system memory usage offloading VRAM depending on batch_size, and still process all stems.
The Colab errors out on full songs also with it, even with 88200 chunks and batch_size 1 and 2. Probably because of a lack of virtual memory available on local machines.

The all-in-one model also works in UVR, but because it probably forces batch_size 2, it might require more memory than MSST (older version of the MSST inference the UVR is based on had a bug where batch_size 1 caused skipping between segments, so it’s probably forced).

Alternatively, use this mask estimator mod by Frazer, so it will process each stem one by one instead of all in the same time DL.

Older 4 stems models continues later below

Chained separation order

With single stem models below, feel free to experiment with different orders of sequential stem separation to enhance separation results:

#1

1) Well sounding result from instrumental model/s first 2) drums or bass 3) piano or guitar 4) strings or horns
Note: If your song has “weird percussion that don't get picked up by drum models and if it's piano-heavy, I would go for piano first, but it sometimes leaves piano behind” - Isling

#1b

1) Instrumental separation 2) drums 3) (other of drums to) bass 4) (other of bass to) 5) guitar/piano/other instrument

“If it's a rap song (...) I have seen better result separating it this way, than using either the other of the bass or the other of the drums to get the melody” - Infisreal

So basically, on an already well sounding instrumental you should start with "separating the drums, using the other from the drums for the bass, and the other of the bass for the melody, whether it's piano, guitar, synths etc.”

#2
1) Instrumental model first 2) drums 3) piano 4) strings or horns 5) bass 6) guitars
This way once someone “ended up with a decent 5th-stage “other” stem that seemed to contain some unknown synth sounds + possible orchestra hits” vs when bass was after drums when “gave me a messy “other” with stray piano & strings content” - SilSinn9821

#3

“After each [removal], you also remove some useful parts for [the] next instruments. So I'd propose to first use the models with the best quality. Anyway, in the end others can be very muddy” - ZFTurbo

#4

1) Bass Ensemble (SCNet XL + BS Roformer + HTDemucs4) 2) Drums MVSep SCNet XL 3) Piano MVSep SCNet Large 4) Organ MelRoformer 5) Saxophone MelRoformer 6) MVSep Wind Ensemble (SCNet + Mel) 7) MVSep Guitar Ensemble (BSRoformer+ MelRoformer) 8) MVSep Strings” - dynamic64

“I think this is my new default [order]”

Example: “Personally, I like putting the instrumental through MVSep bass model (above), then putting the other stem through MVSep drums (specifically SCNet XL), then putting the other stem of that through the 6 stem model [a.k.a. SW]

The 6 stem model is best for piano and guitar, and separating out the drums and bass beforehand helps that model not have to work as hard.

And it also allows you to manually add anything that the first 2 models missed, because its a 6 stem model, and re-separates any missed drum and bass.

Creates close to studio results. As close to studio as I've ever heard”

#5

In case of ensembling various models of the same stem, if the top model in terms of SDR doesn’t have significantly lower metric, sometimes it’s better to use it instead of ensemble [esp. if it doesn’t have any crossbleeding] - dynamic64.

Just be aware that it might vary from song to song

#Tip

“Drums model often bleeds the bass and 808s into the drum track, [using bass model] prevents this issue from happening”

Drums

- Ensemble: MVSep Drums MelBand + SCNet XL + BS Roformer SW 14.35 SDR

- Ensemble: MVSep Drums Mel + MVSep Drums SCNet XL 13.70 SDR ~”Usually works the best” - dynamic64, but you might want to save intermediate files and split it into different fragments to get the best of both. Consider using instrumental model result as input.

- MVSep Drums SCNet XL, SDR 13.72 (drums only/other)

“Very hit or miss. When they're good they're really good, but when they're bad there's nothing you can do other than use a different model” - dynamic64


- BS-Roformer SW 14.05 SDR (e.g. the drums-only variant in MVSep Drums)

- MDX23 by ZFTurbo (fork by jarredou v2.4/+) 11.97 SDR (no drums-only variant, so 4 stems; sometimes better than the SW, esp. for drumsep)

- Huge-SCNet v1.2 model (4 stems) by Aname-Tommy 11.74 SDR

- BS-Roformer-DrumsOther-Duality | yaml by gilliaan | MK Colab fork | 11.64 SDR (av. in testing ver. uvronline; a fullness dual model for drummers to have good other stem to play with, can work well with vocal isolation as preprocessor for the model, or as ensemble with the SCNet or SW to increase their fullness while it's not as accurate; SDR lower than MDX23; incompatible with UVR - only one stem is exported, use MSST for it)

- SCNet XL IHF model (4 stems) by ZFTurbo 11.58 SDR

- SCNet Masked XL IHF (4 stems) by ZFTurbo 11.38 SDR

- Demucs_ft 11.40 SDR (no drums-only variant, more “compressed” sound than MDX23)

- Drums only/other Mel-Roformer model by viperx on x-minus.pro (occasionally might work only with this link), 12.54 SDR

“compared to demucs_ft it was too muddy and had too much bleed at the same time

both MVSep and xminus” - isling

“I got some bad results (that could ruin the ensemble mode). On these tracks, uvronline [x-minus] melband drums model was giving better results” - jarredou
Previously the best

- Drums only/other SCNet Large (“x-minus’ Mel band drums model is better” - drypaintdealerundr)

- Drums only/other Mel-Roformer model by ZFTurbo only on mvsep.com, 12.76 SDR

(older ZF’s model)

- 1053 BS-Roformer drums/bass model by viperx in UVR Roformer beta or Colab.

Very good drums with bass in one stem model - use instrumental as input to avoid vocal residues)
More metrics for drums models.

- MVSep Percussion

- xlancelab Percussion (inference | model | HF inference / #2)

- MVSep Tambourine

- MVSep Timpani

- MVSep Congas

- MVSep Clap

Various lower SDR drum models here like:

- SCNet Tran by ZFTurbo 10.81 SDR

- SCNet Masked Small Weights by ZFTurbo 10.57 SDR

- DTTnet (MUSDB weights) by ZFTurbo 7.58 SDR

- BS-Roformer “MVSep Mega 53 Stems” by ZFTurbo model (1.27GB/77MB) | single models (less VRAM-hungry) - 1.27GB | Colab or MVSepless HF / HF CPU / Colab (you can pick which stems you want in MVSepless)

Drumsep (separating parts of drums)

(kick/hi-hat/snare/toms/…)

Works for drums stem from e.g. Demucs_ft, MDX23 Colab or MVSEP Drums (the SW model sometimes causes issues for Drumsep).

Drums models descriptions here.

Consider using a good instrumental model for the drums model, and then use drumsep.


- MVSEP 8 stems ensemble of all the 4 drumsep models below (so besides the old Demucs model by Inagoy) metrics

- MVSEP’s SCNet 4 stem (kick, snare, toms, cymbals) best SDR for kick and similar to 6s below for toms: -0.01 SDR difference)

- MVSEP’s SCNet 5 stem (cymbals, hi-hat, kick, snare, toms)

- MVSEP’s SCNet 6 stem model (ride, crash, hi-hat, kick, snare, toms) worse snare SDR

- jarredou MDX23C 5 stem model | yaml | MK fixed Colab | UVR

(kick, snare, toms, hi-hat, cymbals)

newer, better model, also SDR-wise (although not vs MVSEP models) | not on MVSEP | worse for debleeding than the older 6 stem below.

"You can retrieve the missing stuff by summing together the 5 stems and phase invert the sum against the input, and you'll get what was missed by the model

(...) it can sound more muffled or show some artefacts (especially for toms stem (...) I never succeeded to make it get the toms fully right)" - jarredou

- jarredou/Aufr33 MDX23C 6 stem model | yaml (kick/hi-hat/snare/toms/ride and crash stems). Available in UVR or MK fixed Colab or MVSep and uvronline.app, and works better for debleeding.

More comparisons and metrics of these models here

Why not the SW model as input for drumsep?

It “gets not just drums but anything percussive/non-melodic. (...) it does cause problems with drumsep models because they're only expecting standard drums.”

“I gotta use acoustic guitar first if I'm using SW drums on Korn or anything with slap bass.

(...) bagpipes before bowed strings is another, but probably also guitar before bagpipes.. thats iffy”

- You can use Lew Apollo Universal on hi-hat and snare from drumsep 6 stems model by jarredou from drums model from e.g. flowersv10 inst model (I've noticed that most of the residual "Fuzz" is usually trapped in the snare or cymbals. - CC Karaoke)

Outdated

- SpectraLayers 11 (Unmix Drums)>OG drumsep by Inagoy/>FactorSynth (depending on how far you want to unmix the drums) >Regroover>UnMixingStation (all the last three paid)>

Virtual DJ (Stems 2.0, barely, or doesn’t pick those instruments at all).

- LarsNet (vs OG drumsep, it also allows separating hi-hats and cymbals, toms might be better)

- RipX (paid)

- SpectraLayers 10 (paid, sometimes worse, sometimes better than OG Drumsep. IDK if it was added in update or main version) "drumsep works a f* ton better when separating on this one song I've tested with the pitch shifted down 2"

- FADR.com (in paid subscription)

- Moises.ai (only for pro)

Compared to OG drumsep, Regroover allows more separations, especially when used multiple times, so allows removing parts of kicks, parts of snares etc, noises etc. More deep control. Plus, it nulls easily. But drumsep sounds better on its own, especially with higher parameters like e.g. shifts 20 and overlap 0.75-0.98. Now it can be replaced by the public MDX23C model.

Bass
- MVSEP Bass SCNet XL (the best 13.81 SDR, “It passes Food Mart - Tomodachi Life test. That's the first model to”,

- Ensemble of SCNet XL, BS and HTDemucs4 models (SDR 14.07); SCNet can be sometimes worse than Demucs which “considers not only spectrograms but also waveforms” - Unwa)

After separation, you might want to then apply Mel-RoFormer De-noise to remove the high noise, and finish with Apollo Universal by Lew (model | Colab) to get more clarity (Tobias51).

- MVSEP Bass BS Roformer 12.49 (integrated to 2 stem and multi/All-in stem ensemble too, worse other stem vs the below - more “empty” than below, problems with getting even results when bass contains a low pass filter with high resonances, picks up more “actual bass guitar than x-minus” model - drypaintdealerundr)

- x-minus.pro BS (much better at treble-heavy bass tones than demucs, better other stem than above, “catches higher end and synth basses. It makes it sound cleaner” although might sound “weird and muddy” compared to demucs_ft at times, and often when it does not capture synth bass, the mvsep bass will do - isling)

- MVSEP Bass BS Roformer SW - might be worse than the above, at least the ensemble with SW model is not better than the 14.07 - isling)

- https://twoshot.app/model/548 (paid)

- MVSep Double Bass model

“The BS-Roformer SW bass model should probably be used first to extract the double bass. Creates a better sound” - dynamic64

- MVSep Synth (also, it can sometimes pick up bass in places where regular models can’t)

- xlancelab Synth (Inference | models | HF inference / #2)

- BS-Roformer “MVSep Mega 53 Stems” model - 1.27GB | single models (less VRAM-hungry) - 1.27GB | MVSepless HF / HF CPU / Colab / newer (you can pick which stems you want there)

4 stems in one models

(continues)

- Demucs_ft (4 stem) - the best single Demucs’ model (Colab / MVSEP / UVR5 GUI)

Multisong dataset SDR 9.48: bass: 12.24, drums: 11.41, other: 5.84, vocals: 8.43
(shifts=1, overlap=0.95)
Better drums and vocals than in Demucs 6 stem model, decent acoustic guitar results in 6s. Good bass stem as Demucs “considers not only spectrograms but also waveforms”.

For 4 stems alternatively check MDX_extra, generally Demucs 6 stem model is worse than MDX-B (a.k.a. Leaderbord B) 4 stem model released with MDX-Net arch from MDX21 competition (kuielab_b_x.onnx in this Colab), and is also faster than Demucs 6s.
For Demucs use overlap 0.1 if you have instrumental instead of mixture mixed with vocals as input (at least it works with ft model) and shifts 10 or higher. For normal use case (not instrumentals input) it will give more vocal residues, overlap 0.75 is max reasonable speed-wise, as a last resort 0.95, with shifts 10-20 max.

- SCNet-large_starrytong model (4 stems) (Colab or MSST-GUI)

Multisong dataset SDR 9.29: bass: 11.28, drums: 11.24, other: 5.58, vocals: 9.06 (overlap: 4)

It’s 3x faster than Demucs (Nvidia GPU or CPU-only) and sounds better for some people, except for bass. SDR-wise, vocals are better than in demucs_ft (which is low vs single vocal/inst models anyway). Better SDR than starrytong’s MUSDB18 and Mamba models.

- Ableton 12.3 update’s high quality option separation (4 stems) - slow, works only on CPU, probably utilizes BS-Roformer. Better bleedless than ZFTurbo SCNet XL undertrained public model below, and better SDR. Might take 20 minutes for 2 minutes separation on slower CPUs (iirc mobile Sandy/Ivy). The default separation mode has very low metrics.

Some of the below can be inferenced using HF inference / #2

- SCNet XL IHF model (4 stems) by ZFTurbo

Multisong dataset SDR 9.93: bass: 11.94, drums: 11.58, other: 6.49, vocals: 9.69

- SCNet XL model (4 stems) by ZFTurbo (Colab | MSST-GUI | MSST | UVR beta patch)
Multisong dataset SDR 9.72: bass: 11.87, drums: 11.49, other: 6.19, vocals: 9.32

Better metrics than the starrytong model but “downgrade to the Large model since it produces a f*** ton of buzzing” due to undertraining.

Only bass is better in Demucs_ft - 12.24, although drums might be still better in demucs_ft.

2 stem model on MVSEP is further trained iirc.


- BS-Roformer model (4 stems) by ZFTurbo | MSST-GUI
Multisong dataset SDR 9.38: bass: 11.08, drums: 11.29, other: 5.96, vocals: 9.19
Trained on MUSDB18HQ

- Mel-Roformer models (4 stems) by Aname

a) Large (4GB): multisong dataset SDR drums: 9.72, bass: 9.40, other: 5.11

b) XL (7GB):

Despite lower AVG SDR on musdb18 dataset (8.54 vs 9), it seems to outperform demucs_ft model (only other stem has better SDR in demucs_ft - all other metrics are better in SCNet/XL/BS-Roformer).  

XL is “heavy and slow, without giving a quality boost compared to existing public 4 stems models trained on musdb18 by ZFTurbo and starrytong (BsRofo and SCNet large/XL [above])” - jarredou (maybe minus buzzing in the public XL model).

“Drums are sounding really good in particular, tested a couple songs with the large model after using unwa's v1e+ for instrumental” “drums are absolutely the standout”.

“The bass stem is definitely the weakest one from this new model. Very, very muddy and inconsistent.” - santilli_

“Large [variant] works in like 99% use case” “Large split sounds amazing so far tho”

XL “result would take so much longer, but the large results sounded better IMO” - 5B

XL model won’t work with default settings on Colab, and very slow on e.g. RTX 3060, “on 4070 Super it took like 6 mins on XL 4 stems compared to 30 seconds on Large 4 stems”

- SCNet Tran a.k.a. small model (4 stems) by ZFTurbo

Multisong dataset SDR bass: 10.99, drums: 10.87, other: 5.63, vocals: 8.42

Outperformed by the above models at least SDR-wise. Cannot be used in UVR.

- KUIELab-MDXNET23C (4 stems) - its first scores were probably from ensemble of its five models, and in that configuration it had better SDR than demucs_ft on its own, and drums had better SDR than “SCNet-large_starrytong” above (so single models’ score of any of these MDX23C models is probably lower than in demucs_ft).
- Lighter “model1” drums sounds surprisingly better than htdemucs non_ft v4 on previously separated instrumental. It handles trap really well and preserves hi-hats correctly, but in cost of other stem bleeding. v4 model can be used to clean it a bit further, but at least using GPU Conversion on AMD and older directml.dll for some GPUs, it adds more noise/artefacts, so use CPU in that case (tested on Roformer as preprocessor for instrumental). It’s relatively fast, but not as mdx_extra (which sounds rather lo-fi in that case).
- Bigger “model2” is heavier and doesn’t work on at least AMD 4GB VRAM GPUs on Beta Roformer patch #2 (before the introduction of the new overlap code).

To run model1/2 in UVR “You must change the model names in "mdx_C" from "ckpt" [name] to model1.ckpt, model2.ckpt, & model3.ckpt. [so simply add the name to the extension]” and then copy the ckpts to models\MDX-Net without yamls. There are actually 3 mdx23c models there (and 2 demucs), but model3 seems to be only for vocals (and with low SDR). So the two of three most important were explained above.
OG KUIELab’s repo.

- model_mdx23c_ep_168_sdr_7.0207 (4 stems)
Multisong dataset SDR bass: 8.40, drums: 7.73, other: 4.57, vocals: 7.36
4 stems, also trained on MUSDB18HQ, but by ZFTurbo, it’s different from the above, similar size to “model2”.

- Aname 4 stem BS-Roformer model | yaml
Multisong dataset SDR bass: 9.79, drums: 10.21, other: 5.27, vocals: 9.13

It has better SDR than the 7.0207 above (as in the SDR metrics link below), but worse than demucs_ft and BS-Roformer 4 stem ZFTurbo model above.

- BS-RoFormer 4 stems model by yukunelatyh / SYH99999 added on x-minus

https://uvronline.app/ai?discordtest
Multisong dataset SDR bass: 8.68, drums: 10.37, other: 5.05, vocals: 8.57
Some people like it more than Demucs, but “it's like demucs v4 but worse, I think.

The vocals have a ton of bleed, the bass is disappointing tbh.

The other stem has a ton of bgv and adlib bleed in it” Isling
It has SDR metrics for all stems worse than 4 stem BS-Roformer by ZFTurbo and demuics_ft.

- v2 of it was added on site with lower metrics for all stems in later period.

Smaller public 4 stem models and all metrics:

https://github.com/ZFTurbo/Music-Source-Separation-Training/blob/main/docs/pretrained_models.md#multi-stem-models

- kuielab_b - lighting-fast, but mediocre results compared to the above (can be installed in e.g. UVR)

Online services

- GSEP AI (2-4-6 stems, sonically it used to have the best other stem vs Demucs, also piano in Demucs is worse, and it picks up e-piano more frequently, GSEP electric guitar model doesn't include acoustic, it's only electric). In general, it used to have a very good piano model before many alternatives existed

- Ripple (defunct; and used to be for US iOS, at one point best SDR for all-in one single 4 stem (besides the other stem), but bad, bleedy other stem. Back then could be the best bass stem, and/or kick in drums, but not the best drums in overall vs demucs_ft, “you need something to get the rest of the drums out of the ‘other’ stem and at that point might as well use a proper drum model”, good vocals. You can minimize residues in Ripple by providing an already well separated instrumental from the section above, and/or minimizing volume of the input file by 3/6dB.

- Bandlab Splitter (4-6 stem - guitar, piano, web and iOS/Android app) - previously could be used e.g. for cleaning stems from other services, 48kHz output, stems can be misaligned, the quality got worse since one of the updates (rather not worth using anymore)

- Audioshake (paid, only non-copyrightes music, or slowed down songs [see workaround in "paid" above]) - sometimes better results than Demucs ft model.

- Spectralayers 10 - mainly for bass and drum separation -

“I think I've got some really comparable samples out of jarredou's MDX23 Colab fork”, but for vocals and instrumentals it’s mediocre [in Spectralayers 10].

- music.ai - “Bass was a fair bit better than Demucs HT, Drums about the same. Guitars were very good though. Vocal was almost the same as my cleaned up work. (...) I'd say a little clearer than MVSEP 4 ensemble. It seems to get the instrument bleed out quite well, (...) An engineer I've worked with demixed to almost the same results, it took me a few hours and achieve it [in] 39 seconds” Sam Hocking

- dango.ai (https://tuanziai.com/en-US) - also has 4 or more stems separation (expensive)

- (old) MDX23 1.0 by ZFTurbo 4 stems (Colab, desktop app, as above, much cleaner vs demucs_ft, less aggressive, but in 1.0 more low volume vocal residues in completely quiet places in instrumentals vs e.g. HQ_3, instrumentals as input should sound similar to the current v. 2.4 fork as only 4 stem separation code didn’t change much since then)

- MVSEP has also single piano, guitar and bass models (in many cases, guitar model can pick up piano better than piano model;

"works great for songs with grand piano, but only grand piano, since that’s what it was trained on.

Same with guitar, which catches more piano than piano model does, ironically").

- BS-Roformer “MVSep Mega 53 Stems” model (1.27GB) | single models (less VRAM-hungry) - 1.27GB | MVSepless HF / HF CPU / Colab / newer (you can pick which stems you want there)

- x-minus/uvronline.app has also single acoustic and guitar models by viperx

- viperx’ piano model is also on x-minus/uvronline.app.

(More on piano and guitar models later below)

To enhance 4 stem results, you can use good instrumental obtained from other source as input for the above (before instrumental Roformers it could be e.g. KaraFan, and its different presets ensembled in UVR5 app with Audio Tools>Manual Ensemble)

For the best results for piano or guitar models, use other stem from 4 stems from e.g. “Ensemble 8 models” or MDX23 Colab or htdemucs_ft as input.

Moises.ai - although drums might be better using e.g. “MVSep Drums” already, probably vs Mel-Roformer on MVSEP or x-minus (not sure) [Moises] can give “better results (...) if the input material is for example cassette-tape sourced or post-FM).

FL Studio - rather nothing better than the solutions above

Older 4 stem models in UVR (for some specific songs, e.g. while trying to fix bleeding across stems):

htdemucs

htdemucs_mmi

mdx_extra

kuielab_b

Strings

- Dango.ai (e.g. violin, erhu; paid, free 30 seconds fragments) - "impressive results" for at least violin

- Music.ai (paid, “Dango.ai and Music.ai [previously] the best strings models [Dango sounds fuller meanwhile Music.ai has more accurate recognition of strings but sounds a bit too filtered]” - from before BS Strings release)

- BS-Roformer Bowed Strings model | yaml by gilliaan | MK Colab fork (fullness duality model, UVR exports only one stem and inverts it, use MSST instead. “Trained on everything except orchestral. (...) Best results at overlap 6. Overlap 1 is also worth trying in some cases! it captures more of the strings, but at the cost of some bleeding/cutting, so it depends on the track”).

- MVSep Bowed Strings (“doesn't disappoint”, SDR 5.41; formerly “MVSep Strings”)

- MVSep Plucked Strings

- MVSep Violin (sometimes does better than the strings model for strings)

- MVSep Synth (has synth strings stems included during training)

- MVSep SATB Choir (works with strings too)

- Moises.ai (paid, not bad)

- Audioshake


- MVSep Strings MDX23C (it’s “weak”, SDR 3.84, perhaps deleted)

- x-minus.pro/Uvronline.app Mel-Roformer model by viperx (SDR 2.87) (sometimes with this link or this link)

- Demix Pro (paid, free trial)

- RipX DeepRemix (once was told to be the best bass model, but it doesn’t score that good SDR-wise, probably it’s Demucs 3 (demucs_extra) and is worse than Demucs_ft and rather also vs MDX23 above; could have been updated) (paid)

- Sometimes Wind model in UVR5 GUI picks up strings

- MVSEP Harp

- MVSep Mandolin

- MVSep Banjo

- MVSep Sitar

- MVSep Ukulele

- MVSep Dobro

- BS-Roformer “MVSep Mega 53 Stems” model - 1.27GB | single models (less VRAM-hungry) - 1.27GB | MVSepless HF / HF CPU / Colab / newer (you can pick which stems you want there)

Violin

- MVSep Violin BS Roformer

“It can even separate violin quartets from cellos, so cool.” - smilewasfound

“Very neat model. (...) Sometimes the model does seem to pick up more than just violins imo, but yeah for separating high strings in particular it is really cool.” - Musicalman

- Dango.ai “impressive”

- MVSEP Viola

- MVSEP Chello (not included in “MVSep Mega 53 Stems” model | single models (less VRAM-hungry) - 1.27GB | MVSepless HF / HF CPU / Colab / newer (you can pick which stems you want there)

Electric guitar
“For better results you might try first removing vocals.”

Audioshake>RipX/>Demix Pro>lalal.ai (e.g. lead guitars; the model got better by the time)

(they’re paid ones)/

Logic Pro>GSEP>Demucs 6s (free)<Moises.ai (paid “holy shit better [vs demucs, but] still pretty bad”)

- Dango.ai (paid)

- Music.ai (paid, free trial)

- Logic Pro (paid, May 2025 update) / BS-Roformer SW 6 stems a.k.a. MVSEP Guitar SW
(“really on point. So far it separated super well, also didn’t confuse organs for guitars and certain piano sounds as well.” - Tobias51

“guitar model sounds better than Demcus, MVsep, and Moises” - Sausum

“guitar in particular was amazing. All other models I tried had trouble with it” - Musicalman)

| | | | | 

- MVSep Electric Guitar (“really neat. One thing I noticed is that it seems to be better than other models at picking up midi/synth lead guitars (...) also gets tripped up a bit more by weird FX and synth sounds being partially flagged as guitar” - Musicalman)

- BS-Roformer “MVSep Mega 53 Stems” model - 1.27GB | single models (less VRAM-hungry) - 1.27GB | MVSepless HF / HF CPU / Colab / newer (you can pick which stems you want there)

- Becruily Melband guitar | yaml (“Not SOTA, but much more efficient and comparable to existing guitar models, and for some songs it might work better because it picks up more guitars [though it can also pick some other instruments].”)

- MVSep Mel-Roformer model and the previous - MDX23C one. Mel is “pretty good but suffers some dropouts where MDX23C doesn't”)

uvronline.app (Mel-Roformer viperx' model, it is not flawless either)

- MVSep Pedal Steel Guitar

- uvronline.app (HQ_5 beta/paid users - places guitars in vocal stem pretty well, might got deleted)

“Rebalance volume of chans before processing” if you have better separation results processing L and R channel separately.


Consider using Apollo Universal by Lew (model | UVR | Colab in HTMYOR) to get more clarity after separation.

Acoustic guitar

- dango.ai (paid, probably the best for now, better in at least some songs than the SW model)

- MVSep Acoustic Guitar (strong competitor, outperforms moises “like crazy”

- uvronline.app (viperx' model for premium users - does a good job too)

- BS-Roformer SW

- BS-Roformer “MVSep Mega 53 Stems” model - 1.27GB | single models (less VRAM-hungry) - 1.27GB | MVSepless HF / HF CPU / Colab / newer (you can pick which stems you want there)

- Demucs 6s - sometimes, when it picks it up

- GSEP - when the guitar model works at all (it usually grabs the electric), the remaining 'other' stem often is a great way to hear acoustic guitar layers that are otherwise hidden.".

- lalal.ai (both paid)>moises.ai (It picks up acoustic and electric guitar together)

- Audioshake (both electric and acoustic)

- moises.ai

Separating electric and acoustic guitar

- MVSep SATB Choir<Phantom center model (on the channel with guitars)<guitar model<good instrumental separation (straturkoise)

Tip - For muddy results use EQ matching (e.g. free Spectrum Thief) from a not muddy fragment of separation, or other guitar play with the same guitar tone playing only guitar (straturkoise).

You could also simply a model from one of the two model categories from the sections above, e.g.:

- MVSep Acoustic Guitar (“it's separating acoustic from electric very well, even in fuzzy, lo-fi recordings” - Input Output)

- BS-Roformer SW (6 stems)/MVSep Guitar SW

Older

- “To separate electric and acoustic guitar, you can run a song [e.g. other stem] through the Demucs guitar model and then process the guitar stem with GSEP [or MVSEP model instead of one of these].

GSEP only can separate electric guitar so far, so the acoustic one will stay in the "other" stem.”

- “⁠medley-vox main vs rest model has worked for me to separate two guitars before”

- moises.ai “it's not perfect, it's good when the solo guitar for example is loud then it can be isolated but when it comes in a balanced lead and rhythm guitar, it can't isolate it”
- MDX23C phantom center model (it was non-giliaan’s model, but the latter might work better already)

- moises.ai (it has electric, acoustic, rhythmic, solo models)

Separating two electric guitars

- MVSep SATB Choir<Phantom center model (on the channel with guitars)<guitar model<good instrumental separation (straturkoise)

Lead and rhythm guitar

- moises.ai (paid)

- MVSep SATB Choir<Phantom center model (on the stem with guitars)<guitar model<good instrumental separation (works better than the single model below, but not for two rhythm or two lead guitars - straturkoise)
- MedleyVox

- MVSep Lead/Rhythm Guitar (1 stage, and 2 stage variant)

- MVSEP guitar models
“I can isolate both guitars with the different models that MVSEP has, especially in rock tracks where the lead guitar is in the center channel and the rhythm guitar is on the right - left side of a stereo track, good results are not always obtained, especially when the lead guitar has long delay effects, tons of reverb or when these effects go from one channel to another, but it also depends on how the song was mixed.” - edreamer 7

- BS-Roformer “MVSep Mega 53 Stems” model - 1.27GB (“rough”) | single models (less VRAM-hungry) - 1.27GB | MVSepless HF / HF CPU / Colab / newer (you can pick which stems you want there)

- MDX23C phantom centre extraction by wesleyr36 model
“First, isolate the guitar, then (...) use phantom centre extraction (...). Here I can find the rhytmn and the lead guitars, as I told before, results can vary” - edreamer 7

- Dry Paint Dealer Undr’s Melband Roformer and Demucs Lead and Rhythm guitar models.

“my own very mediocre model for it that I never shared. it does work but has issues that I imagine any better executed model won't.”

- lalal.ai (“it sucks” - isling, Oct 25)

Wind instruments and wind noises
(trumpet/saxophone/brass/woodwinds/flute/trombone/horn/clarinet/oboe/harmonica/bagpipes/bassoon/tuba/kazoo/piccolo/fluge/horn/ocarina/shakuhachi/melodica/reeds/didgeridoo/mussette/gaida/farts)

- MVSep Wind BS Roforomer (2025.09) 9.77 SDR (+2.64 SDR)

- MVSep Wind BS Roformer (2025.08) (more robust and cleaner than the Mel and detects instruments better, +2.5 SDR)
- Wind BS-Roformer on x-minus.pro by viperx (big step forward vs the old UVR model)

- MVSEP Wind SCNet

- MVSEP Wind Mel-Roformer

- MVSEP Trumpet (“so clean”)

“after testing [Wind 9.77] on a song where trumpet and sax play in unison, doing the trumpet model is cleaner than doing the sax model” - dynamic64

- MVSep Trombone

- MVSep Oboe

- MVSep Clarinet

- MVSep Harmonica (“Hit or miss” - musicbybrooks)

- MVSep French Horn

- MVSep Tuba

- MVSep Bassoon

- MVSep Accordion

- MVSep Brass

- MVSep Woodwind

- BS-Roformer “MVSep Mega 53 Stems” model | single models (less VRAM-hungry) - 1.27GB | MVSepless HF / HF CPU / Colab / newer (you can pick which stems you want there)

- has all types of the MVSep models above, just smaller/muddier [the 53 stems on MVSep is further trained], besides these two:)

- MVSep Bagpipes

- MVSep Braam

- MVSep Whistle

Older

- "Wind" model on UVR5 (Download Center -> VR Models -> select model 17)

(You might have to use it on instrumental separation first, e.g. with HQ_4 or Kim Inst)

- Audioshake

- Music.ai
- karaoke 4band_v2_sn on e.g. MVSEP (worse than Wind model in UVR)

- Adobe Podcast (denoiser and voice enhancer)

- Waves Voice ReGen (5 minutes free per day - denoiser and voice enhancer)
- Acon Extract: Dialogue 2 (update from around April 2026 added de-noise and de-reverb; it sounds better than Waves Clarity VX - nassemo.8805)

- Probably someone had some success with one de-crowd model for wind noises

- Lot of instrumental/vocal models confuses wind instruments with vocals

- Check out also the denoise section

MVSEP Saxophone

- SCNet XL (SDR saxophone: 6.15, other: 18.87)

- MelBand Roformer (SDR saxophone: 6.97, other 19.70)

- Ensemble Mel + SCNet (SDR saxophone: 7.13, other 19.77)

- The 53 stems MVSep model | single models (less VRAM-hungry) - 1.27GB | MVSepless HF / HF CPU / Colab / newer (you can pick which stems you want there)

Piano

Consider using Unwa BS-Roformer Resurrection inst a.k.a. “unwa high fullness inst" on MVSEP as preprocessor - rainboomdash

- Logic Pro (paid; May 2025 update, SDR 7.79) / BS-Roformer 6 stems / MVSEP Piano SW

(“1000 times more efficient than the Lalal.ai piano model”)

- Lalal.ai (paid; no other stem with piano stem attached)

- Demix Pro (paid; “I often combine the two” - Mixman, but it was before the 6 stems above)

- MVSep Piano Ensemble (Mel-Roformer  + SCNet Large piano models; SDR 6.21)
(Mel is viperx' iirc; “a [tiny] bit more bleed during the choruses and whatnot” vs x-minus, “works well maybe 7 times out of 10”; SCNet, “has a less watery sound, but more bleed” vs Mel)

- x-minus.pro (for paid users; cheap subscription; “more consistent than MVSep piano and demucs_6s” it knows well what piano is, but it sounds the best for other stem of piano separation, but e.g. on Carpenters - Yesterday Once More “while not terrible, the dropouts, underwater 'gurgles', and general lack of piano punch/presence remains noticeable” - Chris_tang1, while MVSEP Piano Ensemble: SCNet + Mel, SDR: 6.21, was much better in that case - might vary on a song)

- Music.ai (paid)

- Dango.ai (paid)

- GSEP (formerly best, paid)

- Moises (separate models for piano and keys)

- MVSep Piano MDX23C 2024 & 2023

- htdemucs_6s (not too good)

- MVSEP Digital Piano (much better for epiano, and sometimes also for real or synthesiser when it's picked vs the SW model)

- MVSep Keys

- MVSep Harpsichord

- MVSep Celesta

- MVSep Vibraphone

- MVSep Metal Bars

- MVSep Rhodes

- Piano in BS-Roformer “MVSep Mega 53 Stems” model | single models (less VRAM-hungry) - 1.27GB | MVSepless HF / HF CPU / Colab (you can pick which stems you want there)

- MVSep SATB Choir (it's able to separate piano layers; consider using good instrumental model result as an input passed through e.g. BS-Roformer SW piano stem, and then SATB for fuller sound instead of using “Extract vocals option”)

Synths

- MVSep Synth (it can also pick some bass which some bass models failed to pick up)

- MVSep Organ

- BS-Roformer “MVSep Mega 53 Stems” model (has these two; smaller/muddier) | single models (less VRAM-hungry) - 1.27GB | MVSepless HF / HF CPU / Colab / newer (you can pick which stems you want there)

- Piano or guitar models might work (if the song doesn't have piano or guitar already) depends on a song

- Zero Shot

- lalal.ai (hit or miss, sometimes might not work)

- Some voc/inst models might treat synths as vocals (then you could separate them using better inst/voc model)

Organs

- MVSEP Organ (surprisingly good, and since then SDR doubled since the first version, eliminating some bleed issues or e.g. Hammond organs not being picked in some places)

Idiophones

- MVSep Marimba

- MVSep Glockenspiel

- MVSep Triangle

- MVSep Bells (tubular bells or chimes - for sleigh bells use drums model)

- MVSep Wind Chimes

- BS-Roformer “MVSep Mega 53 Stems” model | single models (less VRAM-hungry) - 1.27GB | MVSepless HF / HF CPU / Colab / newer (you can pick which stems you want there, has all the above in smaller versions besides:)

- MVSep Xylophone

- MVSep Cowbell

De-crowd

Current models struggle with female screams and singing

- MVSEP Crowd MDX23C+Mel-Roformer Ensemble (6.07+6.06)

- MVSEP Crowd BS Roformer (SDR 7.21) (newer model, some people have issues with it: "The 6.27 one removed all kinds of crowd noises, sound effects and general noise, this one only removes random bits of music!")

- UVR-MDX-NET Crowd HQ 1 (UVR/x-minus.pro/model/conf in UVR) (can be more effective than Mel MVSEP’s sometimes; e.g. good for live shows)

- MVSEP Mel-Roformer Crowd by ZFTurbo (SDR 6.07) (MVSEP/files/UVR)

To use it in UVR, Go to UVR\models folder, and paste that folder there.

Then change "dim_t" value to 801 at the very bottom of: “model_mel_band_roformer_crowd.yaml” file in “mdx_c_configs” subfolder. Don’t use overlap above 4.

- Dango decrowd (paid)

- Mel-Roformer De-Crowd by Aufr33/viperx (x-minus.pro/UVR/DL/conf)
For UVR, change the model name to the one from the attached yaml, copy chkpt to models\MDX_Net_Models, and yaml to model_data subfolder, then set overlap 2 or use ZFTurbo inference script] - more effective than MDX below at times)

- MDX23C De-crowd v1/v2 (MVSEP)

- Older MSX23C MVSEP models (applause, clapping, whistling, noise)

 - Aufr33’s Mel-Roformer Denoise average variant (link | yaml | Colab) can be also used as crowd removal

- models for Harmonies - potentially for people singing along

- AudioSep

- USS Bytedance

- Zero Shot Audio Source Separation

- GSEP (sometimes), and e.g. drums stem is able to remove applauses

- Chant model (by HV, VR arch, e.g. works for applauses; may leave some echo to separate with other models or tools below) for Colab usage - you need to copy that model to models/v5 and then use 1 band 44100 param, turn off auto-detect arch and set it to "default". In UVR pick one of 44100 1 band parameter, possibly 512.

“For really difficult live songs (where the crowd is overwhelmingly loud to the point where you can't hear the band properly) sometimes filtering vocals with mel roformer on xminus THEN running the vocals stem through the mdx decrowd model sometimes helps”

SFX

- “You do need to first get an instrumental with a different model, because this isn't really trained to remove vocals. Just SFX” or speech.

- SFX models can be more aggressive than regular vocal models for speech.

Sometimes, some of the regular vocal models may turn out to be better suited for your task, so try out those for speech too.

- MVSep FX (dedicated for music SFX and vinyl scratches; might be potentially useful also after using HyperACE inst v2)

“Finally a model which can remove the cartoon sounds from the Rocky and Bullwinkle score! (..) works really well for removing sfx from old movies” - fal_2067

“works pretty well, I’m a fan” - dynamic64

Removes effects from music videos

FX is the effects, other is vocals and the music. - wancitte

- HyperACE inst v2 (Colab) - good to clean up vocals or voice from SFX, putting them in other stem - better than DNR models at it. A go-to for fun dubs, “SOTA for DNR (...) almost zero muddiness between the sound effects”, a bit better for it than the v1 - wancitte

- BS-Roformer SW drums (6 stems model, Colab, or just drums on MVSep) - “really good to remove some SFX and foley, way better than DnR v3” - erosunica


- MVSep DNnR v3:

a) SCNet

b) MelBand

(better metrics than Bandit v2)


- Bandit v2 (MVSEP | OG weights | yaml) multilingual model (“multi”)

with single ones for EN/GER/FR/SPA/CH/FAR language separately.
All single models converted for ZFTurbo inference | Colab (only multi and EN)
Available in UVR>MDX-Net>Download More Models>Bandit v2/Plus
(v2 is better on speech vs Bandit Plus, not always on SFX - experiment).

“Multilingual model is most of the time giving better results than French model for French content, so I would start with it” - jarredou

“Co-developed by Netflix and Georgia Institute of Technology. The paper is titled "A Generalized Bandsplit Neural Network for Cinematic Audio Source Separation"”

- Mel-Roformer “musicless” SFX and voice isolator by Jasper | MK Colab fork

It removes the music and leaves SFX and voice, so “if you want no voice, you gotta remove voices first or after.”

- mdx23c model_sfxsplitter

 model (yaml) by jazzpear/Jasper (224MB) | nextgen.uvronline.app/experimental

“separates SFX tracks into background and foreground SFX.

So say, you isolated SFX with DnR from a full mix or have a single all SFX track already. You run this and get separated version with BG and FG tracks for easier remastering of the mixed version SFX" - jazzpear

- MVSep Risers

- Sam Audio by Meta (Iirc, on mid side processed files, it got rid of some noise and knocking when separating SFX and instrumental with a prompt “music” and later with a prompt “knocking” on the result file “and selected the part at 0:35 and 0:36 for span prompting” - nicov_na)

- USS-ByteDance (for providing any, at least, proper sample)

- Zero Shot (worse for SFX vs Bytedance)

- Audiosep

- custom stem separation on Dango (paid, 10 seconds for free)
- iZotope RX 12 Scene Rebalance (paid; dialogue, music and effects stems)

Older models

- Bandit Plus (may work good for TV shows and movies: MVSEP, jarredou Colab | UVR
“trained on mono audio, so it's dual mono (...) when content is heavily panned left/right, it's where issues start to say "hello !"” but other than that it might still handle stereo good enough or even better than others depending more on a song)

- “My suggestions from those who want best results, i suggest using either:

Resurrection inst or instfv8 to extract Music and effects audio track” - killdubo

- Jasper (a.k.a. jazzpear94) MDX23C foreground and background SFXs - auxiliary/helper model (yaml) for DNRv3 and vocal model working for speech (it wasn't trained on voice and music; more info)

Mirror (yaml)

- Jasper Mel-Roformer BGM (background music) from a movie model (yaml) (stems: voice with SFX/BGM without singing voices), you might want to use some vocal model as a preprocessor here too.

Compatibility with UVR on both not guaranteed, consider using MSST in case of any issues

- jazzpear94 Mel-RoFormer model (ability to separate specific SFX groups - Ambiance, Foley, Explosions, Toon, Footsteps, Fighting and General - for all in one stem | MVSEP, fixed newer Colabfiles, mirror, instruction, prob. broken Colabs: 1, 2, 3, 4)

- joowon bandit model: https://github.com/karnwatcharasupat/bandit

(better SDR for Cinematic Audio Source Separation (dialogue, effect, music) than DNR Demucs 4 below (SDR 10.16>11.47) - Colab / MVSEP)

- GAudio (a.k.a. GSEP) announced their SFX (DnR) model in their API:
“DME Separation (Dialogue, Music, Effects)”
So far it’s not available for everyone on their regular site:

https://studio.gaudiolab.io/

But the link on their Discord redirects to the site with a form to write an inquiry:

https://www.gaudiolab.com/developers

Shortly after entering the one or both of the links and logging on the first, you might get an email that $20 of free credits to access their API have been added to your account

Older models

- DNR Demucs 4 model (repo: CDX23 repo, MVSEP, Colab) - it used to output fake stereo (at least training dataset was mono).

"I noticed [it] doesn't do well and doesn't detect water sounds, and fire sounds"

Can be used in UVR (the three stems will be labelled wrong, SFX will be bass).

>For UVR, download the model files, put them in the Ultimate Vocal Remover\models\Demucs_Models\v3_v4_repo, Delete "97d170e1-"  from all the three file names, copy this yaml alongside the model files (it won’t work on AMD 4GB VRAM GPUs).

>The Colab might run occasionally on CPU, and then it might be slow to the point that it might take 2.5h for a 15 min audio track (maybe change Google account or retry), and then it might take 2 mins for a similar length once it uses GPU.

- jazzpear’s MDX23C model (files) -  rename the config to .yaml as UVR GUI doesn't read .yml. You put config in UVR’s models\mdx_net_models\model_data\mdx_c_configs. Then when you use it in UVR it'll ask you for params, so you locate the newly placed config file.

- Aufr33 Mel-Roformer denoise average “27.9768” model - dedicated for footsteps, crunches, rustling, sound of cars, helicopters

If it’s not available for paid users of uvronline.app, use this link | MVSEP | model files:

Less aggressive & More aggressive | yaml file | UVR Roformer patch | Colab

- myxt.com (uses Audioshake)

- AudioSep (you can try it to get e.g. birds SFX and then use as a source to debleed or maybe try to invert phase and cancel out)

- Moises.AI (the rumour says it’s better than Bandit v2, but it’s expensive “Dialogue, Soundtrack, Effects”)

- Older DNR model on MVSEP from ‘22

I think the most commonly used recent SFX models discussed in the server before DnR v3 ones are DNR Demucs 4 model and Bandit v2, but I haven't seen any settlement in the community on which model is the best, hence it might simply depend on a song.

- voc_ft - sometimes it can be better than Demucs DNR model (although still not perfect)

- jazzpear94 model (VR-arch) - put the .pth file to: Ultimate Vocal Remover\models\VR_Models. On UVR start set config: 1band sr44100 hl 1024, stem name: SFX, Do NOT check inverse stem in UVR5, 5.1 disabled. Or put that file to model_data subfolder.

- (dl) source by Forte (VR)  (probably setting to: instrumental/1band_44100_hl1024 is the proper config) Might work in Colab “I tried it with the SFX models, and I just uploaded them in the models folder and then placed the model name, and it processed them” and may even work in UVR.

- Or GSEP (sometimes) esp. the new “Vocal Remover” model

Any other stem/instrument/sample if not listed above

- Zero Shot Audio Source Separation

- Bytedance-USS (might be worse for instruments, but better for SFX)

- Dango.ai custom stem separation (paid, free 10 seconds preview)

- Audiosep (separate everything you describe; Colab has unpickling issue)

- Spectral removers (software or VST):

Quick Quack MashTactic (VST), Peel (VST, they say it’s worse alternative of MT), Bitwig (DAW), RipX (app), iZotope Iris (VST/app), SpectraLayers (app, “Problem with RX [Editor's spectral editing] is it doesn't support working in layers non-destructively.”), R-Mix (old 32 bit 2010 Sonar plugin), free ISSE (app, showcase), FactorSynth, Zplane Copycat "but MashTactic also has a dynamics parameter that is really useful (you can isolate attack from longer sounds, or the opposite, coupled with the stereo placement and EQ isolation)"

RipX is “not as good as UVR5 for actual separation, but RipX is very good if you need to edit what's already separated more musically. SpectraLayers is a nicer spectral editor, RipX spectral editor is not as usable”

Consecutive multi-AI separation for not listed instruments

- Extract all other instruments "one by one" using other models in the chain (e.g. remove vocals with voc_ft or now e.g. inst Mel Kim derivative, use what's left to remove drums/bass with htdemucs_ft/MDX23/MVSEP ensemble, use what's left to remove guitars/piano with GSEP/demucs_6s or now any better purpose model, then use what's left to remove e.g. wind instruments (if present) with UVR wind model, or any any other purpose model applicable, even SFX, till you're left with the instrument of your choice, or as few instruments, as possible, for potentially easier work with spectral editor listed above)

- Drumsep - “Using DrumSep on melodic stems can help separate instruments easier if you plan on sampling/editing them, but they are separated based on range rather than actual instrument. Low instruments will often be on the kick/tom stems, mid instruments will be on snare and/or tom, and higher instruments will be on the cymbals.”

De-reverb

- anvuew BS-Roformer dereverb 22.5050 (dereverb_bs_roformer) a.k.a. “DeReverb stereo by anvuew” on MVSep or” | DL | uvronline. New 2026 model. Probably the best for now.

“Very interesting model, sounds like the Mel-Roformer one, but even more aggressive, it's good” it works on “reverb effect on vocals” at least - isling

“Sounds cleaner than the `mono sdr 20.4029` one” - rainboomdash

“The dataset is the same as Mel, all reverb is generated by VST plugins and includes Waves IR1 presets. So it depends on whether you consider IR1 to be room reverb” - anvuew

“anvuew fixed an issue that was causing it to not perform as well as it could - an alignment bug during training (...)

I guess the one labeled "mono" was also after the bug was fixed” - rainboomdash

- Mel-Roformer de-reverb by anvuew v2 (a.k.a. 19.1729 SDR) | DL | config | Colab
Former best de-reverb model.
- RX11's dialogue isolate - probably the best for de-echo for now
“anvuew's models can remove reverb effect only from vocals, “captures early reflections a little”. Old FoxyJoy's model works with a full track [regular songs].”

“sort of reminds of RX11 dialogue dereverb results but doesn't destroy singing voices”
“perfect for rvc”
Both BS and Mel variant “will also remove harmonies or vocal effects that are not in the center channel.”
“it works sort of like that phantom center model, removing sides basically”
Sometimes "noreverb" stem might get empty (e.g. on MVSEP, but similar issues was fixed already once there).

“reminds me of the equivalent de-reverb MDX model (...) cleaner in some ways, though slightly more filtered and aggressive.”

“it's EXTREMELY aggressive, like very aggressive, it seems kinda muddy at a lot of parts, almost NO reverb bleed, it also caught so many effects and removed them (good thing) which is actually insane! I also noticed that when it gets breathy or like it has falsetto, it seems to remove a lot of it, it's very weird at the breathy-ish parts of it lol, will be using this mainly if there are heavy vocal effects I want removing” - isling
In fact, it was “fine-tuned from Kim's Mel” - anvuew.

To make it work with UVR, delete “linear_transformer_depth: 0” from the YAML file, copy the model to MDX_Net_Models and config to model_data\mdx_c_configs.

- Dango Reverb Remover - click (“it's very similar to [RX11] dialogue isolate good/real-time set to 5. Yeah, it's like listening to the same inference files” John; probably also works in mono, you can get 30 seconds for free) but for other purposes, even older FoxyJoy’s models can give better results

Other models than the Dereverb Room below “tend to be more aggressive. But if they aren't working, add some reverb to your input file (make sure the reverb you add is stereo and at least roughly matches the character of the original reverb). That should trick the model into giving you a drier signal. The louder your added reverb is, the more aggressive the separation will be, so it's kind of a trade-off and takes some trial and error” - Musicalman

- anvuew BS-Roformer Dereverb Room model | Colab | MVSEP
(doesn’t work in UVR - it's a mono model, use MSST, if you have stereo errors using MSST on stereo files, update MSST [git clone and git pull commands] or see instructions later below.
“Specifically for mono vocal room reverb.” as most vocals are recorded in mono.

In most cases, it only decreases the reverb from the room, and doesn't remove it entirely, but there are exceptions. On songs, it can only decrease it.

Not that long inference compared to other Roformers.

“Really liking the fullness in the no reverb stem. Virtually all de-reverb Roformers I've tried sound muddy, but this one is just the opposite. (...) Other noises may interfere, and in my experience, makes the model underestimate the reverb. [The previous anvuew’s mono model] is way different [from] this one in every way. So, like I say, [it's] worth a shot.” - Musicalman

Note. Do the below to fix stereo error using that model. It might work with your current MSST version instead of the linked repo below too, but in a different line (fixed in all MSST by ZFTurbo installations from 2025).

“Edit inference.py from my repo line 59:

Replace :

        # Convert mono to stereo if needed

        if len(mix.shape) == 1:

            mix = np.stack([mix, mix], axis=0)

by :

        # If mono audio we must adjust it depending on model

        if len(mix.shape) == 1:

            mix = np.expand_dims(mix, axis=0)

            if 'num_channels' in config.audio:

                if config.audio['num_channels'] == 2:

                    print(f'Convert mono track to stereo...')

                    mix = np.concatenate([mix, mix], axis=0)”

    - jarredou

- Gabox Lead Vocal Mel-Roformer de-reverb | DL | config | Colab

“just use it on the mixture” - Gabox’

“I’ve had great results on vocals with heavy delay” - 5b

- anvuew dereverb_mel_band_roformer_mono_anvuew_sdr_20.4029 model | yaml | x-minus
“supports mono, but ability to remove bleed and BV is decreased” “separates reverb better than [the] v2” below.

- Sucial Mel-Roformer dereverb/echo model #3 called “fused”: model | yaml
“more effective in removing large reverb”

“Specifically targeting large reverb removal. After training, I combined these two models with my v2 model through a blending process, to better handle all scenarios. At this stage, I am still unsure whether my new models outperform the anvuew's v2 model overall [besides large reverbs].” More

- De-reverb MVSep Team (2026.07, BSRoformer) located here on the site | metrics

“It's a universal model, so you can remove reverberation from any stem or vocal track. Currently, it gives the best results for both of our validation datasets.  

Compared to anvuew room model it did nothing to the room echo - real_dyeu_kami

“I used it on an orchestral track (Ba'ku Village from Star Trek Insurrection), and it worked great. Lower volume reverb seems normal to me, as it's usually a distant reflection off walls etc.” - fal_2067

“I tried it on a full song at first just to see and the result was suuuuper quiet and not really accurate.

So then I tried it on a vocal track just now.

Once again the reverb is like 36db quieter than the vocals.

I can hear so much reverb in the dry part in my examples” - dynamic64

- anvuew v2 “less aggressive” variant - a bit lower SDR 18.81 | DL | config

Older
VR models

(added to UVR 5 (GUI), works in NotEddy’s Colab)

- UVR-DeEcho-DeReverb (213 MB) - “it removes reverb, not echo” but you could try it if everything above fails
“use an aggression of 3.0 -5.0 and nothing more than that.” e.g. 4 (0.4 in some CLI code Colabs).
“Results have a frequency ceiling around 17540 Hz and a very high pitched noise above 22000 Hz, you might want to upscale your results with HQNizer or Apollo Model” - more. Or use point #4 to slow down the audio and revert it back.

__

(below are the old ones, which might only work with vocal-remover 5.0.2 by tsurumeso’s default arch settings, [maybe 1band_sr44100_hl1024 or 512? and his nets and layers])

- VR dereverb - only works on tracks with stereo reverb (j48qny.pth, 56,5MB) (dl) (source)

- VR reverb and echo removal model (j48qny.pth, 56,5MB) (dl), works with mono/stereo)

__

MDX model (less aggressive than those at the top)

“I use it when there's not so much reverb, but if it's more intense I will choose VR-Arch DeEcho”

> FoxyJoy's dereverb V2 - works only with stereo (available in UVR's download center and Colab (or eventually via this dl link [dead]); capable of removing reverb from mixtures, not just vocals, but it can spoil singing in acapellas or sometimes removes delay or piano too). “I do think [that] MDX is noticeably more accurate [than VR DeEcho-DeReverb]”

"(the model is also on X-minus) Note that this model works differently from the UVR GUI. I use the rate change (but unlike Soprano mode only by 1 semitone). This extends the frequency response and shifts the MDX noise to a higher frequency range." It's 11/12 of the speed so x 0.917, but actually something else goes on here:

(Anjok)

"The input audio is stretched to 106%, and lowered by 1 semitone using resampling. After AI processing, the speed and pitch of the result are restored."

You'll find slowing down method explained further in "tips to enhance separation"

"De-echo is superior to de-reverb in every way in my experience"

“VR DeEcho DeReverb model removes both echo and reverb and can also remove mono reverb while MDX reverb model can only remove stereo reverb”

"You have to switch [main stem/pair] to other/no other instead of vocal/inst" in order to ensemble de-echo and de-reverb models.

Newer

- UVR Dereverb model by Aufr33 & jarredou for uvronline.app premium | Model files | settings (PS: Dry, Bal: 0, VR 5.1, Out: 32/128, Param: 4band_v4_ms_fulband)

Copy model file to Ultimate Vocal Remover\models\VR_Models (json config is probably already present in lib_v5\vr_network\modelparams and has the same checksum).


- MDX23C UVR Dereverb model a.k.a 6.9096 | uvronline.app premium | Model files | config

(by Aufr33 & jarredou)

Seems to pick up room reverb. Previous Foxy’s model sometimes cut “way too much” than this model.

Copy model file to Ultimate Vocal Remover\models\MDX_Net_Models and yaml config to \model_data\mdx_c_configs subfolder.

Bas Curtiz’ conclusion on both:

- MDX23C “seems to be cleaner, takes the reverb away, also between the words,

whereas (U)VR leaves a little reverb

- VR “seems to sound more natural, maybe therefore actually.

- MDX23C tends to 'pinch' some stuff away to the background, which sounds unnatural.

“This is just based on my experience with 3 songs/comparisons, but both points are a pattern.

Overall, they're both great when you compare them against the original reverbed/untouched vocals.” showcase video.

- You can find older anvuew Mel versions here

- BS-Roformer anvuew 8/256/8 variant (a.k.a. deverb_bs_roformer_8_256dim_8depth) | model | config (info with dead links) - a bit higher SDR than Mel v1 (8/26/6) posted firstly in the ZFTurbo repo.
“not good is all people say” but it might depend on a use case or a song.

Mel-Roformer might turn out to be better more often “works weirdly and leaves some echo for some reason”

“very usable for single singing voices and speech, because it's very precise in eliminating echo and reverb,

but if you have a choir singing or vocals with backing vocals in it, then it'll probably ruin it a bit. For such vocals it's better to use aufr33/jarredou model or dereverb deecho” - John UVR

To fix issues with BS variant of the model in UVR “change stft_hop_length: 512 to stft_hop_length: 441 so it matches the hop_length above” in the yaml file (thx lew), plus delete linear_transformer line in the config like above too.

- 8 384 dim 10 depth BS variant | DL | config

- #2 Sucial Mel-Roformer dereverb/echo model (model | MVSEP).
Fine-tune with more training data.

- #1 Sucial Mel-Roformer dereverb/echo model (model | MVSEP).
It’s good but doesn’t seem to be better than the anvuew's Mel v2 model above here.
Still, it might depend on a use case.

- older V1 de-reverb HQ MDX model by FoxyJoy (dl) (source) (also decent results, but most likely worse).

(“It uses the default older architecture with the fft size of 6144”

“After separation, UVR cuts off the frequencies at 15 kHz, so I found that to fix that is to invert the "Vocals" and mix that with the original audio file.”

Demonstration: Original | Dereverbed | Detected reverb)

- To enhance the result if necessary, you can use more layers of models to dereverb vocals, e.g.:

Demucs + Karaoke model + De-reverb HQ (by FoxyJoy)

"works wonders on some of this stuff".

“Originally I inverted with instrumentals then I ran through deecho dereverb at 10 aggression then demucs_ft then kim vocal 2 then uvr 6_ at 10 aggression and finally deecho normal” (isling)

- For room reverb check out:

Reverb HQ

then

De-echo models (J2)

“From my experience, De-Reverb HQ specifically only really works when the sound is panned in the center of the stereo field perfectly with no phase differences or effects or anything that could cause the sound to be out of phase in certain frequencies.

If the sound doesn't fit that criteria, it only accurately produces the output of whatever’s in the mid”

“I noticed that in some cases the DeEcho normal worked better than the aggressive, which was weird. That's why I ran through both, so to remove as much as possible.”

- For removing reverb bleed left over in the left and right channels of a 5.1 mix from TV shows/movies check out:

Melband Roformer on MVSEP

- https://twoshot.app/model/36 (AI, paid)

Free apps for de-reverb/de-echo/denoise

- Accusonus ERA (was good, but discontinued when Facebook bought them, can be found on archive.org from when they gaveaway it without DRM)

- Voicefixer (CML, only for voice, online)

- RemFX (de: chorus, delay, distortion, dynamic range compression, and reverb or custom)

- Reverser (specifically for “atanh” distortion when level matched, but occasionally worked for other types too)

- Krisp app (paid, free 60 minutes per day) better (same for RTX voice) - free on Discord

- Adobe Podcast (online, a.k.a. Adobe Podcast Enhance Speech, only for narration, changes the tone of voice, so you might want to use only frequencies from it above 16kHz)

- Waves Voice ReGen (5 minutes free per day)
- AI-Coustics (speech enhancement, 30 minutes/5 files for free per month)

- CrystalSound.AI (app)

- Noise Blocker (paid, 60 minutes free per day)

- Steelseries GG (app, classic noise gate with EQ and optional paid AI module, activating by voice in noisy environment may not always work correctly)

- RTX Voice (in NVIDIA Broadcast app, currently for any GTX or RTX GPU)

- AMD Noise Suppression (for RX 6000 series cards, or for older ones using unofficial Amernime Drivers)

- Elgato Wave Link 3.0 - Voice Focus feature (now free for everyone, standalone VST3/AU version is paid 50$)

- AI SWB Noise Suppression (free, currently they give away that Mac/Windows driver only on email requests)

- Audio Magic Eraser shipped with new Google Pixel phones (separate options for cancellation of: noise, wind, crowd, speech, music)


See also denoising plugins

The best paid de-reverb plugins for vocal tracks/stems/separations

- RX 11 Dialogue Isolate (RX Editor/VST, paid) - some people like it more than DeVerberate 3. In RX Advanced variant, there’s additionally “multi-band processing and a high-quality mode as an offline process”, good companion for de-echo along with anvuew v2 model for dereverb

- DeVerberate 3 by Acon Digital (someone while comparing said it might be even better than RX10) "I find it's useful to take the reverb only track and unreverbed track and mix them to a nice level" “Acon is probably best if you can tweak to each stem separated. RX is imo too rough.” Comparison

- Accentize DeRoom Pro ("great" but expensive, available in DxRevive Pro, now 1.1.0)
- prime:vocal (multitool with also dereverb and other vocal enhancers)

- DxRevive Pro 1.1.0 - complete dialogue restoration tool; noise removal, reverb suppression, restoration of absent frequencies, elimination of Codec Artifacts

- Izotope RX <?8-10 Dialogue De-Reverb (RX Editor/VST) for voice and mixtures

(more possible free solutions). Good results not only for room reflections, but also regular reverb in vocals. It picks reverb where even FoxyJoy's model fails (“De-reverb” and “Dialogue de-reverb” options). It’s destructive for mixing raw vocals, but can just work.

- Clear by Supertone “equally good compared to RX10 imho. Smoother imho. It's only good on vocals though” Simple 3 knob plugin - “the cleverest / least-manual to get good results and is AI-based.” previously known as Supertone Voice Clarity and defunct free GOYO.AI) Also destructive for mixing raw vocals, but can just work.

- Acon Dialogue:Extract 2 (voice de-reverb and de-noise added in some update; sounds better than Clarity VX - nassemo.8805 [at least the non-DeReverb variant of VX])

- Waves Clarity Vx DeReverb - cannot perform de-echoing, so you need UVR De-echo (17.7kHz cutoff) or RX Dialogue Isolate for it; simpler than RX (paid; models updated in 12/17/2023 build) - same. Maybe you could even mix the two plugins using less aggressive settings in both.

Others:

- SPL De-Verb Plus

- Audio Damage Deverb

- Zynaptiq UnVeil

- Zynaptiq Intensity

- Thimeo Stereo Tool (one of its modules, not so aggressive/capable like Silk Vocal, can be even used on mixbus lightly)

- Cedar StageVox (voice dereverb and denoiser capable of working in real-time for tracking)

- Cedar VoicEX 2 (voice dereverb and denoiser)

- Waves Silk Vocal (voice has optional Live version too)

- Noiseworks VoiceAssist (doesn't give very dry vocals, but it also has some highs recovery option for them after de-reverb, not always free of artefacts, but very decent)

If you want to use some of these DAW plugins for your microphone in real-time, you can use Equalizer APO.

Go to "Recording devices" -> "Recording" -> "Properties" of the target mic -> "Advanced".

To enable a plugin in Equalizer APO select "Plugins" -> "VST Plugin" and specify the plugin dll. You need a VST2 x64 plugin for the x64 app version, VST3 is unsupported, but using a separate chainer/wrapper/adapter from VST2 to VST3 is still possible.

To run a plugin for a microphone in a simple app and send it to any output device, alternatively, you can download savihost3x64, then edit downloaded exe name to the name of your plugin you want to use, placed nearby, and run the app. Now go to settings and set input and output device (can be virtual card, maybe not necessarily). Contrary to Equalizer APO (which supports only VST2 x64 or x86 for x86 app version) it supports VST3 plugins too. Alternatively you can also check free Cantabile Lite (it incorporates JBridge for 32-bit VST2 plugins). Of course, you can also use DAWs for the same purpose (Reaper, Cakewalk etc., and Audacity maybe received some real-time VST support in newer versions, before it was only for offline use)

More free tools:

https://github.com/FORARTfe/HyMPS/blob/main/Audio/AI-Enhancing.md#dereverbers-

De-echo

- RX11's dialogue isolate (paid) - probably the best for de-echo for now

- UVR-De-Echo-Aggressive (121 MB)

- UVR-De-Echo-Normal (121 MB)

- UVR-DeEcho-DeReverb (213 MB)

(now added in UVR and MVSEP, won't be in Colab for now, but the first too are on HuggingFace)

- delay_v2_nf2048_hl512.pth (by FoxyJoy, all VR arch, source, can't remember if it was one of the above), decent results.

“works in UVR 5 too. Just need to select the 1band_sr44100_hl512.json when the GUI asks for the parameters”

“You [also] can use this command to run it: python inference.py -P models\delay_v2_nf2048_hl512.pth --n_fft 2048 --hop_length 512 --input audio.wav --tta --gpu 0”

They’re also on X-Minus now:

“The "minimum" and "average" aggressiveness settings use the Normal version of the model. The Aggressive one is used only at the "maximum" aggressiveness.”

“What's crazy is maximum aggressiveness sometimes does better at removing BG vox than actual karaoke models”

De-noising (vinyl noise/white noise/general)

Faster/quicker/not so efficient 

- Denoise standard in UVR serving mainly for MDX v2 noise (like in HQ_1-5, iirc it uses HV code; how it works: it separates "twice, with the second try inverted, after separation reinverted, to amplify the result, but remove the noise introduced by MDX, and then deamplified by 6dB, so it still the same volume, just without MDX noise.”

- Denoise model in UVR (it’s using VR’s UVR-DeNoise-Lite, 20kHz cutoff)

(Options>Choose Advanced Menu>Advanced MDX-Net Options>Denoise output)

model dedicated also for filtering noise existing in almost all MDX-Net v2 models in silent or quiet parts, but potentially also for more applications

- Min Spec ensemble of denoise model and denoise disabled results in Advanced MDX-Net Options (Audio Tools>Manual ensemble>Min Spec)

Filters more MDX noise in quieter parts than the denoise standard and the denoise model dedicated option in UVR.

Optimal, all-arounders. for any application, slower
- Phase-corrected de-noise method by n0blink/hitto1st (MVSep paid preset | workflow) -
iirc it uses the model below, with phase from the Aufr’s model aggressive variant

- Gabox Mel denoise/debleed model | yaml | Colab | MVSEP

For noise from fullness models (tested on v5n) - it can't remove the vocal residues - try out denoising on mixture first, then use fullness model.

“It can preserve slightly more high-frequency content in speech [than Aufr33 Mel model]” - Musicalman. “Quite a bit slower”. Impressive “ability to clean up noisy vinyl or cassettes” (padybu), “from my testing it's been better than aufr33's [Mel]” (pipedream, _0nshuk)
It can remove drums in some specific cases - then isolate drums first.
It might remove claps from live videos.

- Mel-Roformer Denoise by Aufr33 | Colab | MVSEP | links later below

a) minimum aggressiveness model called “27.9959” a.k.a. “1”

- good for white noise/static noise

b) average “27.9768” a.k.a. model “2” or a.k.a. “aggressive”
- works for footsteps, crunches, rustling, sound of cars, helicopters, VHS rips

Some people like to use overlap 10 with these models.


minimum - removes fewer effects such as thunder rolls, scratching or sweeping surfaces. It’s not as good at removing louder MDX noise when using AMD GPU instead of CPU on older system’s DirectML.dll in UVR

average a.k.a. aggressive - usually removes more noise than minimum, and also occasionally slight reverb/echo from room in vocals

“The Mel-RoFormer denoise model is amazing at removing 78 RPM record crackle”

It’s much better at higher frequencies than the VR model below (it doesn’t damage them that bad, it's less destructive, but also less aggressive).

~”Incredibly useful as mixing tools, it can pull all kinds of hum out of raw vocals, guitars, room mic's, bass, etc. before mixing with zero artifacts left over”.


If it’s not available for paid users of uvronline.app, use this link | model files:

Less aggressive & More aggressive | yaml file | works with UVR Roformer patch | MSST

For UVR - use Install model option in MDX-Net, or copy ckpt files to models\MDX-Net folder and yaml to model_data\mdx_c_configs subfolder. Choose the new model, press yes to set parameters, enable Roformer option, pick the config file corresponding with the copied yaml name. In case of “use_amp” error (e.g. in MSST), add “use_amp: true” in the yaml under optimizer and other_fix lines.

“From most aggressive to least:

VR Denoise

VR Denoise Lite

[Aufr’s] Mel-Rofo Denoise Aggr(essive)

[Aufr’s] Mel-Rofo Denoise”

- Bas Curtiz
Rather RX11 Spectral Denoise can be more aggressive than all of them at certain settings, and RX12 might be even better at certain cases.

- Apollo Lew Uni model - tends to smooth out some even consistent noise in e.g. higher frequencies, making the spectrum more even there “tends to "clean" audio noise and flatten the sound a bit.” - CC Karaoke/DtN

- UVR De-Noise by aufr33 “minimum aggressiveness” on x-minus.pro/uvronline.app (for premium or using this link) 

(less aggressive than denoise model in UVR,
“The (...) model is designed mainly to remove hiss, such as preamp noise. For vocals that have pops or clipping crackles or other audio irregularities, use the old denoise model“. Grabs “sound effects in old recordings (radio drama”, might make “soft voices sound weak”).

- UVR De-Noise by aufr33 “medium aggressiveness” on x-minus (same as default for free users) - it seems to be even less aggressive than UVR-DeNoise-Lite in UVR

- Mel-Roformer De-Crowd by Aufr33/viperx (x-minus.pro/UVR/DL/yaml)

(“to remove background noise when denoise models were failing [not sure if it was rain or wind”, can remove vinyl noises])
For UVR, change the model name to the one from the attached yaml, copy chkpt to models\MDX_Net_Models, and yaml to model_data subfolder, then set overlap 2 or use ZFTurbo inference script] - more effective than MDX below at times)

- yxlllc’s harmonic noise separation VR model (can be used in UVR: rename “model” to some model name, and pt extension to pth, then use Install model option and set config settings to: VR 5.1, 32/128, 1band_sr44100_hl512; “very good at further remove the noise from a dereverbed vocal yet it is mono. (...) it did have two channels on export but both of them have audible information except one has only some noise that can be easily removed then mix the other channel up to stereo” - mohammedmehditber. Maybe attached CLI code will have better mono model handling)

- mel_band_roformer/mbr_denoise_yuluoye (yaml)

- denoise children phaedrus33 model | yaml

- Mel Rifforge by mesk model final | Colab (might potentially serve also as denoiser, because it can pick up original mixture noise, and filter it out along with vocals)

- Sam Audio by Meta - Iirc, on mid side processed files, it got rid of some noise and knocking when separating SFX and instrumental with a prompt “music” and later with a prompt “knocking” on the result file “and selected the part at 0:35 and 0:36 for span prompting” - nicov_na

- Using de-limiter on mixture before separation with multiband gating with sidechain (alleviates noise in Suno songs - ~“the trick is handful in proactive gating (...) current denoise models don't seem to track it. it only removes resonance and sometimes important frequencies” - mohammedmehditber)

Vocal models as denoisers

- Unwa BigBeta 5e and 6 (5e “good when your mic/pc makes a lot of noise. All the denoise models are a bit too harsh for ASMR” - gilliaan, both “for denoising a conversation, was better than: UVR Denoise, MDX23cInstVocHQ, HQ5, KimVocal2, VocFT, Apollo, MelRof-aufr33-Denoise, GaboxDenoise, BanditV2cinematic, ViperX-BSrof-1297” - mixamillion) | Colab | MVSEP | Colab | MSST-GUI | UVR instruction | Model | yaml: big_beta5e.yaml | fixed yaml for AttributeError in UVR

- MVSep PolarFormer 124 bands (“It picks up microphone noises very well” - pezz23)

- BS-Roformer viperx 1296 / MVSEP BS-Roformer 04.24 / Gabox BS_ResurrectioN (denoising and derumbling working the most efficiently on vocals here too) x-minus/Colab/UVR/MVSEP

- Kim Mel-Roformer (works for denoising and debleeding vocals well)

- Vocal model like Voc_FT or Unwa’s ft2 bleedless (“it can sometimes isolate the vocals without the noise. And has a better result than a normal denoiser model” - Kashi)

- Mel-Roformer Karaoke (by aufr33 & viperx)  (to remove noise from a dialogue, mostly rustling in the background)

x-minus.pro / uvronline.app / mvsep

model file (UVR instruction)
- Mel-Roformer Duality model (excellent for pops and clicks in mixture to get clean vocals out of 45 RPM vinyl mixture - bratmix)

Other tools

- resemble-enhance (available on x-minus, but only as denoiser for voice/vocals, and on HuggingFace, site; works good for wind/outside noise)

- tape.it/denoiser - (“great tool for removing tape hiss. Seems to be free without limitation at this point in time, though it seems to have issues with very large files [20 mins etc])”

- crowdunmix.org/try-rokuon/

- github.com/eloimoliner/denoising-historical-recordings (mono, old 78rpm vinyls, fixed Colab, sometimes deletes SFX, but not as much as UVR De-Noise by aufr33 in old recordings)

- audo.ai

- github.com/sp-uhh/avgen

- github.com/Rikorose/DeepFilterNet | Huggingface (for speech)

- studio.gaudiolab.io (new Noise Reduction feature)

- possibly USS-Bytedance (when similar sample provided)

- Various AI tools - list by FORARTfe/HyMPS | #2

- Free apps

Older models

- UVR-DeNoise (trained by FoxJoy) - DeNoise-lite above is less aggressive

You can use negative values in UVR for that model. -20/-25 - for cleaning vocals

-10/-15 - when some vocals are gone - Gabox
“It's decent, but it needs a little work compared to" RX 10 spectral denoise.

- voc_ft - works as a good denoiser for old vocal recordings

- GSEP 4-6 stem ("noise reduction is too damn good. It's on by default, but it's the best I've heard every other noise reduction algorithm makes the overall sound mushier", it’s also good when GSEP gives too noisy instrumentals with 2 stem option, it can even cancel some louder vocal residues completely)

- UVR-MDX-NET Crowd HQ 1 (UVR/x-minus)

- This VR ensemble in Colab (for creaking sounds, process your separation output more than once till you get there)

Plugins (different types of noise)

Free

- Guide for classic denoiser tools in DAW, e.g. for debleeding (Bas Curtiz):
Tips & Tricks

- Bertom Denoiser Classic (or paid Pro)

- Accusonus ERA 6 (released for free after FB acquisition) - bundle with also de-esser, voice auto-EQ, voice leveller (better than soothe2 for de-essing for some people), deplosive, declipper and more

- Noise Suppression for Voice (a.k.a. RNNoise, worse, various plugin types, available in OBS; now also RRNoise 0.2/1.10 available)

- Airwindows DeNoise - multiband noise gate with various controllable bands
- Alt-Denoiser - “real-time AI audio noise suppression plugin based on DeepFilterNet”; “ I was never impressed by DeepFilter 1 or 2 but this plugin is using version 3 which seems to be an improvement.” - Mogwai

Paid

- Izotope RX 10 Spectral De-noise (“I think RX 10's Spectral De-noise is better at removing the noise MDX [model] makes”)

Actually, the new UVR De-noise model is really good when you combine it with RX 10's Spectral De Noise”, better than Lab 4 and current models, also more tweakable, but takes more time to set (now also RX 11 available - should be even a step forward)

- Acon Restoration Suite 2’s DeNoise “is decent if you can build a good noise profile with the Learn option, I like to have a few in series set to do -3dB of NR.” - theophilus3711

- SOUND FORGE Audio Cleaning Lab 4 (formerly Magix Audio & Music Lab Premium 22

[2016/2017] or MAGIX Video Sound Cleaning Lab - basically the same stock plugin across all of these versions)

- Unchirp VST (for musical noise, artefacts of lossy compression)

- Izotope Dialogue Dereverb (it is also denoiser)

- Izotope Dialogue Isolate in RX11

- Waves Clarity Vx / Pro (designed mainly for vocals)

- Brusfri by Klevgrand

- prime:vocal (multitool with also dereverb and other vocal enhancers)

- DxRevive Pro (mainly for dialogue: denoiser, declipper, dereverb, enhancer, codecs artefacts removal)

- Acon Dialogue:Extract 2 (dereverb, denoise)

- Cedar StageVox - (dereverb and denoiser capable of working in real-time for tracking)

- Waves Silk Vocal - (has optional Live version too)

Visit also Debleeding/cleaning e.g. inverts

Bird sounds

- Google's bird_mixit (code & checkpoint for their bird sound separation algo; more)

De-reverb models, e.g.:

- UVR-DeEcho-DeReverb (doesn't work for all songs)

Vocal models, e.g.:

- MVSEP BS-Roformer 2025.07 (if you already have birds in a vocal stem, as most vocal models do iirc, that may do the trick)

SFX models

Zero shot solutions:

(you can try them to get e.g. birds SFX and then use as a source to debleed or maybe try to invert phase and cancel it out)

- AudioSep

- USS-ByteDance (for providing any, at least, proper sample)

- Zero Shot (currently worse for SFX vs Bytedance)

- custom stem separation on Dango (paid, 10 seconds for free

Technically, if bird noises are in vocals, then equally:

- RTX Voice,

- AMD Noise Suppression or even

- Krisp and

- Adobe Podcast

might get rid of them, but at least the last changes the tone of voice, and the previous may work good only with voice instead of vocals.

Spectral editing

De-clippping/de-limitter/de-compression of dynamics (for loud or brickwalled songs/sounds with overly used compressor/clipper/limiter/distortion - transients/peaks recovery)

- jeonchangbin49’s De-limiter (HF | Colab) (“if you have any squished tracks that Apollo doesn't handle well, try passing it through that AI de-limiter first” - macularguide

“I get less distortion on vocals if I delimited first on a loud mix” - 5b

“this is amazing, it worked wonders for me” [for mixture de-limitting/de-compression, waveforms] - straturkoise

Parameters: “parallel mix - “defines how the normalized input and the de-limited inference will blend together. Where 0 is 100% normalized and 1 means 100% de-limited” - santilli_
For Colab don’t provide the file name in the input field, leave just directory as it was, it will scan its content)

- Ozone’s 12 Delimiter (waveforms) - paid, works also for mixtures, adjustable parameters

Free de-clipper AI tools (not plugins)

- RemFX (contains model to get rid of distortion and compression; “mostly for singular sounds, it won't work for whole mixes like songs”)

- Rukai (for speech and instrumentals)

- Amis (mainly for speech)

- stet-stet’s DDD (speech, req. decent CML knowledge to setup)

- Hendrix-ZT2 pyaudiorestoration’s Spectral Expander - for single band compression or aggressive tape AGC (more)

More GH repos - HyMPS list | #2

Free de-clipper plugins:

- ReLife 1.42 by Terry West (works best for stereo tracks divided into mono, newer versions are paid)

- ERA 6 declipper (released in bundle for free after they were bought by Meta)

- Airwindows AQuickVoiceClip - mainly for streamers yelling into the microphone “It’s not a ‘un-clipper’ but it tames the distortion a bit.”

Paid de-clipper plugins: 

Acon DeClip, ProAudioDeclipper, Declipper in Thimeo Stereo Tool (a.k.a. Perfect Declipper - standalone; both free for Winamp), iZotope RX De-clip (in RX Editor or as plugin), Pure D compressor by Flux Audio, HMD Uncompressor, FX Factory De-Clipper, DxRevive Pro (mainly for dialogue, also denoiser, dereverb, enhancer, codecs artefacts removal), Declipper in Magix/Sound Forge Cleaning Lab, Adobe Audition’s Declipper, sometimes even Fabfilter Pro-MB multiband compressor might be useful

See comparison


Clippers (the opposite, but useful in the whole mastering chain, sometimes in a tandem with the above in the whole chain):

Free

- KClip Zero

- FreeClip (sometimes you can use both in the same session for interesting results)

- GClip

- Limiter6 by vladg (Clipper module)

- Initial Clipper
- Airwindows Hypersoft  - “a more extreme form of soft-clipper”

- Airwindows OneCornerClip - compared to OG ADClip, it retains the character of sound

- Airwindows ADClip8 - “loudenator/biggenator”

- Airwindows ClipOnly - “2-buss safety clipper at -0.2dB with powerful anti-glare processing.”

- Hornet Magnus Lite - clipper and limiter modules

- Razor Clip

- Neutone FX>Clipper, actually an AI plugin (instruction), can be more resource-hungry



Paid: Orange Clip 3 (multiband mode), Gold Clip (widely praised lately), Gold Clip Track, Soundtheory Kraftur, KClip 3, SIR Standard Clip (popular, though KClip 3 may give better results), Izotope Trash 2, DMG Tracklimit, TR5 Classic Clipper (great for a kick), KNOCK (hard & soft clipper), Boz Little Clipper 2, Flatline (clipper), Newfangled/Eventide Saturate (spectral clipper), JST Clip, Brainworx Clipper, Elysia Alpha Mastering Compressor (soft clip module), soft clipper in Cubase, Music Hack Fuel, Music Hack Fuel Clipper (saturation, dynamics, limiting/soft clipper)

De-expliciter (removes explicit lyrics from songs)

https://github.com/tejasramdas/CleanBeats (more recent fork)

De-breath

- Sucial de-breath VR v1/2 models

- Aspiration Mel models by Sucial | config | MVSEP (one variant) (“grabs a lot more than just breaths and other sounds too, de breath gets ONLY breaths” - isling)

- yxlllc’s harmonic noise separation VR model (can be used in UVR: rename “model” to some model name, and pt extension to pth, then use Install model option and set config settings to: VR 5.1, 32/128, 1band_sr44100_hl512;

“it's really useful while making covers and when an OG song is airy/whispery. I'm not the sharpest tool in the shed so I resort to using websites like this” - wancitte

“very good at further removing the noise from a dereverbed vocal yet it is mono. (...) it did have two channels on export but both of them have audible information except one has only some noise that can be easily removed then mix the other channel up to stereo” - mohammedmehditber. Maybe attached CLI code will have better mono model handling)

- Accusonus ERA Bundle (free/gave away plugin after FB acquisition) download

- Dead Duck (free breath removal/gate plugin)

- Noiseworks VoiceAssist (paid plugin supporting ARA, having its own audio editor inside the session highlighting all the breaths allowing to turn them down automatically or with specific threshold for each)

- Izotope RX11’s breath control (paid; VST/Audio Editor)
(“The “remove breaths” preset they have on it Usually works about 95% of the time for me” -5b)

- DNR v3 (Sometimes (...) (without vocal help), the grunts and breathing will be in the SFX, and the dialogue in the speech, while both will be in the music) - fal_2067

- Mel-Roformer de-reverb by anvuew v2 (a.k.a. 19.1729 SDR) | DL | config | Colab

(“when it gets breathy or like it has falsetto, it seems to remove a lot of it, it's very weird at the breathy-ish parts of it lol, will be using this mainly if there are heavy vocal effects I want removing” - isling”)
- MDX23C-InstVoc HQ (“ in some cases, it also removes some airy parts from specific words, and some non-verbal sounds (breathing, moaning).”

_____

Manipulate various MDX settings and VR Settings to get better results

____

Final resort - specific tips to enhance separation if you still fail in certain fragments or tracks

____

Get VIP models in UVR5 GUI (optional donation) - it's if you can't find some of the listed above or in top ensembles chart:

https://www.buymeacoffee.com/uvr5/vip-model-download-instructions

(dead links)

List of VR models in UVR5 when VIP code is entered (w/o two denoise by FoxyJoy yet):

https://cdn.discordapp.com/attachments/708595418400817162/1104424304927592568/VR-Arch.png

List of MDX models when VIP Code is entered (w/o HQ_3 and voc_ft yet and MDX23C):

https://cdn.discordapp.com/attachments/708595418400817162/1103830880839008296/AO5jKyQ.png

More updated list can be found in that UI:
https://huggingface.co/spaces/TheStinger/UVR5_UI

(some models might be not from Download Center/VIP code)

Models repository backup of all UVR5 models in separate links

https://github.com/TRvlvr/model_repo/releases/tag/all_public_uvr_models

Some models might be not available in the repository above, as e.g. 427 model which is available only after entering VIP code.

(just in case, here's the link for 427:

https://drive.google.com/drive/folders/16sEox9Z_rGTngFUtJceQ63O5S9hhjjDk?usp=drive_link

Copy it to UVR folder\models~MDX folder and rename the model name to:
UVR-MDX-NET_Main_427)

_____________

Q: “Hello, we are now getting very good results in turning music that includes human voice into only instrumental. Sometimes there are vocal leaks that we can call just crumbs or whispers, but this is not that important. But now we have another important problem. People who do not want to listen to vocals, that is, who only want to listen to the music that remains when the vocals are deleted, encounter a problem. Sometimes there are big gaps in the songs. Because not every song is arranged in a way that continuous instrumental music is heard, and when the vocal part is deleted, a perception of silence or emptiness can occur. It is as if the music does not have continuity, and everything is cut off in some parts of the song. The reason for this is that when the vocal is deleted, the vocal melody is also destroyed. Although it seems like a good idea at first, when we listen to music that is only instrumental with the vocals deleted, that song loses a lot of its identity. As a result, I want to learn how we can preserve the vocal melody after deleting the vocal. What I mean is, can we divide the song into instrumental and vocal and then turn the melody of the vocal part into an instrument such as piano, bass guitar, flute, etc. Then, I want to combine this vocal melody with the instrumental result.” - sweetlittlebrowncat

A: You could try out some older, less aggressive models than Roformers. Even GSEP.

They can sometimes leave some melody from vocals (in fact, some quiet harmonies), so the song is not so "dead" after separation. Actually, you could try to separate vocals into separate stems to look for something useful to mix with the instrumental quietly.

Check Vocal models, then separate further with BV/Karaoke models or alternatively check GSEP, MDX-Net and maybe even VR models. Open document outline of this document and there you have all the interesting sections.

Also, you can use:

"https://audimee.com/

split instrumental from vocal

use vocal as input

convert it into piano, bass, flute, whatever they offer

merge

profit" Bas Curtiz

_____________

Mixing/mastering

If you already did your best in separating your track, tried out ensembles or manual weighting, also read tips to enhance separation, but if it lacks original track clarity, you can use:

- Demudder added in the beta Roformer patch #14 in UVR (if it won’t increase vocal residues too much; won’t work with even small chuk_size in AMD/Intel 4GB VRAM GPUs with Roformers)

- AI Mastering services (mainly for instrumentals)

- For improving vocals' clarity, you could even “train a RVC model out of clean [artists] audio clips and then inference this audio with the model you made. It takes some time, but the results are worth it” John UVR (examples). Workflow explained later below.

- Aufr33’s expander template for Reaper 7.05 (DL) fixing ducking in instrumentals (explained later below)

- Read Make your own remaster (more below)

Inherited flaws of instrumentals from AI separation models

Lots of audio engineers recommend leaving space for vocals in production. The result is, lots of instrumentals, no matter how good the model is, might have a hole in the midrange.
Sometimes even adding vocals to the mixture makes psychoacoustic effect making an impression that the instrumental sounds fuller. Similarly to the fact that adding artificial noise can increase fullness metrics of models. On headphones, or by sticking one of your smartphone speakers to your ear, you might identify the constant buzzing of models easier, while even on nearfield monitors it might just make an impression that the model is fuller, and the noise won’t be noticeable. Also, current models have troubles with distinguishing noise of the original mixture as background of instrumentals with vocals. Buzzing pop-ins are sometimes remnants of constant noise in the mixture, sometimes louder, sometimes quieter - in places when vocals appear and disappear.
Some model results will be muddy, or also too noisy, even for the best model result or ensemble. You need a result which is not destructive for any of the instruments in order to be able to be picked by a specialised single instrument model to mix it further. If you already have good instrumental result:

Mixing track from scratch using various AIs/models

Now if you're not afraid of mixing, and e.g. if you have clear instrumental already or whole track to remaster, then you can use the following for such a task:

- very quiet mixture (original file; so instrumental mixed vocals - if you remaster whole OG song)

- stems from demucs_ft or BS-Roformer SW (both MDX23 Colab or Ensemble of various models on MVSEP can be even better than Demucs, and vs SW esp. for bass - check out 4 stems section) mixed with also:

- drumsep MDX23C free model result (but you can also test out stems from the old drumsep and LarsNet [although they have worse SDR], or newer MVSEP drumsep models)

- GSEP result for piano or guitars (MVSEP models can be handy too, now the SW model for those stems are much better, and previously the only downloadable decent guitar model released becruily, and demucs_6s is mediocre, now we have SW)

- for bass both GSEP and Demucs ft/MDX23 aligned and mixed together (or simply from MVSEP ensemble or MDX23 Colab) or see bass models for more recent list incl. the SW

- "other" stem could be paired like above too (but drums remained only from e.g. Demucs_ft - they were cleaner than GSEP and good enough)

- Actually in one of those guitars weren't recognized in guitar stem, but were in other stem, so I mixed that all together (it wasn't busy mix)

- If it's not instrumental, probably mixing more than one vocal model might do the job, check various vocal ensembles (but it’s essentially what MDX23 and ensembles on MVSEP do, but the latter with private models, it’s not exactly the same - you can add different effects for every of such tracks, having fuller sound and change their volume manually).

The all above gave me an opportunity for a very clean mix and instruments using various plugins while setting correct volume proportions vs mastering just instrumental separation result or plain 3 stems from Demucs.

For example, demucs_ft or other single or incorporated drums model provides much higher quality of drums than the old Demus drumsep during mixing, so in such case you won’t use its stems on its own, but you will use drumsep more to overdub the specific parts of instruments more (e.g. snares - that’s the most useful part of using drumsep as normally it’s easy to bury snare in a busy mix when hi hats kick in overly in a heavily processed instrumental stem or drums stem - not you won’t have to push drums stems from demucs_ft or MDX23 so drastically).

Sam Hocking’s method for enhancing separated instrumentals from a mixture (song containing instrumental and vocals):

“I think looking at spectrally significant things like snares can work. We can already do it manually by isolating the transient audio/snare pattern as MIDI and then triggering a sample from the track itself to reinforce, but it's time-consuming and requires a lot of sound engineering to make it sound invisible.”

You can probably use Cableguys Snapback plugin for that, or maybe UVI Drum replacer.

Sam’s method will work the best in songs with samples instead of live recordings (if the same sounds repeat across the whole beat). More of those plugins.

PS. In the late 2025 we received an info about Apple Music rejecting Atmos mixes of some legacy music made with separation models (ensembles, and then probably some for 4-6 stems), even though the mixes sounded good, and were accepted by labels and artists. Also, we know that these separation methods worked fine in the past, at least for some other engineers. We suspect that they might use automated tools catching specific artefacts usually seen in separation models on spectrograms of extracted channels from the whole Atmos mix.

“I don't think there's a need of really advanced and expensive method to detect source separated stems, most of the time, just looking at the background noise is enough to tell, original stem vs separated one [click]

+ kind of "aliasing" artifacts and/or dither residues popping here and there...

There are lots of patterns than can make separated audio stems identifiable, I don't think it's hard to develop a model to spot them with quite good accuracy (even if not really audible to human ears)” - jarredou

To sum up:

Tips to enhance separation

Demudder in UVR/x-minus
(increases vocal residues)

AI mastering services

Blending with RVC model

AI audio upscalers list

Make your own remaster:

More clarity/better quality/general audio restoration of separated stem(s)

Have complete freedom over the result, using (among others) spectral restoration plugins to demudd the results of separations freely with plugins. Then you can use the result further with e.g. AI upscaler or in reverse.

E.g. from plugins, you can start by using Thimeo Stereo Tool which has a fantastic re/mastering chain feasible for spectral restoration useful for instrumentals sounding too filtered from vocals and lacking clarity. Also use Unchirp which states great complement to Thimeo Stereo Tool, although focuses more on the already existing spectrum.

You can also play with free Airwindows Energy/Energy2 and Air/Air2 (or Air3, MIA Thin) plugins for restoration, and furthermore some compressors or other plugins and effects mentioned in the link above.

If you're not afraid of learning a new DAW, Sound Forge Cleaning Lab 4 has great and easy built-in restoration plugins too (Brilliance, Sound Clone>Brighten Internet Sources) with complete mastering chain to push even further what you already got with Unchirp and Stereo Tool.

Izotope RX Editor and its Spectral Recovery may turn out to be just not enough, but the rest of RX plugins also available as VST can become handy, although Cleaning Lab has lots of substitutes for filtering various kinds of noise. Working comfortably in real-time with all the plugins opened simultaneously while combined is more comfortable than RX Editor workflow. But you can use some plugins from RX Editor as separate VSTs in other DAWs including Lab 4. Ozone Advanced might turn out useful too.

Actually, once you finish using the plugins above, now you can try out some of the mastering services and not in the opposite way (although you might want to meet some basic requirements of AI mastering services to get the best results first, e.g. in terms of volume).

Q: AI vocal remover did not "normalize" (I don't think it's the right word) the track on the moment where the vocal was removed, so it's noticeable, especially on instrument-heavy moments.

I make things better by creating a backup echo track by combining stereo tracks with inverted ones and adding this to the main track with -5db, but it's still not good enough. Are there any technics that separate track with not noticeable effects or maybe there is some good restoration algorithm that I can use

A: If vocals are cancelled by AI, such a moment stands out from the instrumental parts of the song.

Sometimes you can rearrange your track in a way that it will use instrumental parts of the song when there are no vocals, instead of leaving AI separated fragments. Sometimes it's not possible, because it will lack some fragments (then you can use only filtered moments at times), and even then, you will need to take care about coherence of the final result in the matter of sound as you said.

At times, even fade outs at the ends of tracks can have decent amounts of instrumentals which you can normalize and then use in rearrangement of the track. E.g. you normalize every snare or kick and everything later in fade out, and then till the end, so it will sound completely clean.

Generally it's all time-consuming, not always possible, and then you really have to be creative using normal mastering chain to fit filtered fragments to regular unfiltered fragments of the track.

You can also try out layering, e.g. specific snare found in a good quality in the track. May work easier for tracks made with quantization, so when the pattern of drums is consistent throughout the track. Also, you can use 4 stem Demucs ft or MDX23 and overlap drums from a fragment where you don’t hear vocals yet, so drums are still crispy there.

Ducking effect eliminator

You can also check Aufr33 Reaper 7.05 project aimed at alleviating this issue:

“the music volume is reduced where there are vocals”. Instruction:
“Just place two stems: vocals and music. Adjust the Expander if necessary.”

“It's just an expander side-chained to vocals. You can replicate this in any other DAW.”

Src | mirror

- Nice chart (>moved to “Advanced chain processing chart” at the bottom of Karaoke section (use search)

describing process for creating AI cover (replace kim vocal with voc ft there, or MDX23 vocals/UVR top ensemble/Roformers).

Blending with RVC model 
(by Gabox & dubpluris a.k.a. Mark | Avalaunch - text)

My use Case:

Restoring older, lower-quality vocal recordings (e.g., camcorder recordings from the 90s) using RVC models trained on clean studio vocals from the same artist.

Practical Workflow for Using RVC in Vocal Restoration

The idea is not to replace old performances entirely, but to enhance them. A few key points came out of the discussion:

- Blending, not replacing: Using only the RVC output will usually sound artificial. The better approach is to run the old vocal stems through the trained RVC model and then blend the AI-generated stem with the original. This preserves natural performance qualities while adding clarity.

- Input quality matters: Even “decent but rough” camcorder audio can work. Extremely degraded sources, however, will still produce artifacts (“bad input = bad output”).

Complementary tools:

FlashSR – an audio super-resolution method that restores high frequencies and improves fidelity before running RVC. (https://mvsep.com/en/demo?algorithm_id=60)
[AudioSR might potentially give better results, but it’s much slower;
“Imo much better candidates are: AP-BWE (Colab | new repo [old]) and Clearer-Voice-Studio's Clear Voice (my favorite is the 2nd one - codename0; more simplified version by codename0 - DL”]

Matchering – matches EQ/tonal balance of rough recordings to studio references, either standalone or integrated into UVR5. Using a clean studio version of the artist as the reference and the old performance as the target is recommended.

(https://sergree.github.io/matchering / online) - Available on UVR

[You can also try out https://masterknecht.klangknecht.com/]

General workflow:

   

1. (Optional) Pre-process low-quality audio with FlashSR.

2. Train RVC on clean studio stems.

3. Run inference on the old stems with the trained model (i.e., feed the cleaned original vocal through the trained RVC model to get a converted stem.)

4. Blend, align and mix original + RVC stem (RVC as enhancement, not replacement) until it feels natural.

5. Use Matchering or other mastering techniques to polish.

For a comprehensive remastering workflow, see How to make your own remaster.

The overall takeaway: RVC can be used for restoration, but it works best as part of a chain of tools (super-resolution, EQ matching, mastering), with the human performance always kept at the center through blending rather than full replacement.”

(old) More descriptions of models

and AIs, with troubleshooting and tips
(most models here are dated as it lacks Roformers)

(Instruction here was moved to Reading advice)

Older models descriptions

- Inst fullband (fb) HQ_3/4/5 x-minus, MVSEP, Colabs

HQ_4 vs 3 has some problems with fadeouts when occasionally it can leave some vocal residues

HQ_3 generally has problems with strings. mdx_extra from Demucs 3/4 had better result with strings here, sometimes 6s model can be good compensation in ensemble for these lost instruments, but HQ_3 gives some extra details compared to those.

HQ_3/4 are generally muddy models at times, but with not much of vocal residues (near Gsep at times, but more than BS-Roformer v2).

For more clarity, use MDX23C HQ model (HQ_2 can have less vocal residues at times).

Another possibly problematic instruments are those wind ones (flute, trumpet etc.)

- use Kim inst or inst 3 then

HQ3 has worse SDR vs:

- voc_ft, but given that HQ_3 is an instrumental model, the latter can leave less vocal residues at times.

https://mvsep.com/quality_checker/leaderboard2.php?id=4029

https://mvsep.com/quality_checker/leaderboard2.php?id=3710

These are SDR results from the same patch, so the voc_ft vs HQ_3 comparison is valid.

- MDX23C_D1581 (narrowband) - usually worse results than voc_ft and probably worse SDR if evaluation for both models was made on the same patch

Can be a bit better for instrumentals

“The new model is very promising

although having noise, it seems to pick vocals more accurately and the instrumentals don't have that much of the filtering effect (where entire frequencies are being muted).”

While others say it’s worse than demucs_ft

- GSEP AI an online closed source service (cannot be installed on your computer or your own site). mp3 only, 20kHz cutoff.

Decent results in some cases, click on the link above to read more about GSEP in the specific section below. This SDR leaderboard underestimates it very much, probably due to some kind of post-processing used in GSEP [probably noise gate and/or slight reverb or chunking). As a last resort, you can use 4-6 stems option and perform mixdown without vocal stem in e.g. Audacity or other DAW. 4-6 stem option has additional noise cancellation vs 2 stem.

GSEP is good with some tracks with a busy mix or acoustic songs where everything else simply fails, or you’re forced to use the RX10 De-bleed feature.

- GSEP is also better than MDX-UVR instrumental models on at least tracks with flute and possibly duduk/clarinet or oriental tracks, and possibly tracks with only piano, as it has a decent dedicated piano model.

- To address the issue with flute using MDX-UVR, use the following ensemble: Kim_Inst, HQ1, HQ2, INST 3, Max Spec/Max Spec (Anjok).

- Sometimes kim inst and inst3 models are less vulnerable to the issue (not in all cases).

- Also, main 406 vocal model keeps most of these trumpets/saxes or other similar instruments

- Passing through a Karaoke model may help a bit with this issue (Mateus Contini method).

- inst HQ_1 (450)/HQ_2 (498)/HQ_3 MDX-UVR fullband models in the Download center of UVR5 - great high quality models to use in most cases. The latter a bit better SDR, possibly a bit less vocal residues. Not so few like inst3 or kim ft other in specific cases, but a good point to start.

What you need to know about MDX-UVR models is that they're divided into instrumental and vocal models and that instrumental models will always leave some instrumental residues in vocals and vice versa - vocal models will more likely to leave some vocal residues in instrumentals. But you can still encounter specific cases of songs when breaking that rule will benefit you - that might depend on the specific song. Usually, instrumental model should give better instrumental if you’re fighting with vocal residues.

Also, MDX-UVR models can sometimes pick up sound midi effects which won’t be recovered.

- kim inst (a.k.a. ft other) - cutoff, cleaner results and better SDR than inst3/464 but tends to be more noisy than inst3 at times. Use:

- inst3/464 - to get more muddy, but less noisy results, although it all depends on a song, and sometimes HQ_1/2/3 models provide generally less vocal residues (or more detestable).

- MDX23 by ZFTurbo v1 - the third place in the newest MDX challenge. 4 stem. Already much better SDR than Demucs ft (4) model. More vocal residues than e.g. HQ_2 or Kim inst, but very clean results, if not the cleanest among all at the time. Jarredou in his fork fixed lots of those issues and further enhanced the SDR so it’s comparable with Ensemble on MVSEP, which was also further enhanced since the first version of the code released in 2023, and also has newer models and various enhancements.

- Demucs 4 (especially ft 4 stem model; UVR5, Colab, MVSEP, 6s available) - Demucs models don't have so aggressive noise cancellation and missing instruments issue like in GSEP. Check it out too in some cases (but it tend to have more vocal bleeding than GSEP and MDX-UVR inst3/464 and HQ_3 (not always, though), and 6 stem has more bleeding than 4 stem, but not so much like the old mdx_extra 4 stem model).

- Models ensemble in UVR5 GUI (one of the best results so far for both instrumentals and vocals SDR-wise). Decent Nvidia GPU required, or brace for 4 hours processing on 2/4 Sandy Bridge per whole ensemble of one song. How to set up ensemble video.

General video guide about UVR5.

"UVR-MDX still struggles with acoustic songs (with a lot of pianos, guitars, soft drums etc.)" so in this case use e.g. GSEP instead.

Description of vocal models by Erosunica

"That's my list of useful MDX-NET models (vocal primary), best to worst:

- MDX23C-8KFFT-InstVoc_HQ (Attenuates some non-verbal vocalizations: short low-level and/or high-frequency sounds)

- Kim Vocal 2

- UVR-MDX-NET-Voc_FT

- Kim Vocal 1

- Main (Attenuates some low level non-verbal vocalizations)

- Main_340 (Attenuates some non-verbal vocalizations)

- Main_406 (Attenuates some non-verbal vocalizations)

- Kim Inst (Attenuates some non-verbal vocalizations)

- Inst_HQ_3 (Attenuates some non-verbal vocalizations)

- MDXNET_2_9682 (Attenuates some non-verbal vocalizations)"

and it’s also worth to check HQ_4.

“UVR BVE v2 model [currently on x-minus] is actually full band. There is, however, a small nuance. This model uses MDX VocFT preprocessing, which is not full band. MDX VocFT model is rebalancing the song. The music is slightly mixed with the vocals (25% music + 100% vocals). This mix is then processed by the BVE model. A small amount of music can help the model better understand the context (it's important for harmony separation). We train the model on a rebalanced dataset. It contains 25% of music.” aufr33

_____

All the tips moved to Tips to enhance separation section

_____

Screenshot and video showcase

MDX settings & ens. explanations in UVR5
(and also Demucs/VR/MDX v2/23C inferencing parameters)

In one of the pre-5.6 UVR updates, the following min/avg/max features for single models got replaced by a better automated alternative, and you might still get cleaner results of e.g. voc_ft with max_mag on X-Minus or in this Colab still utilizing it (or downgrade your UVR version).

Now it’s only applicable for Ensemble and Manual Ensemble in Audio Tools.
Manual Ensemble is very fast, can be used on even old dual-core CPU, as it uses already separated files and simple code - not model.

Ensemble algorithm explanations

Ensemble - a way to use multiple models to potentially get better results.

Rules to be broken here, but:

Max Spec is generally for vocals
(is maximum result of each stem, e.g. in a vocal you'll get the heaviest weighted vocal from each model, and the same goes for instrumental, giving a bit cleaner results, but more artefacts)

Min Spec for instrumentals in most cases
(it leaves the similarity from the models)

Avg Spec is something in between
(gets the average of vocals/instrumentals)

E.g. following the above, we get the following setting:

“Max Spec / Min Spec”

Left side = about the Vocal stem/output

Right side  = about the Instrumental stem/output

"Max takes the highest values between each separation to create the new one (fuller sounding, more bleed).

Min takes the lowest values between each separation to create the new one (filtered sounding, less bleed).

Avg is the average of each separation."

More

For ensemble, avg/avg got the highest SDR, then worse results for respectively max/max, min/max and min/min.

For single MDX model, min spec was the safest for instrumental models and gave the most consistent results with less vocal residues than others.

Max spec - is the cleanest - but can leave some artifacts (if you don't have them in your file, then Max Spec for your instrumental like now might be a good solution).

Avg - the best of the both worlds and the only possible to test SDR e.g. at least for ensembles, maybe even to this day if it wasn't patched

“Max Spec/Min Spec” option

For at least a single instrumental model, it's the safest approach for instrumentals and universal for vocals. E.g. Min Mag/Spec in kae Colab using the old codebase for MDX models gives me the only acceptable results with hip-hop. I usually separate using a single model, but I cannot guarantee that Min Spec in UVR and manual ensemble will necessarily work exactly like Min Mag in Colab for a single model. But the explanation remains the same. The best option might even depend on a song.

TL;DR

For vocals bleeding in instrumentals

You can use Spectral Inversion for alleviating problems with bleeding in instrumentals.

Max Spec/Min Spec is also useful in such scenario.

You want less bleed of Vocal in Instrumental stem?

Use Max-Min

For bleeding instruments in vocals

Phase Inversion enabled helps to get rid of transients of the kick which might be still hearable in vocals in some cases.

Set Ensemble Algorithm: Min/Avg when you still hear bleeding.

If still the same, try Min/Max instead of Avg/Avg when doing an ensemble with Vocals/Instrumental output.

Also, you can resign from ensemble setting, and simply use only one clean model on the models list if the result is still not satisfactory.

Further explanations


Why not always go for Min-Max when you want the best acapella?

Why not always go for Max-Min when you want the best Instrumental?

So far, I hear Max-Min on Instrumental sounds more 'muddy/muffled' compared to Avg-Avg.

I bet this will be the same for acapella, but it's less noticeable (I don't hear it).

Hence, I think the best approach would be always going with Avg-Avg.

Then based on the outcome - after reviewing, tweak it based on your desired outcome,

and process again with either Min-Max or Max-Min.”

Min = less bleeding of the other side/stem (into this side/stem), but could get sound muddy/muffled

Max = more full sound, but potential it will have more bleeding

Avg = average, so a bit of all models combined

Average/Average is currently the best for ensemble (the best SDR - compared with Min/Max, Max/Min, Max/Max).

“Ensemble is not the same as chopping/cutting off and stitching, it blends/removes frequencies. If song 1 has high vocals in the chorus, and song 2 has deep vocals in the chorus, max will mash them together, so the final song will have both high and deep vocals

while min will remove both vocals”

"If I ensembled with max, it would add a lot of noise and hiss, if I ensemble with min it would make the overall sound muted gsep."

Technical explanation on min/avg/max

Max - keeps the frequencies that are the same and adds the different ones

“Max spec tends to give more artifacts as it's always selecting the loudest spectrogram frequency bins in each stft frames. So if one of the inputs have artifacts when it should be silent, and even if all other inputs are silent at the same instant, max spec will select the artifacts, as it's the max loud part of spectrogram here.” jarredou

Min - keeps the frequencies that are the same and removes any different ones

"if the phases of the frequencies are not similar enough min spec and max spec algorithms for ensembles will create noisy artifacts (IDK how to explain them, it just kinda sounds washy), so it's often safer to go with average"

by Vinctekan

"Min = Detects the common frequencies between outputs, and deletes the different ones, keeps the same ones.

Max = Detects the common frequencies between outputs, and adds the difference to them.

Now you would think that Max-Spec would be perfect since it should combine the all of the strengths of every model, therefore it's probably the best option

That would be the case if it wasn't for the fact that the algorithms that are used are not perfect, and I posted multiples tests to confirm this.

However, it still gives probably the cleanest results, however, there are a few issues with said Max_Spec:

1. Lot of instrumentals are going to be left within the output

2. If you are looking to measure quality by SDR, don't expect it to be better than avg/avg

The average algorithm, basically, combine all the outputs and averages them. Like the average function in Excel.

The reason why it works best is that it does not destroy the sound of any of the present outputs compared to Max_Spec and Min_Spec

The 2 algorithms still have potential for testing, though."

More on how the ensemble in UVR works

"Max takes the highest values between each separation to create the new one (fuller sounding, more bleed).

Min takes the lowest values between each separation to create the new one (filtered sounding, less bleed).

Avg is the average of each separation."

“[E.g.] HQ 1 would be better if the ensemble algorithm worked how I thought it did.

It was explained to me that [ensemble algorithm] tries to find common frequencies across all the outputs and combines them into the result, which to me doesn't actually seem to happen when HQ1 manages to bring vocals to the mix in an 8 model ensemble, how is it not like "okay A those are vocals, and B you're the only model bringing those frequencies to me trying to imply that they are not vocals" and discard them. I mean I am running max/max, but I swear all avg/avg and min/min do is lower the volumes [see enemble in DAW], It's hard to know without days of testing”

“If u try avg/avg it will get quite muddy on instr result than max/max. But some song if you put kim vocal 1 will get vocal residue on the result”.

4-5 max ensemble models rule
Q: Why I shouldn’t use more than 4-5 models for UVR ensemble (in most cases)

A: It's easier to get, when you separate the same song using some models. Get the best 4-5 models out of the most recommended currently, plus make some more separations, using some random ones. Then try to reflect avg spec from UVR by importing all of these results to your DAW.

You'll do it by decreasing volume by 3dB per one stem, so for a pair you need to decrease the volume of two stems by 6dB (possibly 6.02 as well). Decrease the volume by the same value further for more than a pair for all stems accordingly, so you'll get pretty much similar result like avg spec in UVR.

You can also maybe apply a limiter on the master. In the second variant, manipulate the volume of all stems by your taste instead of keeping the same volume. By this process, you can observe that the more results imported above 4-5 results, the worse result you have when you don't decrease volume of worse results. When you have control over the volume of single results, you'll end up decreasing the volume of bad results (or deleting them completely). You don't have this opportunity in UVR using avg spec - so like in the first variant in your DAW when you set the same volume for all results. The only way to not deteriorate the final result further, is to delete such worse results from the bag entirely, to not worsen the final outcome when you have too many models ensembled. Without the possibility of decreasing volume of such a result when all volumes are equal, the more results you'll import to the bag of the 4-5 the best models, the worse final result you'll get. Because you cannot compensate for bad results in the bag by decreasing their volume like in avg spec - all tracks are equally loud in the bag of avg to the models with good results - hence, good models sound quieter if they are in minority and the final outcome is worse.

The 4-5 max models ensemble rule is taken from long-conducted tests of SDR on MVSEP multisong leaderboard. When various ensembles were tested in UVR, most of these combinations didn't consist of more than 4-5 models, because above that, SDR was usually dropping. Usually due to all the reasons I mentioned.

Even using clever methods of using only certain frequencies of specific models, like in ZFTurbo, jarredou and Captain FLAM code from MDX23 (don't confuse with MDX23C arch) and its derivations, which minimize the influence of "diminishing returns" when using too many models I think they never used more than 4-5 in their bags, and they conducted impressive amount of testing, and jarredou even focused on SDR during developing his fork (actually OG ZFTurbo code too).

_____


For vocal popping in instrumental issue, read about chunks or update UVR to use a better option used automatically (called batch mode) if you didn't update to 5.6/+ for a long time already, but the issue might still occur on GPUs with less than 11GB VRAM (and earlier patches doesn’t have Roformers support).

_______

MDX v2 parameters (e.g. HQ_1-5, Kim inst, Inst 1-3, NET, Crowd)
(self.n_fft /  dim_f / dim_t inference parameters later below)

Segments 512 had better SDR than many higher values on various occasions (while 256 has lower SDR, and has almost the same separation time).

Segments 1024 and 0.5 overlap are the last options before processing time increases very much.

Don't exceed an overlap of 0.93 for MDX models, it's getting tremendously long with not much of a difference.

Overlap 0.7-0.8 might be a good choice as well.

Segments can also ditch the performance AF -

segments 2560 and 2752 (for 6GB VRAM) might be still a high, but balanced value, although not fully justified SDR-wise, as 512 or 640 can be better than higher values for many songs.

In UVR and Not Eddy’s Colabs you can change segment size from 512 to 32 in order to possibly get better results with some older models like e.g. 438 (it tremendously increases separation time).

Overlap: 0.93-0.95 (0.7-0.8 seems to be the best compromise for ensembles, with the biggest measured SDR for 0.99)

Best measured SDR on MVSEP leaderboard have currently following settings (but it was measured on 1-minute songs, so it can be potentially different for your song):

Segment Size: 4096

Overlap: 0.99

with 512/0.95 worse by a hair (0.001 SDR) and 0.9 overlap for as long, but still not tremendously long processing time (1h30m31s vs 0h46m22s for multisong dataset on GTX 1080 Ti).

Also, segments 12K performed worse than 4K SDR-wise (counterintuitively to what it is said, that higher means better result, but maybe diminishing returns at some point here, so too big values maybe cause SDR drop in some cases)

It seemed to be correlated with set overlap.

For overlap 0.75, segments 512 was better than 1024,

but for overlap 0.5, 1024 was better, but the best SDR out of these four results has 0.75/512 setting, although it’s a bit slower than 1024, but for 0.99 overlap, 4096 segments were better than 512.

SDR difference between overlap 0.95 and 0.99 for voc_ft in UVR is 0.02.

Segment size 4096 with overlap 0.99 (here) vs 512/0.95 (here) showed only 0.001 SDR difference for voc_ft and vocals in favour of the first result.

Difference between segment size 512 with overlap 0.25 (here) vs 0.95 (here) is 0,1231 SDR for the latter.

The difference between default segment size 256 with overlap 0.25 (here) vs 512/0.95 (here) is 0,1948 SDR for vocals, and 0,1969 with denoiser on (standard, not model), and 0.95 is longer by triple.

1024/0.25 vs 256 has not much longer processing time (7 vs 6 mins) than default settings, and better SDR by 0.0865

For overlap 0.75, segments 512 were better than 1024 (at least on 1 minute audio).

Measurement is logarithmic, meaning that 1 SDR is 10x difference.

Be aware that increasing only overlap to e.g. 0.5 from default 0.25, when segments are still at default 256 will muddy the result a bit (might be more noticeable with denoise model enabled), while increasing segments (at least up to 480/512) suppose to add more clarity.

At least on the second beta Roformer patch, max supported segment size on 4GB AMD/Intel GPUs is 480 (at least for 4:58 and HQ 1-3 can sometimes only work with lower 448 - higher overlap and segment size crashes).

256/0.5 also works, at least with HQ 4 (but crashes with 480 segments)

480/0.38 works too, but you can settle on e.g. 0.31 if it’s too muddy.

Try not to keep too many opened apps during separation, as drawing their interface also eats up VRAM on the GPU.

MDX-Net v2 max balanced settings:

Segment Size: 2752 (1024 if it’s taking too long as it’s the last value before processing time increases really much; at least SDR-wise, 512 is better in every case than default 256 unless overlap is increased, and still gets good SDR results)

Overlap: 0.7-/0.8

Denoising

Denoise option used to increase SDR for MDX-Net v2, but instrumentals get a bit muddier (result).

Denoise model has slightly lower SDR (result).

For MDX23C models it somehow changed and using standard denoiser doesn’t change SDR.

Spectral Inversion

On bigger dataset like Multisong Leaderboard decreases SDR, but sometimes you can avoid some e.g. instrumental residues using it - can be helpful when you hear instruments in silent parts of vocals.

Explanation:

"When you turn on spectral inversion, the SDR algorithm is forced to invert the spectrum of the signal. This can cause the SDR to lose signal strength, because the inverse of a spectrum is not always a valid signal. The amount of signal loss depends on the quality of the signal and the algorithm used for spectral inversion.

In some cases, spectral inversion can actually improve the signal strength of the SDR. This is because the inverse of a spectrum can sometimes be a more accurate representation of the original signal than the original signal itself. However, this is not always the case, and it is important to experiment with different settings to find the best results.

Here are some tips for improving the signal strength of the SDR when using spectral inversion:

* Use a high-quality input. The better the quality of the signal, the less likely it is that the SDR will lose signal strength when the spectrum is inverted. (...)"

Further, there is also about picking a good inversion algorithm and experimenting with different ones, but UVR seems to have one to pick anyway.

Q: I noticed https://mvsep.com/quality_checker/leaderboard2.php?id=2967

has Spectral Inversion off for MDX but on for Demucs. The Spectral Inversion toggle seems to apply to both models, so should it be on or off?

A: Good catch.

Once u put it on for one or the other, both will be affected indeed.

I've enabled it (so for both, actually) [for this result].

_____

MDX v3 parameters (e.g. MDX23C-InstVoc HQ and 2 and MDX23C_D1581)

(biggest measured SDR)

Segment Size: 512

Overlap: 16

(default)

Segment Size: 256

Overlap: 8

The “512/16 is slightly better for big cost of time” vs the default 256/8.

- On a GPU with lots of VRAM (e.g. 24GB), you can run two instances of UVR, so the processing will be faster. You only need to use 4096 segmentation instead of 8192.


It might be not fully correct to evaluate segment and overlap SDR-wise based on measurements done on multisong dataset, as every single file in the dataset is shorter than average normal track, and that might potentially lead to creating more segments and different overlaps than with normal tracks, so achieved results won’t fully reflect normal separation use cases (if e.g. number of segments is dependent on input file). Potentially, the problem could be solved by increasing overlap and segments for a full length song to achieve the same SDR as with its fragment from multisong dataset.

Recommended balanced values for various archs

between quality and time for 6GB graphic cards:

_____

VR

Window Size: 320 (best measured SDR)

Faster value for slow PCs: 512

Slower, might give more artefacts: 272

Worse: 768, 1024

Read more in VR settings

_____

Demucs

Segment: Default

Shifts: 2 (def)

Overlap: 0.5

(experimental: 0.75,

default: 0.25)

The best SDR for the least time for Demucs (more a compromise, as it takes much longer than default settings ofc - “best SDR is a hair more SDR and a sh*load of more time):

Segments: Default

Shifts: 0

Overlap: 0.99 (max can be 0.999 or even more, but it’s getting tremendously long)

Best results for instrumentals as input (tested in Colab):

Segments: Default

Shifts: 10 (20 is max possible)

Overlap: 0.1

"Overlap can reduce/remove artifacts at audio chunks/segments boundaries, and improve a little bit the results the same way the shift trick works (merging multiple passes with slightly different results, each with good and bad).

But it can't fix the model flaws or change its characteristics"

In case of Voc_FT it's more nuanced... there it seems to make a substantial difference SDR-wise.

The question is: how long do you wanna wait vs. quality (SDR-based quality, tho)”

In UVR and Not Eddy’s Colabs you can change segment size from 512 to 32 in order to possibly get better results with some older models like e.g. D1581 (but it tremendously increases separation time).

For lack of spectrum above 14.7kHz

E.g. in such ensemble:

5_HP-Karaoke-UVR, 6_HP-Karaoke-UVR, UVR-MDX-NET Karaoke, UVR-MDX-NET Karaoke 2

Set Max Spec/Max Spec instead of Min Spec/Min Spec, and also hi-end process (both need to be enabled for fuller spectrum).

Karaoke models are not full band, even VR ones are 17.7kHz and MDX are 14.7kHz IRC. Setting Max Spec with hi-end process will give around 21kHz output in this case.

Cutoff with min spec in narrowband models is a feature introduced at some point in UVR5 GUI for even single MDX models in general, and doesn't exist in CLI version. It's to filter out some noise in e.g. instrumental from inversion. Cutoff then matches model training frequency (in CLI MDX, vocal model after inversion with mixture gives full band instrumental). Also, similar filtering/cutoff is done in ensemble with min spec.

More settings explanation

Leaving both shifts and overlap default vs shifts 10 decreases SDR by only 0.01 SDR in ensemble, but processing time is much faster - 1.7x for each shift. Also, 0.75 overlap increases SDR at least for a single model when even shift is set to 1)

 

It takes around 1 hour 36 minutes on a GTX 1080 Ti for 100 1-minute files.

“And 18 hours on i5-2410M @2.8 for 5:04 track.

Rating 1 Ensemble on a 7-min song to compare.

Time elapsed:

1080Ti = 5m45s = 345s = 100%

4070Ti = 4m49s = 289s = 83,8%

4070Ti = ~16% faster

1080Ti = ~€250 (2nd hand)

4070Ti = €909 (new)

Conclusion: for every 1% gain in performance, u pay €41 extra (€659 extra in total).” Bas

More min/max explanations moved to MDX/Ensemble settings

Compensation values for MDX v2
(no longer necessary since MDX23C)

“Volume compensation compensates the audio of the primary stems to allow for a better secondary stem.''

For the last Kim's ft other instrumental model, 1.03 or auto seems to do the best job.

For Kim vocal 1 and NET-X (and probably other vocal models), 1.035 was the best, while 1.05 was once calculated to be the best for inst 3/464 model, but the values might slightly differ in the same branch (and compensation value in UVR5 only changes secondary stem - changing compensation value in at least UVR GUI for inst models doesn't change SDR of instruments metric)

self.n_fft /  dim_f / dim_t parameters

These parameters directly correspond with how models were trained. In most cases they shouldn't be changed, and automatic parameter detection should be enabled.

- Fullband models:

self.n_fft = 6144 dim_f = 3072 dim_t = 8

- kim vocal 1/2, kim ft other (inst), inst 1-3 (415-464), 406, 427:

self.n_fft = 7680 dim_f = 3072 dim_t = 8

- 496, Karaoke, 9.X (NET-X)

self.n_fft = 6144 dim_f = 2048 dim_t = 8 (and 9 kuielab_a_vocals only)

- Karaoke 2

self.n_fft = 5120 dim_f = 2048 dim_t = 8

- De-reverb by FoxyJoy

self.n_fft = 7680 dim_f = 3072 dim_t = 9

___

Roformers (located in MDX-Net menu;

only in UVR Roformer beta patches)

chunk_size

“most of the time using higher chunk_size than the one used during training gives a bit better SDR score, until a peak value, and then quality degrades.

For Roformers trained with 8 sec chunk_size, 11 sec is giving best SDR (then it degrades with higher chunk size)

For MDX23C, when trained with ~6 sec chunks, iirc, peak SDR value was around 24 sec chunks (I think it was same for vit_large, you could make chunks 4 times longer)

How much chunk_size can be extended during inference seems to be arch dependant.” - jarredou

Be aware that increasing chunk_size consumes much more VRAM, and for 4GB VRAM AMD/Intel GPUs, the max supported will be chunk_size = 112455 (2,55s), sometimes chunk_size = 132300 (3s). CUDA has garbage collector which might make VRAM usage more efficient.

“Conversion between dim_t and chunk_size [dim_t was used in the old Roformer beta 2 UVR patch]

dim_t = 801 is chunk_size = 352800 (8.00s) - maximum value working on AMD/Intel 8GB GPUs and 900MB, at least Mel models

dim_t = 1101 is chunk_size = 485100 (11.00s)

dim_t = 256 is chunk_size = 112455 (2,55s) - maximum value for AMD/Intel 4GB GPUs

dim_t = 1333 is chunk_size = 587412 (13,32s)

The formula is: chunk_size = (dim_t - 1) * hop_length)” - jarredou

Unless you turn off Segment default in Options>Advanced MDX-Net>Multi Network Options, chunk_size is being read from the yaml of the model.

Inference mode

Can be found in the menu Multi Network Options menu above. Turning it off will fix the issue of silent separations on older GTX GPUs (iirc GTX 900 and older), but it might make separation slower for other, at least Nvidia GPUs.

It was implemented in one of the latest beta Roformer patches, so if you noticed any slowdowns since updating UVR, try enabling it (now it’s disabled by default).

batch_size

Inference Colab by jarredou forces 1 (clicks with that setting were fixed in MSST later), and using above 2 might increase VRAM usage. In newer patches, Anjok started to use MSST inference code for Roformers and MDX23C, hence it might have inherited its usage.

Technical explanation of how it works near the end of this document (scroll down a bit).

Overlap

4 is a balanced value in terms of speed/SDR according to measurements (since the beta patch #3 or later used above, overlap 16 is now the slowest (not overlap 2 anymore) and overlap 4 has a bigger SDR than overlap 2 now.
Some people still prefer using overlap 8, while for others it’s already an overkill.
There’s very little SDR improvement for overlap 32, and for 50 there’s even a decrease to the level of overlap 4, and 999 was giving inferior results to overlap 16.

Compared to overlap 2, for 8 “I noticed a bit more consistency on 8 compared to 2 (less cut parts in the spectrogram).”
Instrumentals with overlap higher than 2 can get gradually muddier.


So the highest SDR gives overlap 32, but for drastic separation time increase vs 16.
Calculations above were based on evaluations conducted on multisong dataset on MVSEP. Search for e.g. overlap 32 and overlap 16 below, and you will see the results to compare:

https://mvsep.com/quality_checker/multisong_leaderboard?algo_name_filter=kim

“overlap=1 means that the chunk will not overlap at all, so no crossfades are possible between them to alleviate the click at edges.”
The setting in GUI overrides the one in the model's yaml.

Check out also MVSep/ZFTurbo analysis on overlap here.

Refer to UVR Roformer beta patch section for more detailed information

>>Tips to enhance separation results<<

If you cannot achieve good separation, you can conduct the following experiments.
Some tricks might be outdated when using Roformers.

1. De-bass

Turn down all the bass to stabilize the voice frequencies of your input song (example EQ curves: 1 and 2).

Male setting: cut all below 100Hz + cut all above 8kHz.

Female setting: cut all below 350Hz + cut all above 17kHz.

This works, because jitter is reduced a lot.

2. De-reverb

You can also test out the de-reverb e.g. in RX Advanced 8-10 on your input song. One or both combined in some cases may help you get rid of some synth leftovers in vocals. Alternatively (not tested for this purpose), you can also try out this or this (dl is in UVR's Download Center) de-reverb model (decent results). Currently, the VR dereverb/de-echo model in UVR5 GUI seems to give the best results out of the available models (but RX or others described in the models list section at the top can be more aggressive and effective with more customizable settings).

3. Unmix drums 

(mainly tested on instrumentals)

Separate an input song using 4 stem model, then mix the result tracks together without drums and separate the result using strong ensemble or single vocal or instrumental model (doesn't always give better results).

Alternatively, unmix bass as well. There’s great bass+drums BS-Roformer model released for UVR (currently in beta)

4. Pitch it down/up

(soprano/tenor voice trick + ensemble of both)

- You can use https://github.com/JoeAllTrades/SpectraDownshift for it
(it’s based on scipy, it’s lossless, so fully reversible and nulls), or:

- Already implemented option in newer versions of UVR under “Shift Conversion Pitch” in Settings>Choose Advanced Menu>Advanced [Arch] Options>

And there are positive and negative values when you scroll up and down
(lossy, even more than soxr).

Negative value will slow down the track before separation, so e.g. model with cut-off will be compensated for its band lost a bit after speeding up again.

If you slow down the input file, it may allow you to separate more elements in the “other” stem of 4-6 stems separations of Demucs or GSEP (when it’s done manually).

It works either when you need an improvement in such instruments like snaps, human claps, etc. The soprano feature on x-minus works similarly (or even the same), it’s also good for high-pitched vocals.

Be aware that low deep male vocals might not get separated while using this method (then use tenor voice trick instead - so pitch it up instead of pitching it down.

Also, it serves the best for hard paned songs (e.g. 1970 and pre era, e.g. The Beatles, etc). Also, it works great for drums. While evaluation on multisong dataset on MVSEP, it decreases SDR by around 1.

"Basically lossless speed conversion a.k.a. soprano voice trick done manually:

Do it in Audacity by changing sample rate of a track, and track only (track > rate), it won't resample, so there won't be any loss of quality, just remember to calculate your numbers

44100 > 33075 > 58800

48000 > 36000 > 64000

(both would result in x0.75 speed)

etc." (by BubbleG)

Q: Won't the result be sped up?

A: “No. Because when you first slow it down, after processing with said model it gets converted to 44100 again (only the sample rate, not the actual speed), so speeding it up brings the speed back to normal” becruily

Q: I don't quite get what I'm supposed to do though, just slow down the file to 0.75x and then export in 58800?

A: “change the sample rate to 33075 Hz,

then export at whatever sample rate

process then,

change the sample rate of the processed file to 58800 Hz

key word being change, not resample

like this, click other and the pick the correct samplerate” Dry Paint Dealer

4b*. If you have a mix of soprano and baritone voices, you possibly can do:

"1. Soprano mode (slow down sample rate), then bring back to normal

after that

2. Tenor mode (speed up sample rate), then bring back to normal

and finally combine the two with max algorithm"

Making an ensemble of such results can also increase the quality of separation.

“I have tried faking the sample rate both in 48k and 96k files. For 48k usually the results are satisfactory, because most models already are trained on lower pitches from augmentation.

But from 96k to 44.1k is such a long road, it simply doesn't work” - Michael

5. Use 2 stem model result as input for better 4-6 stem separation 

You may get better results in Demucs/GSEP/MDX23C Colab using previously separated good instrumental result from UVR5 or elsewhere (e.g. MDX HQ3 fullband or Kim inst narrowband in case of vocal residues, or BS-Roformer 1296)

6. Debleed

If you did your best, but you still get some bleeding here and there in instrumentals, check RX 10 Editor with its new De-bleed feature. Showcase

More methods of debleeding stems.

7. Vocal model>karaoke model

You might want to separate the vocal result achieved with a vocal model with MDX B Karaoke afterwards to get different vocals (old model).

8. The same goes for unsatisfactory result of instrumental model - you can use MDX-UVR Karaoke 2 model to clean up the result, or top ensemble or GSEP like for cleaning inverts (old models)

9. Mixdown of 4 stems with vocal volume decreased for final separation 

An old trick of mine. Used in times of Spleeter to minimize vocal residues.

Process mixture to 4 stems and then mix stems in a way that vocal is still there, but quieter, so lower their volume, and set drums louder, then send the mixture from it to one good isolation model/ensemble, so in result drums after separation will be less muddy, and possible vocal residues will be less persistent.

But it was in times when there wasn't even Demucs (4) ft or MDX-UVR instrumental models, where such issues are much less prevalent.

10. If you use UVR5 GUI and 4GB, you may hear more vocal residues using GPU processing than e.g. while using 11GB GPU (tested on NVIDIA). In this case, use CPU processing instead.

11. Fake stereo trick

Aufr33: “process the left channel, then the right channel, then combine the two. [Hence] the backing vocals in the verses are removed” (it still may be poor, but better). “I'm having to process as L / R mono files otherwise I get about 3-5% bleed into each channel from the other channel, but processing individually, totally fixes that” -A5

On an example of Audacity: import your file, click on down arrow in track selection near its label, click Split Stereo Track, go to Tracks>Add New>Stereo Track.

Mark the whole channel, copy and paste on one of the tracks you divided before.

It will overlap the same mono track in stereo track, so the same across both channels.

Do the same for both L and R separately. Then separate with some model both results separately. Then import both files and join their separate channels by method above. Don’t confuse L and R channel while joining both.

12. Turn on Spectral Inversion in UVR 

it can be helpful when you hear instruments in silent parts of vocals, and sometimes also using denoiser might help for it (although both can make your results slightly muddier)

13. Chain separation

For vocal residues in instrumental, you can experimentally separate it with e.g. Kim vocal (or inst 3) model first and then with instrumental model. You might want to perform additional steps to clean up the vocal from instrumental residues first, and invert it manually to get cleaner instrumental to separate with instrumental model to get rid of vocal residues. Tutorial 

14. To not clean silences from instrumental residues in the vocal stem manually, you can use a noise gate in even Audacity. Video

In some cases, using noise reduction tool and picking noise profile might be necessary. Video

15. Choice of good models for ensemble

Use only instrumental models for ensemble if you have some vocal residues (and possibly vice versa - use only vocal models for ensemble for vocals to get less instrumental residues) - mainly used in times when there was still strong division between vocal and instrumental models (before MDX23C release). Now it can narrow down to picking only models which doesn’t have bleeding - listening all the separate models results carefully, and pick the best 2-5 results to make an ensemble.

16. For vocals with vocoder

You can use 5HP Karaoke (e.g. with aggression settings raised up) or Karaoke 2 model (UVR5 or Colabs). Try out separating the result as well (outdated models).

"If you have a track with 3 different vocal layers at different parts, it's better to only isolate the parts with 'two voices at once' so to speak"

Be aware that BS-Roformer model ver. 2024.04 on MVSEP is better on vocoder than the viperx’ model.

17. Find some leaked or official instrumental for inversion

 

To get better vocals

If you're struggling hard getting some of the vocals:

"I used an instrumental that I don't remember where I found it (I'm assuming most likely somewhere on YouTube) and inverted it and then used MDX (KAR v2) on x-minus and then RX 10 after.

I Just tried the one-off Bandcamp and funnily enough it didn't work with an invert as good as the remake that I used from YouTube, but I don't remember which remake it was I downloaded because it was a while ago"

18. Fix for ~"ah ha hah ah" vocal residues

Try out some L/R inverting, try out to separate multiple times to get rid of some vocal pop-ins like this

19. Center channel extraction method 

by BubbleG using Adobe Audition:

(may not work on Roformers anymore, see after point 72)

"The idea is that you shift the track just enough where for example if you have a hip hop track, and the same instrumental tracks the drums will overlap again in rhythm, but they will be shifted in time so basically Center Extract will extract similar sounds. You can use that similarity to further invert/clean tracks... It works on tracks where samples are not necessarily the same, too…”

>

Step-by-step guide by Vinctekan (video)

1. You take your desired audio file

2. Open it in Audacity

3. Split Stereo to Mono

4. Click the left speaker channel (now mono), and duplicate it with Ctrl+D.

*: If the original and duplicate is not beside eachother, move it so that it's next to eachother

5: Select the original left speaker channel and it's duplicate, and click "Make Stereo Track"

6: Solo it.

7. Export it in Audacity, preferably in 44100hz since UVR doesn't output in higher frequencies. Format, and bit depth don't really matter, I prefer wav always.

8: Do the same thing for the right speaker channel.

9: Open UVR

10: Navigate to Audio Tools>Manual Ensemble.

11: Make sure to choose Min Spec (since that function is supposed to isolate the common frequencies of 2 outputs)

12: Select the 2 exported fake stereo files of both the left and right speaker channels.

13: Hit process

___

20. Q&A for the above

Q: For the right channel are you doing the same with the duplicate and moving the file next to the original or just duplicating and making that stereo?

A: Those 2 steps go hand in hand. These reason I mentioned it is because if you try to make a Stereo Track with those 2 (the left/right channel speaker, and it's duplicate mono]) when there is a track between them, it doesn't work. Even if you select those 2 with Ctrl held down.

Take that 1 channel (left/right), Ctrl+C, Ctrl+V, now you have 2 of the exact same audio. Hold Ctrl select the 2, click "Make Stereo Track". Finally, export.

_____

21. Passing through lot of models one by one

"I usually do ensemble to make an instrumental first, then demucs 4_ft… sometimes I do it once, then take that rendered file and pass it back through the algo a few more times, depends until it strips out artifacts."

It can be beneficial also in case of more vocal residues of MDX23 or Demucs ft model compared to current MDX models or their ensembles.

22. If you still have instrumental bleeding in vocals using voc_ft, process the result further with Kim vocal 2

23. Rearrange cleaner parts

When a verse starts, and you start having muddy drums and their pattern is consistent (e.g. some hip-hop), and you have cleaner drums from fragments before the verse starts, you can rearrange drums manually, using 4 stems model and paste that cleaner fragments throughout the track. Sometimes fade outs or intros can have clean loops without vocals, which can be rearranged without even the need of separation. Listen carefully to the track. Such moments can be even briefly in the middle of the song.

24. arigato78 method for lead vocal acapella

1) Try to make the best acapella (using mvsep.com site or using UVR GUI). I recommend the MDXB Voc FT model for this with an overlap setting set to at least 0.80 (I used 0.95 for this example). The overlap for this model at mvsep.com is set to 0.80. Speaking of the "segment size" parameter in UVR GUI - changing it from 320 to 1024 doesn't make much of a difference. It acts randomly, but we're working on a beta version of UVR GUI - remember that. (...)

I noticed all the "vocal-alike" instruments still remaining on the acapella track, but wait...

2) The second part is to process the acapella thru the mdx karaoke model (I did it using mvsep.com). I prefer the file with "vocalsaggr" in the name. It has more details than the file with "vocals" in it. The same goes to the background vocals in this case - I prefer the "instrumentalaggr" one.

One important thing - all (maybe almost) of the residue instrumental sounds were taken by mdx karaoke model to the backing vocals stem, leaving the lead vocal almost studio quality ("studio"). But - it may be helpful for all you guys trying to make good acapellas. I was just playing with all the models and parameters and I accidentally came across this. Please, let me know what you think about it. I'm gonna try this on some tracks with flutes, etc. And I realize that this method is not perfect - we get nice lead vocals, but the backing vocals are left with all that sh*tty residues.

So the track is called "Reward" by Polish singer Basia Trzetrzelewska from her 1989 album "London, Warsaw, New York".

__

25. Uneven quality of separated vocals

You can downmix your separated vocal result to mono and repeat the separation (works for e.g. BVE model on x-minus).

26. Experimental vocal debleed with AI for voice

Sometimes for instrumental residues in vocals, AIs for voice recorded with home microphone can be used (e.g. Goyo [now paid Supertone Clear], or even Krisp, RTX Voice, AMD Noise Suppression, Adobe Podcast as a last resort) it all depends on the type of vocals and how destructive the AI can get.

27. Minimize vocal residues for very loud songs

For very loud tracks between -2.5 and -4 iLUFS, try to decrease volume of your track before separation. E.g. for Ripple, -3dB for loud tracks is a good choice. If your track you’re trying to separate is already quiet and around -3dB, then the step is not necessary.

27b. You could try out attenuate volume of the mixture before separation (-3/6 dB), but I can't remember whether current MSST uses normalization before anyway. UVR maybe not.

28. Brief (old) models summary

MDX-Net HQ_3 or 4 is a more aggressive model for instrumentals, with usually fewer amounts of residues vs MDX23C HQ models or sometimes even vs KaraFan or jarredou’s MDX23 Colab v2.3. HQ_3 can give muddier results vs competition, though.

The most aggressive are BS-Roformer models, but they can sound filtered and even muddier at times, but cleaner. It’s good to use them with ensemble with e.g. MDX23C model.

voc_ft is pretty universal for vocals (with residues in instrumental, but not less muddy results), while people also liked Ripple/Capcut, although they give more artefacts (use the released BS-Roformer models now for vocals instead). Consider using MDX23C HQ model(s) as well, but they tend to have more instrumental residues.

29. Cleaning up bleeding between mics in multitracks

(by SeniorPositive)

"Demucs bleed "pro" tip that I figured out now, and I didn't see mentioned, that I will probably try to use every time I hear some bleed between. (...) I was cleaning multitrack from bleed between microphones in conga track, and used demucs for separation drums/rest pair, and [the] other [stem] had some of those bongos still, very very low, but it existed, and I heard it just enough.

- So I took rest signal, boosted it +20db (NOT NORMALISE! Other value but make note how much of it you boosted, go few dbs less to 0db threshold). If you do not boost it to sensible levels, the algorithm will skip it.

- Do separation once again (this time I've done it using spectralayers one, but it's also demucs)

- lower result -20dB add this result to first separation result

[The] result [is -] better separation, fewer data in other/bleed and with proper proportions.

It looks like AI is not yet perfect with low volume information and, as seen in ripple Bas Curtiz discovery, too hot content also."

Showcase

30. For clap leftovers in vocal stem

Methods suggested in debleeding

31. (paraphrase of point 17)

Use the traditional phase inversion method and then feed them to the UVR models if you have a chance of finding any official instrumental or vocal, but it doesn’t invert perfectly. This way, the models will have less noisy data to work with. But it sometimes happens that the official instrumental and the vocal version of tracks have slightly different phasing. This makes isolating vocals via phase inversion difficult, or even sometimes impossible ~Ryan_TTC

Sometimes only specific fragments of song will align, and in further parts of the track it will stop and require manual aligning. You may try to use Utagoe or possibly UVR with Aligning in Audio Tools as it shares some similar functionalities.

Why official stems don’t invert?

“Very rarely will the vocal or instrumental fully invert out of the master. This is because of master bus processing and non-linear nature of that processing. I.e. part of the masters sound is the processing reacting to the vocal and instrumental passing through the same chain.

Sidechaining and many limiters are also looking ahead to the signal. Also, some processing is non-linear so even if you set it up identically re. settings, each bounce will be slightly different in nature. Stuff like saturation/distortion. Some reverbs, limiters and transient shapers etc are not outputting the same signal / samples every time you bounce, so instrumental bounce is not the same as the master bounce in terms of phase inversion.” - Sam Hocking

32a. Muddiness in instrumentals of some BS-Roformer models

Invert (at best lossless) mixture (original song - instrumental mixed with vocals) with vocal result of separation. It might increase vocal residues outside busy mix parts.

Inverting vocals instead of mixture will result in less residues, but more artificial results in busy mix parts.

A similar trick might even increase SDR for MDX23C models irc.

How to perform inversion is explained somewhere in this doc by Bas Curtiz.

It might be unnecessary to use in UVR - it might use this trick for BS-Roformer models already, but for 2024.02 on MVSEP it was beneficial.

The trick is not necessary for the 04.2024 BS-Roformer model (it sounds worse after inverting).

Furthermore, for some muddiness in this model, you can use the premium’s feature - ensemble. The default output without intermediates should be enough (min_fft is very muddy, and max_fft very noisy). Strangely, the result from Roformer from intermediates might sound v. slightly better (maybe it was something random). The ensemble is kinda mimicked in jarredou’s MDX23 v2.4 Colab and to some extend it can be mimicked in UVR by using 1296+1297+MDX23 HQ ensemble (or copy of 1296 result via Manual ensemble instead, for faster processing).

Now also x-minus has a drums ensemble feature for Roformer models.

32b. Fixing muddiness for MDX-Net (on example of HQ_3 model) - inverting trick

It's less muddy when mixture is inverted and mixed with separated vocals in louder parts, but vs the instrumental stem, it's worse in silent parts with less busy mix - then it has more vocal residues than the instrumental stem.

When vocals were inverted instead of mixture, it was more muddy, but still more residues were present vs OG inst. stem, just a bit less. Can't tell how it's SDR-wise.

So you can combine various fragments for the best results.

33. Descriptions of models, pt. 2

Muddiness of instrumentals in specific archs

Beside changing min/avg/max spec for MDX ensembling (or in Colab for single models), plus aggression for VR models, or manipulating shifts and overlap for Demucs models, you need to know that some models or AIs sound usually less muddy than others. Like e.g. VR tends to have less muddiness vs MDX-Net v2 arch, but the first tends to have more vocal residues. Consider using HQ2/3/4/inst3/Kim inst for fewer residues than in VR arch or BS-Roformer.

For less muddiness than in MDX-Net, consider using MDX23 Colab 2.0/2.1 or 2.2 (more residues) or KaraFan (e.g. preset 5).

34. Muddiness of 4/+ stem results after mixdown

UVR5 supports even 64 bit output for Demucs, eventually you can use Colab or CML version for 32-bit float, but mvsep.com supports 32 bit output in MDX23 model when you choose WAV. It has better SDR vs Demucs, anyway, but sometimes more vocal residues.

Then, on MVSEP beside 4 stems, you have also instrumental - ready mixture of the three for instrumental in 32 bit provided, which is not bad, but you can go to extreme, and download e.g. Cakewalk, and 3 stems separately, and now in Cakewalk:

1) Don't use splash screen project creation tool, close it

2) Go to new

3) Pick 44100 and 64 bit

4) Make sure that double 64 bit precision is enabled in options

5) Import MDX23 3 stems (without vocals)

6) Go to file>Export

7) Pick WAV 64

Output files of 64 bit mixdown are huge, but that way you get the least amount of muddiness as possible. If only MDX23 model doesn't give you much more vocal residues vs MDX-UVR inst models or top ensemble which you wouldn't accept.

Be aware that 32-bit float vs 16 bit outputs can sound more muddy. Probably due to the fact that most sound cards/DACs don’t have native 32-bit float output support in drivers and additional downsampling must be done in-fly during playback, probably even if some drivers allow using 32-bit output in Sound settings in Control Panel for the same device (while other version might not).

Spectrum-wise, instrumentals downloaded from MVSEP vs manual mixdowns are nearly identical. The only difference in one case I saw was in an instrumental intro in the song where the site's instrumental had more high end, maybe noise, but besides, spectrum looks identical at first glance without zooming it. Still, when I performed mixdown to anything lower than 64 bit, I didn't get comparable clarity to the site's instrumental. Maybe I'd need to change some settings, e.g. change project bit depth to the same 32 bits as stems and later perform mixdown to 64 bit. Haven't tested it yet.

35. Debleeding of drums in vocals by Sam Hocking

“For drums, I usually try and do some kind of sidechained denoise using the demixed Drum stem itself as the signal to invert with. If you shape 'shape' the sidechained input using spectral tools/filters/transient tools etc, you can often null more of the drum out of the vocal. My favourite tool for this is Bitwig Spectral Split, but there's several FFT Spectral VSTs out there. The key is the tools has smoothing to extend the transients in time a bit so they null more.

Difficult to audibly hear on a video, but here's a vocal stem with a lot of residue I've exaggerated in a passage without singing. I turn on a sidechain bass, drums and other stem to phase invert them out the vocal a bit via the spectral transient split in Bitwig. I then take a spectral noiseprint in Acon Digital of what's left and that works as a mild denoiser, but only after the inversion has done its thing. Don't take the noise print until you're happy, everything else is inverting out as much as you can get it, and it's not noticeable.”

36. Manual MDX23 stems mixdown issues

It can happen that after importing three stems from MDX23 or other arch, into the same session, all combined they sound so loud that they clip on the master fader. I’d rather suggest that, in many cases it can be ignored, as after mixdown it will be fine in most cases and better than with using limiter, but it also depends on a song loudness of how much clipping even the instrumental from single model will have:

37. Q: Why sometimes separated instrumentals have clipping?

A: “Mixture doesn't clip, but instrumental is clipping.

This is because where the instrumental is clipping in positive values, the vocals are in negative values, and so vocals are lowering instrumental peak value when mixed together.

If you separate a song peaking at 0 with high loudness, the instrumental will probably clip because of this (and the more loudness, the more chances this clipping can happen, as waveform is brickwalled toward boundaries values). It's the laws of physics, as that's because of these laws that audio phase/polarity inversion works.

That's why Demucs is using the "clamp" thing, or can also lower the volume of the separated stem to avoid that clipping.

- Most of the time, lowering your input by 3dB solves that issue

- Saving your audio to float32 can be a solution, as "clipped" audio data is not lost in this case” (jarredou)

So theoretically in a 32-bit float, the volume can be decreased after separation and still nothing is lost, and clipping should be fixed.

- UVR has normalization option which turns down the output volume to avoid clipping if necessary, if you don't plan to use 32 bit float and potentially turn down the volume manually later

38. Separated audio using MDX-Net arch has noise when mixture has no audio and is silent

Use denoise standard (or denoise model) in Options>Choose Advanced Menu>Advanced MDX-Net Options>Denoise output

39. MDX23C/BS-Roformer models ringing issue

“It was reported that maybe a DC offset can amplify it. Fixing it with RX before separation was said to alleviate the issue” See the screenshot how to do it,

“Don't forget to use "mix" pasting mode” - jarredou

It serves to alleviate the issue of horizontal lines in specific frequencies across the whole track, cause most likely by bandsplitting neural network artifacts. Problem presented above.

Q: Mine is 0.047% for the DC offset, so I would just do 0.047 or 0.04

A: “0.047% is kind of a normal value, it's even a great one. No need to fix that.

I don't know at what value it could become problematic for source separation models.

On some raw instrument recordings, I have seen 20%~30% DC offset sometimes, which can become a real issue for mixing then, as it's reducing headroom” - jarredou

40. Ensemble of pitch shifted results (point 4 continues)

So you follow point 4, and “change sample rate before each separation and restore it after for each, then ensemble them all”.

“on drums it was really working great, where sometimes you have sudden muffled snare because other [stem] masked it, the SRS ensemble [irc used in MDX23 2.x Colab and KaraFan] was helping a lot with that, making separation more coherent across the track.”

41. A5 method for clean separations

Consider the fake stereo trick fist from point 11, separate with BS-Roformer 1296, clean the residues in vocals manually, put the vocals back into mixture - so perform mixdown to have a mixture again, and then separate this mixture with demucs_ft (old models)

42. Using surround versions of songs

Sometimes you can get vocals separated easier from the center channel from the surround version of the song. Perhaps you might also get different separations of instrumentals from such versions, also with possibility of manipulating the volume of specific tracks before mixdown to 2.0/stereo file. It might be necessary anyway, because otherwise you might run into some errors on an attempt of separation of 5.1 file or with more channels.

Use Dolby Atmos/360 Reality Audio/5.1 version of the song

Multichannel mixes can give better results for separation. For more on Atmos read.

Be aware that the center may contain not only vocal, but also some effects.

Consider separating every channel separately, or one pair of channels at the time (rear, front, center, sides separately) or only separate center channel separately and all the rest separately.

Visit this section for more information.

43. Matchering as substitute of ensemble (UVR>Audio Tools)

If the result of some separation is too noisy, but it preserved the mood and clarity of the instrumental much better than some cleaner, but muddy result, you can use that noisy result

as the reference for more muddy target file. E.g. voc_ft used as reference for GSEP 2 stem instrumental output.

44. Retry separation 2–3 times

At least for MDX23C models it happened for someone, that every separation made in UVR differed in terms of muddiness and residues, and someone received satisfactory result after the second or third attempt of separating the same song. Consider turning on Test mode in UVR, so the few digits number will be added to the output file name, so the results won’t be overwritten during the process, and you’ll be able to listen and compare them.

45. Ensemble instrumental result with drums with max/max

Can help to fix muddiness of vocal BS-Roformer models, but drums can sound too loud in the end. Consider decreasing their volume before ensemble if necessary.

Drums can be obtained from e.g. demucs_ft (and mixture as input or from some less muddy model) or from MDX23 Colab/MVSEP (which already uses its own input from model ensemble for 4 stems)

46. Use EQ on your song before separation (e.g. for too weak “s” sounds in separated vocals)

It’s an old method used in times when models didn’t give good quality yet, might no longer be necessary after release of Roformers (e.g. Vinctekan claims it doesn't even do anything anymore). Theoretically you could use EQ on a mixture to stress vocals in the mix more, so the separation might turn out to be better.

47. Bas Curtiz video tutorial for tips and tricks and document

48. Aufr33’s demudder (more for Roformers than HQ 4)

49. Volume compensation fine-tuning for MDX-Net models

It can slightly enhance the result, helping fighting muddiness a bit.

It’s no longer beneficial for MDX23C and Roformer models.

Volume compensation generally differ for every song. E.g. for HQ_3 model, sometimes 1.035 can be the best, but sometimes 1.022. By default, it affects only vocals, but when you switch primary stem in model settings, so vocals are labelled as instrumentals and vice versa (so how MDX kae Colab works), it can be used also to fine tune instrumental stems.

50. Picking correct models for ensemble (by dca100fb8)

“I'm seeing a certain pattern, if the Mel-Roformer model from x minus leaves faint vocals in the background during silent parts of the song, then it means MDX23 & Demucs 4 htdemucs_ft models should not be used for ensemble because vocals can be heard in the background too using these models, while MDXv2 models will not leave those vocals. So it's either UVR Mel/BS-Roformer 1296 + 1297 + MDXv2 or MDX23 + Demucs + Mel-Roformer X Minus + BS-Roformer 1296 + 1297.

I excluded VitLarge because it always leaves faint vocals”

51. Ensemble only extra higher frequencies 

from e.g. HQ 3 model with narrowband inst 3 model - guide

52. Use some BV/Karaoke model first, to potentially get cleaner instrumental with dedicated model afterwards

53. Set vocals to center with stereo plugin (guide by Musicalman)

Trick [working] with the [now outdated]  BS-Roformer karaoke model, though it may work on other karaoke models too (I suspect you might have some mileage with MDX for instance). Anyway, the trick has to do with separating one voice from other sounds. If the voice you want to separate is panned centrally, you're already in luck; the model should expertly separate it. If not, you can rotate the stereo field so that the voice is as close to the center as possible (I use the Reaper js stereo field manipulator plug in for this). Process the rotated sound with the karaoke model and the voice you're looking for will magically be separated, even from other voices! If you need the original stereo image back, simply perform the opposite rotation.

54. Method for cleaner vocals (by YAZKEN*)

“Basically you do 2 vocal extractions, invert the polarity of one of them and render it, after that you invert the rendered audio and choose one of the extractions you’ve made and listen [to] what is cleaner” sounds familiar to what TTA actually does.

55. Spectral editing in Audacity explained (by CC Karaoke)

I typically use Audacity... But [...] RX11 has some nice shiny toys, so maybe try that. I do things the hard way.

Here's a great basic example when using the (Roformer) Karaoke model. It can really sometimes struggle with the hard consonants.

So for best results, you'll often need to isolate those in the main vocal stem by muting out all the surrounding sound, and then mixing them back in with the backing vocal stem. Of course by doing this the hard consonants will often be too loud, so you can de-amp the volume on them and then play back till you get a level that sounds like it blends properly.

https://imgur.com/a/dyc9olh

Slightly more complex example; Where the vocal lines are overlapping. I tried drawing green over one of the lines to show the difference. Might be a couple mistakes lol, as I haven't checked, but you get the idea. The previous sound is a carrying note, whereas the next line is a 'HA' kinda hard hit punching words, so it has a different shape to it. This kind of more obvious difference is easier than say... reverb…

https://imgur.com/a/BFGqN7P

jarredou’s hint: That's a case where I would go SpectraLayers as while the 2 vocals are not on the same pitch, you can separate them manually (with harmonic selection tool). At least for that small part shown here.

In SpectraLayers, you can change FFT resolution, higher value will give you more defined freq "picture", and it can help when 2 parts are really close in pitch, like here.

The downside is that with high FFT values, you lose time resolution. So to use SpectraLayers manual selection efficiently, you often need to switch that FFT resolution value depending on the elements you are targeting, like you would zoom/dezoom in photoshop while editing a picture.

56. Sequential stem separation (by dynamic64/isling)

With single stem models, feel free to experiment with sequential stem separation -

Instrumental model first, then drums or bass, piano or guitar, strings or horns. It depends on the song whether better results will give e.g. drums or bass when separated first, the same to piano vs guitar and strings vs horns first.

57. Advanced chain processing chart (image)

It’s a method utilizing old models, and e.g. Kim Vocals 2 can be potentially replaced by unwa’s BS/Mel-Roformer models in beta UVR (or other good method for vocals) or ensembles mentioned in this document. Check the best current methods for vocals in one stem to find what works the best for your song to get all vocals before splitting to other stems using this diagram.

htdemucs v4 above can be replaced by htdemucs_ft, as it's the fine-tuned version of the model (or MDX23 Colab). Even better, you can use some of the methods for 4 stems in this GDoc (like drums on x-minus).

De-echo and reverb models can be potentially replaced by some better paid plugins like:

DeVerberate by Acon Digital, Accentize DeRoom Pro (more in the de-reverb section).

UVR Denoise can be potentially replaced by less aggressive Aufr33 model on x-minus.pro (used when aggressiveness is set to minimum), and there’s also newer Mel-Roformer (read de-reverb section).

As for Karaoke models, there's e.g. a Mel-Roformer model on x-minus.pro for premium users or MVSEP/jarredeou inference Colab.

"If the vocals don't contain harmonies, this model (Mel) is better. In other cases, it is better to use the MDX+UVR Chain ensemble for now.". It is possible to recreate to some extent this approach while not using BVE v2 models, by processing the output of main vocal model by one of Karaoke/BVE models in UVR (possibly VR model as the latter) using Settings>Additional Settings>Vocal Splitter Options, so it separates using one model, then it uses the result as input for the next model (see the Karaoke section).

MedleyVox (not available in UVR) will be useful in the end in cases when everything else fails after you obtain all vocals in one stem, as it's very narrowband. But you can use AudioSR on it afterwards.

58. See here for more on cleaning/debleeding

59. Reverse polarity and/or remove DC offset of the input file

60. Find fragments of instrumentals in your song and overlap them inverted across the whole song before separation (heauxdontlast)

60. Method for better quality of instrumental leaks on YT by theamogusguy

“I did something really odd. (...) since you can only rip max 128 kbps I did something really odd to get a higher quality instrumental:

I inverted the 128 kbps AAC YouTube rip into the original to get the acapella

I took the subtracted acapella and ran it through AI (Mel-Roformer 2024.10) to reduce the compression artifacts

I then inverted the isolated acapella and mixed it with the lossless to get an... unusual lossless instrumental file?
Also, the OPUS stream goes up to 20 kHz, but I feel like the sample rate difference is going to cause issues, so I ended up ripping AAC (OPUS is 48 kHz while most music is 44.1 kHz)”

61. Join the best fragments from various models

E.g. unwa inst models might be noisy at times, so you might want to use specific fragments of v1e/v1/v2 fitting across the song, or e.g. beta 4 vocal model in certain fragments where it’s not enough, though it is more muddy, but less noisy than unwa’s inst models. In some cases, if it’s still not enough, you might want to use BS-Roformer models like unwa’s Large or e.g. 24.10 on MVSEP. Just find which model on the list in this document has the least amounts of residues and experiment with the rest starting from models listed at the top.

62. Lowpassing lossless file to 20kHz

Sometimes it’s a bit useful in getting rid of some constant faint noise/residues from vocals in instrumentals. It might muffle some unwanted parts of the instrumental, but some more difficult fragments with more residues than usual might sound better that way. Tested on FLAC 16 compressed to mp3 320 kbps, but it should work better with lowpassing using EQ instead of compressing. Other example values you might want to try out using are 19 kHz (mp3 VBR V0 cutoff)/17.7 kHz (cutoff of some narrowband models)/16kHz (cutoff of mp3 and AAC 128 kbps)/14.7 kHz (D1581 model cutoff).

A possible explanation of why it might sometimes work is: sometimes, e.g. more oldschool hip-hop beats might have less higher tones, or even none above 16 kHz, so most of the information in this area might come from vocals in a mixture. You can recognize it especially if vocals lose much more clarity than beat in the mixture once you compress it to e.g. mp3 VBR V0 (19kHz cutoff) or lower.

63. Refrain from excessively stacking models (e.g. for RVC)

“Inst Voc, Kim Vocals, Denoise, ensemble mode, and so forth can introduce noises to your dataset as it rips away frequencies from your audio. This harms the model's fidelity and quality.” more

64. Get cleaner vocals with vocal and instrumental model mixdown (e.g. of Mel becruily models) by Havoc/mrmason347

Separate with becruily Mel Vocal model and its instrumental model variant, then get vocals from the vocal model, and instrumental from instrumental model, import both stems for the DAW of your choice (can be Audacity) so you’ll get a file sounding like original file, then export - perform a mixdown of both stems, then separate it with vocal model

65. Less vocal bleed with dim_t 256 or corresponding chunk_size (cypha_sarin)


Small difference observed on 6GB NVIDIA GPU and Gabox instv5 model where “one little vocal glitching sound from the song that only gets picked up when the segment size is lower [256]”

66. If you set 24-bit output in UVR>Options>Additional settings (or ev. 64-bit) for e.g. demudder, the results might be slightly less muddy

67. Clean loop of the instrumental used for Matechering and full separation

You can use a well sounding fragment of single instrumental model separation with high fullness metric as a reference for Matchering in UVR for phase-fixed muddy result set as target. It will have less bleeding than models with low bleedless metric, but still fuller than phase-fixed results.

You can even use fade-in and/or fade/out to reconstruct full loops of instrumental. It can be done by normalizing fading out peak with everything later, and again, and again. So you start normalizing from most of the peaks starting to fade out (so already below certain normalizing threshold, e.g. 0dB). Then correct volume of specific fragments to avoid any pumping or sudden changes of volume across the loop.

68. chunk_size 112455 and overlap 50

To have the best SDR for Roformers, use chunks not lower than 11s, which is usually training chunks value (rarely higher). Although, at times people get better results with 2,55s chunks (called chunk_size 112455 since UVR Roformer patch #3). But be aware that e.g. using becruily Karaoke model, using a low 2,55s chunk will lead to crossbleeding. dim_t to chunk_size conversion is later here. Sometimes even go to extremes and use e.g. overlap 50 claiming that it was better with 112455 and becruily inst model (thx gustownis). Generally every model might have a different chunk_size value at which the SDR is the highest, but maybe not the best for specific bleedless/fullness metric.

69. If you want smoother vocals from e.g. Beta 5e, use negative values of Shift Pitch Conversion in UVR Advanced MDX-Net settings (explained more thoroughly above).

“Tried it on a regular model (bigBeta5e) - the spectrogram looks a little more cut off at the high end than without the pitch adjust and overall the vocal sounds a little rounder and not quite as harsh (so the transients are not so nuclear)” - cristouk

70. Fixing missing sound after separation of multistem models

With certain at least 4 stem models, you might find out that the inversion of a mixdown of those 4 stems vs original mixture is different. So you might get an additional 5th stem that way - your own “other”. It might be useful if some instruments got missed, or simply for remastering purposes where not having any missed bits of audio is critical for your work.

71. Start separation in a different place of the song

Cut it manually. The result might resemble changing chunks setting a bit.

72. Use instrumental model result as pre-processor for vocal model

It’s one of suggested RVC workflows

73. Don't use 96kHz and higher audio files in UVR for separation

For some reason it yields bad results and clipping. 48kHz are allright.

As a workaround, you could use: https://github.com/rorgoroth/mingw-cmake-env/releases/tag/latest ffmpeg -i "C:\input96or48.wav" -af asf2sf=dblp,ardftsrc=44100:quality=61656210:bandwidth=0.9941249 -c:a pcm_f32le C:\output44.wav 
It will give you a 32-bit float file after downsampling, so you could turn down the volume before separation (probably you could even settle on some volume attenuation in the command itself). Ardftsrc with these parameters was #1 on the resamplers chart on hydrogenaudio at the time. It requires 24GB of RAM to process faster (if you have, turn on two SSDs for pagefile if you have 16GB RAM or less). Be aware that the script doesn't track the progress, you need to just wait after executing tipo it finishes. In a limited RAM scenario, it can take even 10 minutes or more for a 5-minute song. Ardftsrc can be also used in the latest Foobar2000 versions.

74. Pre-process your loud song/mixture with  de-limiter to get less distortion in vocals after separation with model

75. Try out using MSST instead of UVR. We had at least one report with BS single stem phantom center that UVR gave more bleeding despite using the same settings. You might try out Colabs as well (they're based on older MSST code before refactoring).

76. Taming top end artefacts

At least some Roformers tend to have some top-end buildup noticeable on the spectrum, which you could just lowpassed in EQ like TDR Nova. Although it might be not noticeable, it might be useful in further post-processing.


—--

Vinctekan write-up on some tips since the release of Roformers and MVSep Demucs bass fine-tune:

“- Extracting center channel for things like bass guitar now on its own produces worse results when separated with modern models like BS-Rofo, the SW variant, the finetuned Demucs Bass model on MVSEP, and the SCNet model as well. In order for this to be worth it you would have to check with something like Spectralayers to compare spectrogram of the original separation, and the separation with center channel extraction beforehand. There may be a few harmonics here or a transient there that the center channel separation did better than the original, and in order to combine these results while maximizing clarity you will have to perform a maximum spectrogram ensemble.

- Normalizing the audio to ANY DB peak does nothing now, except maybe makes the noise floor different (which is useless).

- Upward compressing/limiting/maximizing only changes the output stems dynamic, not the quality.

- Separation of other stems from the song beforehand is also almost completely useless, especially if it comes from the same multistem model.

- Adding static EQ does nothing, does not matter what the settings are.

- Putting the song through a spectrogram STFT mask suffers from the same fate as Center Channel Extraction, the results on average are a lot worse, and you would have to perform a max spectrogram ensemble in order to get the best of both outputs. It super rarely adds anything different/useful now (...)

You would think: "Oh, If I extract the center channel, then the model has less audio to dig through, and the desired stem (vocals or bass or whatever) is overlapping with less useless audio and therefore is more audible, cleaner sounding. Therefore separation will be better, YAY!"

But this stopped being the case for well over 2 years now…”


___

Get UVR VIP models (optional donation)

https://www.buymeacoffee.com/uvr5/vip-model-download-instructions

If you still see some missing models in UVR5 GUI, which are mentioned in this document, get them from download center (or here, expansion pack) and click refresh in model list if you don't see some models.

_______________________________________________________