Google Doc

KaraFan by Captain FLAM

Page 15 of 28 · Edit this page in Google Docs ↗

(2 stems)

It was made before Roformers models (potentially deprecated)

Colab w/ more models (AI Hub fork, also fixed), fixed org. Colab, org. Colab (slow), GUI, GH documentation

How to install it locally (advanced), alt. tutorial,

or easy instruction

Should work on Mac with Silicon or AMD GPU (although not for everyone)

& Linux with Nvidia or AMD GPU

& Windows probably with at least Nvidia GPU, or with CPU (v. slow)

- For Colab users - create “Music” in the main GDrive directory and upload your files for separation there (the code won’t create the folder on the first launch).

- Sometimes you’ll encounter soundfile errors during separation. Just retry, and it will work

KaraFan (don’t confuse with KaraFun) is a direct derivative of ZFTurbo’s MDX23 code forked by jarredou, but with further tweaks and tricks in order to get the best quality of instrumentals and vocals sonically, but without overfocusing on SDR only, but the overall sound.

Its aim is to not increase vocal residues without making instrumentals too muddy like e.g. sometimes HQ_3 model does, but without having so many vocal residues as MDX23C fullband model (but it depends on chosen preset).

Since v. 4.4 and 5.x you have five presets to test out.

Presets 3 and 4 are more aggressive in canceling vocal residues (P4 can be good for vocals).

Preset 5 (takes 12 minutes+ on the slowest setting for 3:25 track on T4) has more clarity of instrumentals over presets 3 and 4, but also more vocal residues (although less than P1 and P2 (takes 8 minutes for 3:24 track on the slowest setting).

On 23.11.24 “Preset 5 was corrected to be less aggressive as possible”. All the below Preset 5 descriptions refer to the old P5. The original preset 5 is here, and is less muddy, but has more vocal residues (at least the original preset contains more models and is slower).

Speed and chunks affect quality. The slower, the muddier, but also slightly less vocal residues, although they’ll be still there (just slightly quieter). I’d recommend the “fastest” Speed setting and 400K chunks for the current P5 (tested on 4:07 song, may not work for longer tracks).

- If you replace Inst Voc HQ1 model by HQ2 using AI Hub fork in current P5, the instrumental will be muddier.

- To preserve instruments which are counted as vocals by other MDXv2 models, use these preset’s 5 modified settings - they have more clarity than P5 and preserve hi-hats better. But to preserve the same processing time as in P5, but setting “Speed” slider to medium, in this case will result in more constant vocal residues vs P5 with the slowest setting (too much at times, but it might serve well for specific song fragments). It will take 12 minutes+ for 3:24 track on medium. Debug and God mode on the screenshot are unrelated and optional.

- To fix issues with saxophone in P5 use these settings. They even have more clarity than the one above, but also more hearable vocal residues. It helps to preserve instruments better than the setting from the above. It can be better than P2 - less hearable consistent vocal residues, but in similar amount, while on other artists sax preset even gives more vocal residues than P2. Sax setting is worse in preserving piano than the setting above.

- Using the slowest setting here in sax fix preset will result in disconnection of runtime with free T4 after 28 minutes of processing, but it should succeed anyway (result files might be uploaded on GDrive after some time anyway).

Vs medium, the slowest setting gives more muffled sound, but not always less vocal residues. It can be heard the best in short parts with only vocals. 18 minutes for 4:07 track on Fast setting (God Mode and Debug Mode are disabled in KaraFan by default).

After 3-4 ~18 minutes separations (in this case not made in batch, but with manually changed parameters in the middle), when you terminate and delete environment, you might be not able to connect with GPUs again as the limit will be reached unless you switch Colab account (mount the same GDrive account as Colab to avoid errors)

- Preset 5 provides more muffled results than the two settings above, but with good balance of clarity and vocal residues. Sometimes this one has less vocal residues, sometimes 16.66 MDX23C model on MVSEP (or possibly a bit older HQ_1 model in UVR), it can even depend on a song fragment.  Using newer MDX23C HQ 2 in P5 instead of MDX23C HQ doesn’t seem to produce better results

After 5th separation (not in batch) you must start your next separation very fast because or you’ll run out of Colab free limit when GUI is in idle state. In such case, switch Colab account, and use the same account to mount GDrive (or you might encounter error).

Comparisons above made with normalization disabled and 32-bit float setting

The code handles mono and 48kHz files too, 6:16 (preset 3) tracks, and possibly 9 minutes tracks too (but can’t tell if with all presets). It stores models on GDrive, which takes 0,8-1,1GB (depending on how many models you’ll use). One 4:07 song in 32-bit float with debug mode enabled (all intermediate files will be kept) will take 1,1GB on GDrive. Instrumentals will be stored in files marked as Final (in the end), Music Sub (can sound a bit cleaner at times, but with more residues), and Music Extract (from specific models).

Older 1.3 version Colab fork by Kubinka was deleted.

Colab fork made by AI HUB server members also includes MDX23C Inst Voc HQ 2 and HQ_4 models, and contains slow separation fix from the “fixed Colab”.

KaraFan used to have lots of versions which differ in these aspects with an aim to have the best result in the recent Colab/GUI version. E.g. v.3.1 used to have more vocal residues than in 1.3 version and even more than in HQ_3 model on its own, and it got partially fixed in 3.2 (if not entirely). But 1.3 irc, had some overlapped frequency issue with SRS disabled, which makes the instrumentals brighter, but it got fixed later. The current version at the time of writing this excerpt is 4.2, with pretty good opinions for v.4.1 shortly before.

Colab troubleshooting

- (no longer necessary in the fixed Colab) If you suffer from very slow or unfinishable separations in the Colab using non-MDX23C models (e.g. stuck on voc_ft without any progress), use fixed Colab (the onnxruntime-gpu line added in the end of the first cell)

- Contrary to every other Colab in this document, KaraFan uses a GUI which launches after executing inference cell. It triggers Google’s timeout security checks frequently esp. in free Colab users, because Google behaves like the separation is not being executed where you do it in GUI, and it’s generally against their policies to execute such code instead pasting commands to execute in Colab cells directly. The same way many RVC Colabs got blocked by Google, but this one is generally not directly for voice cloning, and is not very popular yet, so it wasn’t targeted by Google yet.

- Once you start separation, it can get you disconnected from runtime quickly, especially if you miss some multiple captcha prompts (in 2024 captchas stopped appearing at all, so the user inactivity during separation process seems to be no longer checked).

- After runtime disconnection error, output folder on e.g. GDrive can be still constantly populated with new files, while progress bar is not being refreshed after clicking close or even after closing your tab with Colab opened. At certain point it can interrupt the process, leaving you with not all output files. Be aware that final files always have “Final” in their names.

- It can consume free "credits" till you click Environment>Terminate session. It happens even if you close the Colab tab. You can check “This is the end” option so the GUI will terminate the session after separation is done to not drain your free limit.

- (rather fixed) As for 4.2 version, session crashes for free Colab users can occur, due to running out of memory. You can try out shorter files.

Currently, if you rename your output folder with separation, and retry separation, it will look for the old folder with separation to delete, and return the error, and running the GUI cell again may cause disappearing of GUI elements.

it's a default behavior of Colab and IPython core : Sync of files the Colab sees is not real time

Two possible solutions:

  • wait until sync with Google Drive is done
  • restart & run Colab

- Sometimes shutting down your environment in Environment options and starting over might do the trick if something doesn't work. E.g. (if it wasn't fixed), when you manipulate input files on GDrive when GUI is still opened, and you just finished separation, you might run into an error when you start separating another file with input folder content changed.

In order to avoid it, you need to run the GUI cell again after you've changed the input folder content (IRC it's "Music" folder by default). Maybe too low chunks (below 500k for too long tracks if something hasn't changed in the code). Also, check with some other input file you used before and worked before first.

Also, be more specific about what doesn't work. Provide screenshot and/or paste the error.

- You can be logged to a maximum of 10 Google accounts at the same time. You can’t log out of any of these single accounts on PC in browser. The only way is to do it on your Android phone, but it might not fix the problem, as it will tell “logged out” on that account on PC, and logging into other one might not work and the limit will be still exceeded. At this situation you can only logged out from all accounts (but it will break accounts order, so any authorizations set to specific accounts in your bookmarked links will be messed up - e.g. those to Colab, GDrive, Gmail, etc. I mean: /u/0 and in Colab authuser= in links. Easier way to access to extra Google account will be to log into it from Incognito mode.

If you possess lots of accounts and you don’t log for some for 2 years, Google can delete it. To avoid it, create YT channel on it, and upload at least one video, and the account won’t be deleted.

Tests of four presets of KF 4.4 vs MDX-UVR HQ_3 and MDX23C HQ (1648)

(noise gate enabled a.k.a. “Silent” option)

Not really demanding case, so without modern vocal chain in the mix, but probably enough to present the general idea of how different presets sound here. So, more forgiving song to MDX23C model this time, and less aggressive models with more clarity.

Genre: (older rap) Title: (O.S.T.R. - Tabasko [2002])

BEST Preset : 3

Music :

Versus P4, hi-hats are preserved better in P3.

Snare in P3 is not so muffled like in P4.

HQ_3 has even more muffled snares than in P4.

P3 still had less vocal residues than MDX23C HQ 1648 model, although the whole P3 result was more muffled, but residues are smartly muffled too.

MDX23C had like more faithfully sounding snares than P3, to the extent that they can be perceived brighter (but vocal residues, even on a more forgiving song like this, are more persistent in MDX23C than in P3).

Sometimes it depended on specific fragment where P4 and where P3 has more vocal residues in that specific case, so P3 turned out being pretty much balanced, although P4 had less consistent vocal residues, although still not so few like HQ_3, but it's not that much of a problem (HQ_3 is really muffled). If it was 4 stems, then I'd describe P3/4 as having very good "other" stem but drums too as I mentioned.

WORST Preset (in that case) : 1

Music : Too much consistent vocal residues

There's a similar situation in P2, but at least P2 has brighter snares than even MDX23C.

In other songs, P1 can be better than P2, leaving less vocal residues in specific fragments for a specific artist, but noticeably more for others.

Preset 4 with setting slow (but not the slowest) takes 16 minutes for 5 minutes song on T4 in free Colab (performance of ~GTX 3050). For 3:30 track, it takes 13:30 for the slowest setting. In KF 5.1 with default chunks 500K and slowest setting, for 4:50 song and preset 2 it took <10 minutes, preset 3, 12 minutes.

VS preset 3, the one from the screenshot (now added as preset 5) is more noisy and has more vocal residues, mainly in quiet places or when there is no instrumental. Processing time for 6:16 track on medium setting is 22:19 minutes. But it definitely has more clarity over preset 3. And there is still less vocal residues than in Preset 1 and 2, which have more clarity, but tend to have too many vocal residues in some tracks. Hence, preset 5 is the most universal for now.

For future: “To add or remove some models u need to edit the .csv file https://github.com/Eddycrack864/KaraFan/blob/master/Data/Models.csv

with the model info (Only MDX23C or MDX-NET) u can found the model info on the model_data_new.json: https://raw.githubusercontent.com/TRvlvr/application_data/main/mdx_model_data/model_data_new.json u need to find the hash of the model. And.... that's it! (Not Eddie)