Ripple/Capcut/SAMI-Bytedance/Volcengine/BS-RoFormer (2-4 stem)
(Ripple is discontinued since 31 January 2026)
Output quality in Ripple is: 256kbps M4A (320kbps max) and lossless (introduced later). 50MB upload limit, 4 stems
Min. iOS version: 14.1
Ripple is only for US region (which you can change, more below)
Ripple no longer separates stems (there's an error "couldn't complete processing please try again")
Ripple for iOS: https://apps.apple.com/us/app/ripple-music-creation-tool/id6447522624
Capcut for Android: https://play.google.com/store/apps/details?id=com.lemon.lvoverseas
(separation only for Pro, Indian users sometimes via VPN)
Capcut a.k.a. Jianying (2 stems) works also on Windows (only in Jianying Pro, separation option is available)
Can be used instead of Ripple if you're on unsupported iOS below 14.1 or don’t have iOS. To get Ripple you can also use a virtual machine remotely instead (instructions below). Ripple can also be run on your M1 Mac using app sideloading (instructions below).
Ripple = better quality than CapCut as of now (and fullband)
with fixed the click/artifacts using cross-fade technique between the chunks.
Capcut = “the results are really low quality but if you export the instrumental and invert it with the lossless track, you will get the vocals with the noise which is easy to remove with mdx voc ft for example, then you can invert the lossless processed vocals with the original and have it in better quality.
The vocals are very clean from cap cut, almost no drum bleed”
Ripple and Capcut uses SAMI-Bytadance arch (later known as BS-Roformer), but it’s a different model with worse SDR than on the leaderboard. It was developed by Bytedance (owner of TikTok) for MDX23 competition, and holds the top of our MVSEP leaderboard. It was published on iOS and for the US region as “Ripple - Music Creation Tool” app. Furthermore, it's a multifunctional app for audio editing, which also contains a 4 stem separation model. Similar situation with Capcut (which is 2 stems only IRC). The model itself is not the same as for MDX23 competition (SAMI ByteDance v1.0), as they said, models for apps were trained on 128kbps mp3 files to avoid copyright issues, but it’s the same arch, just scores a bit lower (even when exported losslessly for evaluation). SDR for Ripple is naturally better than for Capcut.
Seems like there is no other Pro variant for Capcut Android app, so you need to unlock regular version to Pro.
At least the unlocked version on apklite.me have a link to the regular version, so it doesn't seem to be Pro app behind any regional block. But -
"Indian users - Use VPN for Pro" as they say, so similar situation like we had on PC Capcut before. Can't guarantee that unlocked version on apklite.me is clean. I've never downloaded anything from there.
Bleeding
Bas Curtiz found out that decreasing volume of mixtures for Ripple by -3dB (sometimes -4dB) eliminates problems with vocal residues in instrumentals in Ripple. Video
This is the most balanced value, which still doesn't take too many details out of the song due to volume attenuation.
Other good values purely SDR-wise are -20dB>-8dB>-30dB>-6dB>-4dB> /wo vol. decr.
The method might be potentially beneficial for other models, and probably work best for the loudest tracks with brickwalled waveforms.
The other stem is gathered from inversion to speed up the separation process. The consequence is bleeding in instrumentals.
- If you suffer from bleeding in other stem of 4 stems Ripple, beside decreasing volume by e.g. 3/4dB also “when u throw the 'other stem' back into ripple 4 track split a second time, it works pretty well [to cancel the bleeding]”
The forte of the Ripple is currently vocals - the algo is very good at differentiating what is vocals and what is not, although they can sound “filtered” at times.
Currently, the best SDR for public model/AI, but it gives the best results for vocals in general. For instrumentals, it rather doesn’t beat paid Dango.ai (and rather not KaraFan and HQ_3 or 1648/MDX23C fullband too).
It's good for vocals, also for cleaning vocal inverts, and surprisingly good for e.g. Christmas songs, (it handled hip-hop, e.g. Drake pretty well). It's better for vocals than instrumentals due to residues in other stem - bass is very good, drums also decent, kicks even one if not the best out of all models, as they said some fine-tuning was applied to drums stem. Vocals can be used for inversion to get instrumentals, and it may sound clean, but rather not as good as what 2 stem option or 3 stem mixdown gives as output is lossy.
Capcut (2 stems only)
It is a new Windows and Android app which contains the same arch as Ripple inst/vocal, but lower quality model, and without an option of exporting 4 stems.
It normalizes the input, so you cannot use Bas’ trick to decrease volume by -3dB to workaround the issue of bleeding like in Ripple (unless you trick out the CapCut, possibly by adding some loud sound in the song with decreased volume).
“At the moment the separation is only available in Chinese version of Windows app which is jianyingpro, download available at capcut.cn [probably here - it’s where you’re redirected after you click “Alternate download link” on the main page, where download might not work at all]
Some people cannot find the settings on this screen in order to separate.
Separation doesn't require sign up/login, but exporting does, and requires VIP, which is paid depending on whether you’re from rich or poor country, then it’s free.
- There’s a workaround for people not able to split using Capcut for Windows in various regions.
- Bas Curtiz' new video on how to install and use Capcut for separation incl. exporting:
https://www.youtube.com/watch?v=ppfyl91bJIw
"It's a bit of a hassle to set it up, but do realize:
- This is the only way (besides Ripple on iOS) to run ByteDance's model (best based on SDR).
- Only the Chinese version has these VIP features; now u will have it in English
- Exporting is a paid feature (normally); now u get it for free
The instructions displayed in the video are also in the YouTube description."
- mitmproxy script allowing to save to FLAC instead of AAC (although it just reencodes from AAC 113kbps with 15.6kHz lowpass filter). It’s a bit more than script. See the full tutorial.
- For some people using mitmproxy scripts for Capcut (but not everyone), they “changed their security to reject all incoming packet which was run through mitmproxy. I saw the mitmproxy log said the certificate for TLS not allowed to connect to their site to get their API. And there are some errors on mitmproxy such as events.py or bla bla bla... and Capcut always warning unstable network, then processing stop to 60% without finish.” ~hendry.setiadi
“At 60% it looks like the progress isn't going up, but give it idk, 1 min tops, and it splits fine.” - Bas
“in order to install pydub within mitmproxy, you additionally need to:
open up CMD
pip install mitmproxy
pip install pydub”
- IntroC created a script for mitmproxy for Capcut allowing fullband output, by slowing down the track. Video
Older Capcut instruction:
The video demonstration of below:
0. Go offline.
1. Install the Chinese version from capcut.cn
2. Use these files copied over your current Chinese installation in:
C:\Users\(your account)\AppData\Local\JianyingPro
Don’t use English patch provided below (or the separation option will be gone)
3. Now open CapCut, go online after closing welcome screen, happy converting!
4. Before you close the app, go offline again (or the separation option will be gone later).
! Before reopening the app, go offline again, open the app, close welcome screen, go online, separate, go offline, close. If you happen to missed that step, you need to start from the beginning of the instruction.
(no longer works after 4.6 to 4.7 update, as it freezes the app) The only thing that seems to enable vocal separation without requiring replacing everything, is to replace that SettingsSDK folder contents inside User Data. It's probably the settings_json file inside responsible for that.
FYI - the app doesn’t separate files locally.
The quality of separation vs Capcut is not exactly the same as Ripple. Seeing by spectrograms, there is a bit more information in vocals in Capcut, while Ripple has a bit more information in spectrum in instrumentals.
Separated vocal file is encrypted and located in C:\Users\yourusername\AppData\Local\JianyingPro\User Data\Cache\audioWave”
The unencrypted audio file in AAC format is located at \JianyingPro Drafts\yourprojectname\Resources\audioAlg (ends with download.aac)
“To get the full playable audio in mp3 format, a trick that you can do is drag and drop the download.aac file into Capcut and then go to export and select mp3. It will output the original file without randomisation or skipping parts”
(although it resulted in VIP option disappearing but Bas somehow managed to integrate it in his new video tutorial, and it started to work, English translation isn't the culprit of the problem, but if you use both language pack and SettingsSDK folder from above)
You can replace the zh-Hans.po file with English one to have English language on Chinese version of the app possessing separation feature in:
jianyingpro/4.6.1.10576/Resources/po
While you can’t use that language pack, you can always use Google Translate to transform Chinese into your own language on a screen of your smartphone.
https://support.google.com/translate/answer/6142483?hl=en&co=GENIE.Platform%3DDesktop
“Trying out capcut, the quality seems the same as the Ripple app (low bitrate mp3 quality)
at least the voice leftover bug is fixed, lol”
Random vocal pops from Ripple are fixed here.
Also, it still has the same clicks every 25 seconds as before in Ripple.
Capcut adds 1024 extra samples at the beginning, and 16 extra samples at the end of the file.
How to change region to US
in order to make Ripple work on iOS
in Apple App Store to make "Ripple - Music Creation Tool" (SAMI-Bytedance) work.
https://support.apple.com/en-gb/HT201389
- Bas' guide to change region to US for Ripple on iOS
https://www.bestrandoms.com/random-address-in-us
Or use this Walmart address in Texas, the number belongs to an airport.
Do it in App Store (where you have the person-icon in top right).
You don't have to fill credit cards details, when you are rejected,
reboot, check region/country... and it can be set to the US already.
Although, it can happen for some users that it won't let you download anything forcing your real country.
"I got an error because the zip code was wrong (I did enter random numbers) and it got stuck even after changing it.
So I started from the beginning, typed in all the correct info, and voilà"
If ''you have a store credit balance; you must spend your balance before you can change stores''.
It needs (an old?) a simcard to log your old account out if necessary
Ripple on Windows or MacOS
- Another way to use Ripple without Apple device -
virtual machine
Sideloading of this mobile iOS app is possible on at least M1 Macs.
- Saucelabs
Sign up at https://saucelabs.com/sign-up
Verify your email, upload this as the IPA: https://decrypt.day/app/id6447522624/dl/cllm55sbo01nfoj7yjfiyucaa
Rotating puzzle captcha for TikTok account can be tasking due to low framerate. Some people can do it after two tries, others will sooner run out of credits, or completely unable to do it.
Mobile device cloud
- Scaleway
"if you're desperate you can rent an M1 Mac on scaleway and run the app through that for $0.11 an hour using this https://github.com/PlayCover/PlayCover”
IPA file:
https://www.dropbox.com/s/z766tfysix5gt04/com.ripple.ios.appstore_1.9.1_und3fined.ipa?dl=0
"been working like a dream for me on an M1 Pro… I've separated 20+ songs in the last hour"
More info:
-https://cdn.discordapp.com/attachments/708579735583588366/1146136170342920302/image.png
- “keep in mind that the vm has to be up for 24 hours before you can remove it, so it'll be a couple bucks in total to use it”
Fixing chunking artefacts (probably fixed)
- Every 8 seconds there is an artifact of chunking in Ripple. Heal feature in Adobe Audition works really well for it:
https://www.youtube.com/watch?v=Qqd8Wjqtx-8
-The same explained on RX10 example and its Declick feature:
https://www.youtube.com/watch?v=pD3D7f3ungk
Volcengine (a.k.a. The sami-api-bs-4track - 10.8696 SDR Vocals)
https://www.volcengine.com/docs/6489/72011
Ripple/SAMI Bytedance's API was found. If you're Chinese, you can go through it easier -
you need to pass the Volcengine facial/document recognition, apparently only available to Chinese people
We already evaluated its SDR, and it even scored a bit better than Ripple itself.
"API from volcengine only return 1 stem result from 1 request, and it offers vocal+inst only, other stems not provided. So making a quality checker result on vocal + instrument will cost 2x of its API charging.
Something good is that volcengine API offers 100 min free for new users"
API is paid 0.2 CNY per minute.
It takes around 30 seconds for one song.
It was 1.272 USD for separating 1 stem out MVSEP's multisong dataset (100 tracks x 1 minute).
"My only thought is trying an iOS Emulator, but every single free one I've tried isn't far-fetched where you can actually download apps, or import files that is"
So far, Ripple didn't beat voc_ft (although there might be cases when it's better) and Dango.
Samples we got months ago are very similar to those from the app, also *.models files have SAMI header and MSS in model files (which use their own encryption), although processing is probably fully reliable on external servers as the app doesn't work offline (also model files are suspiciously small - few megabytes, although it's specific for mobilenet models). It's probably not the final iteration of their model, as they allegedly told someone they were afraid that their model will leak, but better than the first iteration judging by SDR with even lossy input files.
Later they told that it’s different model than the one they previously evaluated, and that time it was trained with lossy 128kbps files due to some “copyright issues”.
"One thing you will notice is that in the Strings & Other stem there is a good chunk of residue/bleed from the other stems, the drum/vocal/bass stems all have very little to no residue/bleed" doesn't exist in all songs.
It's fully server-based, so they may be afraid of heavy traffic publishing Ripple worldwide, and it's not certain whether it will happen.
Thanks to Jorashii, Chris, Cyclcrclicly, anvuew and Bas, Sahlofolina.
Press information:
https://twitter.com/AppAdsai/status/1675692821603549187/photo/1
https://techcrunch.com/2023/06/30/tiktok-parent-bytedance-launches-music-creation-audio-editing-app/
Site:
BS-RoFormer
Used architecture in Capcut/Ripple (now defunct). Their paper was published and later reimplemented by lucidrains for training and inferencing:
https://github.com/lucidrains/BS-RoFormer
Later, Mel-Band RoFormer based on band split was released, which is faster, but doesn't provide such high SDR as BS. Mel variant might require some revision of the code, and its paper might lack some features need to keep up SDR-wise with extremely slow BS original variant. On paper, it should be better than BS-Roformer, but for some reason, models trained with Mel have worse results than with BS-Roformer (so probably problem with reimplementation from paper). Kim reworked her config, so the results with Mel models improved, but still are a tad lower than BS-Roformer. ZFTurbo includes training and inference of Roformers in his repository on GitHub.
For more information, check the training section.
About ByteDance
Winners of MDX23 competition. They said at the beginning, that it utilizes novel arch (so no weighting/ensembling of existing models). In times of v.0.1 seemingly the best vocals, not so good instrumentals, as it was once said by someone who heard samples, but they came a long way lately. It's all about their politics. It's a Chinese company responsible for TikTok, famous for d**k moves outside China - manipulating their algorithms - encourage of stupidity outside China, and greedy, wellness-centered attitudes for users in China (the app is currently banned in China), manipulating their algorithms to promote only black-white relationships in western countries, spying on users copying their clipboard, spying even on journalists to find their sources of information about the company, unauthorized remote access to TikTok user data from China, and also, a subject to ban in US and other countries for bad influence on children, data infringement by storing non-China users data directly on their servers which is against the law of many countries (there were some actions taken on it later). Decompiling TikTok analysis (tons of spying improper behavior of the app). Currently, Bytedance is only around 40% owned by founders, Chinese investors, and their employees and the rest (60%) state global investors (incl. lots of American) and is pushed to sale more stakes to US risking US ban on the app.
They said, the CEO, told them to hold this ByteDance arch for two years for themselves. Initially they had plans to release it in some kind of app, firstly at the end of June, later something was planned at the end of year, later they said something about two years (maybe more about open sourcing, but we can't have our hopes high). Previously, they said the case of open sourcing/releasing was stuck in their legal department. Later they told they used MUSDBHQ+500 songs for their dataset. These 500 songs could have been obtained illegally for training (although everyone does it), but they might be extremely cautious about it (or it's just an excuse). Eventually, they released Ripple and Capcut. Then they released the paper for the Bs/Mel archs, and it was implemented and coded by Lucidrains, so later could be used for training.
Later, they seemingly spread information among users privately, that despite the similarities in SDR, the 18.75 score is a result of a trolling, someone other than ByteDance. Some people favoring ByteDance were rumored for disruptive, trolling behavior on our server too, harassing other users, or just being unkind to others etc. Besides, the same person responsible, was also the most informed about ByteDance next moves, and was also changing nicknames or accounts frequently. Also possessed great ML knowledge. Many coincidences. In the end, the same user, zmis (if you see the details of the account above), was behind a lot of newly created, accounts, which were banned on our server.
The same day or in very similar period, a new account was created, conducting the same behavior, when previous was banned.
The main core of their activities, was spreading misinformation about SDR metrics, telling that is the most important thing in the world, because their own arch is good at it, hence the narration.
So don't bother, and do your good job not feeding troll from other company. They don't like competition, doing their own moves behind backdrops and become better.
It’s not impossible to fake SDR results in the MVSEP leaderboard. For current public archs, you’d need to feed your dataset both by the songs in the evaluation dataset, keeping your regular big dataset in place, so you simply lose evaluation factor of this leaderboard, or you can simply mix your result stems with original stems. “SDR focuses more on lower frequencies, it can easily be fooled into giving a higher score if the lower frequencies are louder, Bas tested this theory and confirmed it”. you can boost the bass, and it will score +1 or +2 sdr higher or something, that's why It's not always reliable” - becruily
Those results, which are not faked, are at least those, which were uploaded by various users evaluating the same public, available for offline use models, but usually uploaded with various parameters which affects SDR (so usually the better parameters, the higher SDR, but not always), remain consistent among various users evaluations with similar parameters and inference code, so scalability is correct and preserved, thus the results weren’t faked, and can be reproduced with similar SDR. For the other scores from unpublic inferences/models/methods, we simply trust ZFTurbo and rather viperx too, as they’re/were our trusted users for years. Also, the leaderboard in the current multisong dataset tends to give better SDR to the results with more residues on different occasions before, so the chart is simply not fully reliable for that, but rather not manipulated in its core either. It’s more a nature of SDR measurement and/or used dataset.
ViperX trained the first community BS-Roformer model similar to SAMI v1.0 model, although 2 stems (and lower scoring Mel at the time). His BS model sounds similar to Ripple (although it's only 2 stem, while 4 stem Ripple variant scores a bit higher than the 2 stem variant, but still lower than ViperX and v1.0). Then there were a lot of community trained models like private one by ZFTurbo, and fine-tunes by various users (Kim, Unwa, Bas, ZFTurbo)
Bas tried to train a model purely on multisong dataset only, but failed to surpass the SDR score of a 1.1 Bytedance’s model anyway. v1.1 has new arch enhancements to the arch, and will be presented on ISMIR2024 (white paper is already out; link in the Training section).