AudioShake vs.
Demucs, BS-Roformer & MVSep

Sound Separation, Speech Separation, & Transcription
11.7
8.8
5.9
AudioShake
Demucs v4
Spleeter

Music separation vs. Demucs & BS-Roformer

AudioShake leads the field in high-quality source separation, which is typically measured via something called the Signal-to-Distortion (SDR) score. AudioShake has repeatedly demonstrated its ability to achieve the highest SDR scores and set several state of the art benchmarks.

But SDR scores alone are not a good way to measure source separation–it's quite possible to achieve a high SDR score while producing output that doesn't sound very good. That's why we focus even more on perceptual quality. In Meta's SAM Audio evaluation, AudioShake was the highest-performing discriminative model on perceptual quality, ahead of Moises / Music.ai, FADR, Lalal and MossFormer. AudioLabs Erlangen reached the same conclusion on piano concertos, finding AudioShake produced the highest-quality piano and orchestral stems.

11.7
Quality or SDR Score, averaged over four stems
15
Number of Separations — against 6 and 4
SDR Scores

The music information retrieval community typically uses something called the Signal-to-Distortion (SDR) score to measure quality, and AudioShake has repeatedly demonstrated its ability to achieve the highest SDR scores. However, Even though we hold many of the state-of-the-art benchmarks, we caution people from solely using SDR score as the best indicator of a high-quality model, because it is quite possible to achieve a high SDR score while the actual results sound poor.

That's why we have developed perceptual metrics that we use for evaluating model performance, and sometimes choose lower-SDR scores models that perform better on different tasks.

With all that said, below is a sample of our stem separation scores for music.

Separations available
15
AudioShake
6
Demucs v4
4
Spleeter
Stem separation scores
SDR, higher is better
AudioShake
Demucs v4
Spleeter
0
7
14
Average
11.7
8.8
5.9
Vocals
13.5
8.9
6.9
Drums
12.1
10.0
6.7
Bass
11.9
9.8
5.5
Other
9.2
6.4
4.6

Hear the difference: vocal isolation compared

Models from 10x-400x faster than real-time
Process hours-long files
Process high-resolution (192kHz) files
Stem
Bass
Drums
Other
Vocals
AudioShake
Solo
SAM Audio
Solo
Demucs v4
Solo
Spleeter
Solo
Source
“Future” by Torches

Speech separation: Multi-Speaker 2.0 vs. open-source baselines

The comparisons above cover music. Speech is measured separately: Multi-Speaker 2.0 against Multi-Speaker 1.0 and open-source baselines on separation, diarization, and downstream transcription, with the datasets and methodology written out.

32%
better separation between voices than Multi-Speaker 1.0 on broadcast-style speech (LRS2-2Mix)
75%
more noise and interference removed than the 1.0 pipeline (Libri2Mix + WHAM!)
76%
fewer word errors than TF-Locoformer on LibriCSS, transcribed with Whisper large-v3

Lyric transcription accuracy (WER)

To evaluate transcription models, researchers use industry-wide benchmarks that measure word error rate (WER) alongside punctuation and formatting metrics. Alignment is evaluated separately, measuring how closely predicted word and line timings match the audio. At ISMIR, AudioShake’s research team presented a new benchmark based on the JamendoLyrics dataset that captures the finer nuances of written lyrics, called Jam-ALT.

Below are the results of running various transcription and alignment systems against these benchmarks. AudioShake sets the state-of-the-art in both categories — including 95% accuracy on lyric alignment — making it the most accurate system available for transcription and for the word-level timing that synced lyrics, karaoke, and captioning depend on.

16.1
Word Error Rate across English, Spanish, German and French
95%
accuracy on lyric alignment
Word Error Rate
English, Spanish, German and French — lower is better
Scale to 70
AudioShake
16.1
Whisper v3 (OpenAI)
33.5
OWSM v3.1
66.5
Get in touch.
Thank you! Your submission has been received!
Oops! Something went wrong while submitting the form.