Separate overlapping speakers into clean voices

The world's first high-fidelity, multi-speaker voice technology. Multi-Speaker Separation pairs high-accuracy diarization with industry-leading sound isolation, turning overlapping voices into clean, individual stems with confidence scores built in.

Available via AudioShake Live and via our API.
PLAY
00:00
SPEAKER 1
SPEAKER 2

What is Multi-Speaker Separation?

From podcast interviews and reality TV, to phone calls, interviews, and field recordings—much of the world’s audio is filled with conversational audio and overlapping speakers.

Multi-Speaker Separation takes a fully mixed track and pulls each person onto their own isolated, clean stem.
Built for overlapping speech
Speaker diarization
Confidence scores

What makes AudioShake's Multi-Speaker Separation different

01
We're probably better than the diarizer you're using
30% lower diarization error than pyannote community-1, and lower error than the paid service we replaced in every domain we test — meetings, telephone, broadcast and web video.
02
A diarizer can't give you the overlapped voices
Labels can mark that two people spoke at once. They can't recover the two voices as separate audio. Multi-speaker separation does — so each interjection is preserved.
03
AudioShake also delivers confidence scores to make review faster
The separated audio comes back scored per frame and per file, so you can process the confident material automatically and route the rest to a human. More on confidence scores.
01

Measurably better separation quality

Our latest Multi-Speaker model combines high-accuracy diarization with AudioShake's sound isolation. Here's what that changes, measured against Version 1.0 and the strongest open-source systems available.
75%
more noise and interference removed than Multi-Speaker 1.0
32%
less bleed between voices than Multi-Speaker 1.0
fewer transcription errors than the open-source separator
30%
lower diarization error than pyannote community-1

The more people talk over each other, the more it matters

Send a conversation straight to a transcription tool and the errors climb as people interrupt one another. Run Multi-Speaker 2.0 first and that climb flattens — at the heaviest overlap, half the errors.
Errors per 100 words010203010%20%30%40%How much of the recording is people talking at onceNo separation27.8Multi-Speaker 1.018.8Multi-Speaker 2.014.1
Transcription errors on LibriCSS, a standard set of recorded group conversations. Same recordings, same transcription tool, in every case.
SEE THE FULL BENCHMARKS
02

Speaker separation built for real-world audio

Recordings don't arrive clean. AudioShake's Multi-Speaker Separation holds up across the range of audio you actually work with — different resolutions, channel formats, and messy acoustic conditions.

Sample rates

8 kHz – 48 kHz

From compressed phone audio to full studio capture.

Channels

Mono & stereo

Native support for both, from single-mic captures to stereo masters.

LANGUAGES

Language agnostic

Acoustic-models not language-models, trained on how voices sound vs. what they're saying.

Metadata

Confidence scores

Per-segment certainty you can act on at scale.

03

You get the labels, not just the audio.

Alongside the separated tracks you get a structured results file — an overlap-aware timeline of who spoke when. Meaning when two people speak at the same moment, both speakers are active in the timeline instead of one of them being dropped.

That's what lets the output go straight into captioning, subtitling, dubbing and edit prep, where losing an interjection is the failure mode you can least afford.

Example: One mixed recording in, two labeled tracks out
Speaker 1
Speaker 2
Both talking at once
Input
OUTPUT
Speaker 1
Speaker 2
both active
both active
Confidence
0:00
0:15
0:30
0:45
1:00

What teams do with the output

01
Send a reviewer to the areas that need an ear
The timeline and the per-frame scores point at exact timestamps, so a QC pass is minutes on the moments that matter instead of on the whole file.
02
Process the confident material automatically, route the rest
File-level scores rank a batch, so the clean majority moves through the pipeline untouched and only the uncertain files queue for attention.
03
Filter a training corpus by score instead of by ear
The same scores work as a dataset filter, which is what makes it possible to keep real conversation — interruptions, laughter, backchannels — instead of discarding it as unusable.
04

Built for high-stakes audio workflows

01
Film, TV & post-production
Isolate overlapping dialogue for the edit
Pull apart crosstalk and overlapping dialogue into separate speaker tracks for cleaner editing, ADR, and re-mixing. Recover a single line buried under a crowd, or lift one actor's voice out of a busy scene without touching the rest of the mix.
02
Podcasting
A clean, separate stem for every guest
Turn multi-mic or single-track recordings into per-speaker stems for editing, leveling, and cleanup. Fix one guest's audio, remove crosstalk between hosts, or balance a remote guest against the room — without re-recording.
03
Dubbing & localization
A clean stem for every speaker to dub
Get an isolated stem per speaker across the whole recording, so localization teams can re-voice and translate speaker by speaker instead of wrestling with a mixed track. Clean separation makes each voice easier to replace, sync, and balance for global distribution.
04
Transcription & ASR
Cleaner input, better machine intelligibility
Feed cleaner per-speaker audio into speech-to-text so a single-speaker ASR system can attribute words to the right person. At up to 20% crosstalk, our separated tracks transcribe as accurately as individually mic'd speakers. Use the built-in confidence scores to route high-confidence segments straight through and flag the rest for review.
05
VOICE AI & TRAINING DATA
Real conversations, turned into structured data
Turn real recordings into speaker-isolated stems with turn boundaries and confidence scores attached. Filter a corpus by confidence instead of by ear, and keep the overlapped regions that diarization alone forces you to discard.
06
ACCESSIBILITY
Isolate and preserve an individual voice
Extract a single, clean voice from a noisy multi-speaker recording — useful for captioning, assistive workflows, and voice-banking where fidelity to the original speaker matters.
05

How to use AudioShake's Multi-Speaker Separation

01
Upload your recording
Upload or stream your recording as it is. Anything from 8 to 48 kHz — no pre-processing or format conversion required.
02
Run Multi-Speaker
Select Multi-Speaker Separation to access individual recordings per speaker in any piece of audio.
03
Download separated audio
Export the stems and metadata, or send them into your editing, transcription, or data pipeline.
WEB PLATFORM
Upload and process recordings directly
AudioShake Live is an intuitive, drag-and-drop web platform designed for companies, film studios, and media production teams to create high-quality stems on demand.
REQUEST ACCESS
API
Integrate Multi-Speaker Separation into your workflow
Connect Multi-Speaker Separation to your own product or pipeline, with diarization and confidence scores included.
DEVELOPER DOCS

Processing audio at scale?

Separate your audio, then filter by confidence to keep only the cleanest outputs — a programmatic way for labs and enterprises to QA large volumes and build reliable datasets.

Explore Data Services
06

Frequently Asked Questions

No items found.
Get in touch.
Thank you! Your submission has been received!
Oops! Something went wrong while submitting the form.