Make any audio trainable.
The Refinery turns raw, mixed recordings into clean, training-ready data — diagnosed, separated, scored, never synthetic. Bring us the audio you've licensed, acquired, or created, and we return isolated, confidence-scored stems your models can actually learn from.
One engine. Any audio in, training-ready data out.
Most audio a lab licenses and acquires is a black box — multi-party, uneven, and difficult to trust at scale. The Refinery tells you what's in it, what's good, and then makes it usable.
See inside the black box
Point us at your corpus. We assess every hour for what actually limits training — background music, speaker overlap, noise, degradation, bleed — and determine which audio separation models will deliver the best results.
Isolate the real signal
Source separation pulls clean stems out of finished, mixed recordings — dialogue, per-speaker tracks, music-free speech — from a single track, no session files needed.
Put a number on quality
Every output ships with confidence scores, so you filter, weight, and QC your dataset programmatically instead of listening your way through it.
Ready to train
Structured, training-ready data delivered over API or inside your own environment when the audio can't leave your walls.
Your audio, refined.
Refine cleans and scores what you bring. Send us the recordings you've licensed, acquired, or created — even if you don't yet know what's usable. We diagnose the whole corpus, then return clean, isolated, confidence-scored stems, ready to train.
Diagnose what's usable, fixable, or junk before you spend compute on it
Works from finished, mixed recordings — no original stems required
Music, noise, and bleed removed; speakers split into individual streams
It's your real audio, isolated — never synthesized or hallucinated
Turn licensed archives into training data
Turn the licensed archives you already own into training-ready data for ASR, diarization, speaker-ID, and TTS. Diagnose quality before you spend compute and labeling budget, and separate everything you need from finished recordings — so your existing corpus becomes a data source.
Upgrade raw inventory into sellable datasets
Upgrade raw audio inventory into isolated, scored, training-ready datasets you can sell. Refine adds the separation and quality layer buyers trust — turning mixed recordings into structured, per-source data with confidence scores attached.
Usable, structured data.
Send us finished, mixed recordings — no original stems or multitrack sessions required. We return training-ready components, each isolated from the real signal.
Spoken dialogue pulled out of real-world audio — crowd environments, on-location recordings, mixed broadcast content. AudioShake's various speech isolation models work across a range of resolution and content types.
Multi-speaker conversation – including overlapping speech and filler words – split into individual speaker tracks.
Isolate 10+ individual instruments — vocals, bass, drums, guitar, lead and backing vocals, and more — from fully mixed tracks.
All background music removed — including sung vocals — leaving a pure speaker signal where music would otherwise contaminate the training set.
Know the quality before you train.
Every multi-speaker output ships with confidence scores — signals indicating how cleanly a source is isolated across the track. Filter, weight, or QC your dataset on numbers, not subjective listening.
Watch how confidence scores work →Drop the 0.62 stem, or route it to human QC — before it reaches your pipeline.
We collect audio in the real world, which means real-world noise and interference. AudioShake has helped us process thousands of hours of clean, speaker-separated data, enabling the world's leading labs to build better models.
Built for the way labs actually train.
Diagnose before you spend
We assess your corpus upfront to determine the separation models that will get the most out of it — so compute and labeling budget go to audio that will actually improve training, not audio that was never going to help.
From a single mixed track
No original stems, no multitrack sessions, no isolated recordings required. We separate finished audio using our state-of-the-art models — which means your existing archive is already a data source.
Never generative
Every stem is real audio, isolated — not synthesized, not hallucinated, not filled in. What you train on is exactly what was recorded.
Consistent and repeatable
The same input yields the same output at scale, so no distribution shift creeps in from inconsistent preprocessing — the failure mode that destabilizes benchmarks.
Cloud or on-prem
Run on our API, or deploy inside your own environment when the data can't leave your walls.
Cleared to train
Your audio never becomes someone else's dataset. We process what you bring and return it to you — and in on-prem deployments, it never leaves your environment. Provenance you control.
Frequently Asked Questions
Structured, training-ready components — isolated stems and per-speaker tracks, each with confidence scores — delivered in our cloud or your environment.
Yes. Marketplaces use Refinery to turn raw audio inventory into isolated, scored, training-ready datasets with a quality layer buyers can trust.
Yes. The Refinery can be deployed on-prem so audio never leaves your walls.
You do. The Refinery processes the audio you bring and returns it to you; we don't repurpose your data and all data is deleted upon completion of the job. The Refinery can also be deployed on-prem.
No. AudioShake's Refinery works from finished, mixed recordings — a single track is enough.
Mixed audio entangles signals — overlapping speakers, music, and noise — so a model can't cleanly learn any one of them. Separated, per-source stems give your model the isolated signal it's actually meant to learn.
Yes. The Refinery is built for corpus-scale throughput; we've processed over one billion minutes of audio to date.
That's what the Diagnose step is for. We assess your whole corpus and flag what's usable, fixable, or junk before you commit compute or budget.