https://arxiv.org/abs/2508.14623
Accepted for IEEE ASRU, 2025
by Simon Dahl Jepsen, Mads Græsbøll Christensen, and Jesper Rindom Jensen
Welcome to this interactive listening demo, which complements the research described in the above paper. On this page, you will find side-by-side audio excerpts processed by the three different models, trained under different conditions, so that you can hear the effect of training-data design on separation outcome.
This demo lets you listen to the outputs of the three speech-separation models evaluated in the paper. They differ only in the type of target (reference) signals used during training:
- Baseline model – trained on the standard mixtures from the wsj0-2mix [1,2] corpus, where the target references contain some residual noise.
- Enhanced-reference model – trained on the same data, but with denoised references to remove that residual noise.
- Enhanced-reference + noise-augmented model – trained on denoised references and mixtures with extra added noise to improve robustness in noisy conditions.
For each selected mixture (from the test sets described in the paper: e.g., WSJ0‑2Mix and Libri2Mix) we provide three listening tracks:
- Mixture (two speakers)
- Reference (target
- Output of Model 1 (baseline)
- Output of Model 2 (enhanced references)
- Output of Model 3 (enhanced + noise augmentation)
Use your headphones for better comparisons!
WSJ0-2Mix samples
Two examples are provided from this test set.
Example 1:
Example 2:
| [1] Garofolo, John S., et al. CSR-I (WSJ0) Complete LDC93S6A. Web Download. Philadelphia: Linguistic Data Consortium, 1993. [2] Hershey, J. R., Chen, Z., Le Roux, J., & Watanabe, S. (2016, March). Deep clustering: Discriminative embeddings for segmentation and separation. In 2016 IEEE international conference on acoustics, speech and signal processing (ICASSP) (pp. 31-35). IEEE. |