arXiv:2509.26580cs.SDcs.LG2025-09

解决无伴奏合唱中歌手分离难题,提升复杂场景下的分离效果。

Source Separation for A Cappella Music

  • 用幂集增强数据,指数级扩充多歌手训练样本。
  • 在JaCappella上达到顶尖性能,支持不同人数混合的泛化分离。
  • 适配周期激活与复合损失,静音声部也能准确分离。

本文研究无伴奏合唱音乐中的多歌手分离任务,其中混合音轨中活跃歌手数量不固定。为应对这一挑战,我们采用基于幂集的数据增强策略,将有限的多歌手数据集扩展为指数级更多的训练样本。为实现歌手分离,提出SepACap,这是对当前先进说话人分离模型SepReformer的改进版本。通过引入周期性激活机制和一种在声部静音时仍有效的复合损失函数,模型具备更强的鲁棒性。在JaCappella数据集上的实验表明,该方法在全乐团及子集歌手分离场景下均达到最先进水平,优于基于频谱图的基线模型,并能泛化到具有不同歌手数量的真实混合音频中。

原文摘要 · Abstract (English)

In this work, we study the task of multi-singer separation in a cappella music, where the number of active singers varies across mixtures. To address this, we use a power set-based data augmentation strategy that expands limited multi-singer datasets into exponentially more training samples. To separate singers, we introduce SepACap, an adaptation of SepReformer, a state-of-the-art speaker separation model architecture. We adapt the model with periodic activations and a composite loss function that remains effective when stems are silent, enabling robust detection and separation. Experiments on the JaCappella dataset demonstrate that our approach achieves state-of-the-art performance in both full-ensemble and subset singer separation scenarios, outperforming spectrogram-based baselines while generalizing to realistic mixtures with varying numbers of singers.

语音分离无伴奏合唱数据增强深度学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。