解决歌声混叠分离难题,提升复杂声线分离效果
UNMIXX: Untangling Highly Correlated Singing Voices Mixtures
- 用音乐化混音策略生成逼真训练数据
- 通过反向注意力机制分离高度相关声线
- 结合幅度惩罚损失,显著提升分离精度
我们提出 UNMIXX,一种用于多声部歌声分离(MSVS)的新框架。相较于语音分离,MSVS面临数据稀缺与歌声高度相关两大挑战。为此,UNMIXX引入三个关键组件:(1) 音乐引导的混音策略,构建高度相关且类音乐的混合信号;(2) 跨源注意力机制,通过反向注意力促使两位歌手表征分离;(3) 幅度惩罚损失,对误分配的干扰能量进行抑制。UNMIXX不仅通过模拟真实训练数据缓解数据稀缺问题,更在架构与损失层面实现跨源交互,有效分离高度相关混合信号。大量实验表明,其性能显著优于现有方法,SDRi 提升超过 2.2 dB。
原文摘要 · Abstract (English)
We introduce UNMIXX, a novel framework for multiple singing voices separation (MSVS). While related to speech separation, MSVS faces unique challenges: data scarcity and the highly correlated nature of singing voices mixture. To address these issues, we propose UNMIXX with three key components: (1) musically informed mixing strategy to construct highly correlated, music-like mixtures, (2) cross-source attention that drives representations of two singers apart via reverse attention, and (3) magnitude penalty loss penalizing erroneously assigned interfering energy. UNMIXX not only addresses data scarcity by simulating realistic training data, but also excels at separating highly correlated mixtures through cross-source interactions at both the architectural and loss levels. Our extensive experiments demonstrate that UNMIXX greatly enhances performance, with SDRi gains exceeding 2.2 dB over prior work.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。