通过随机排列麦克风通道提升分离模型抗串音能力
Learning Input-Channel Permutation Equivariance for Multi-Channel Source Separation: Reducing Bleeding in Small Music Ensembles

- 训练时对输入通道和目标信号做相同随机排列,强制模型不依赖固定通道顺序
- 在未见过的混响环境和真实录音上,音质指标SDR提升且串音显著减少
- 适合音乐制作中麦克风布局多变的场景,尤其小乐队现场录制
麦克风串音是小型乐团和管弦乐录音中的长期难题,近距离麦克风虽针对单个乐器,却会拾取邻近声源的泄漏。这种重叠降低音轨分离效果并增加混音难度。本文通过将通道排列等变性作为核心学习原则来解决该问题:训练时对输入麦克风通道及其对应参考信号施加相同的随机排列,从而避免模型依赖固定的通道-乐器对应关系,提升对录音布置甚至乐器变化的鲁棒性。所提模型在包含多样化模拟混响与麦克风位置的合成合奏数据上训练,并在未见的模拟条件及真实URMP录音上评估。结果表明,相比非排列感知基线,该方法在未见条件下始终提升信号失真比(SDR)并减少串音。研究证实,排列等变性是一种简单、以数据为中心的策略,可有效实现音乐制作中的鲁棒去串音与多通道声源分离。
原文摘要 · Abstract (English)
Microphone bleed is a persistent challenge in small ensembles and orchestral recordings, where close microphones intended for individual instruments also capture leakage from nearby sources. This overlap degrades track isolation and complicates mixing. This paper addresses the bleeding problem by making channel-permutation-equivariance a core learning principle. During training, we apply the same random permutation to the input microphone channels and their corresponding reference targets. This discourages reliance on fixed channel-instrument associations and improves robustness to changes in the recording setup and even in the recorded instruments. The proposed model is trained on synthetic ensembles with diverse simulated room acoustics and microphone placements, and evaluated on unseen simulated conditions and real URMP recordings. The results show that permutation-aware training consistently improves SDR and reduces bleeding under unseen conditions compared with non-permutation baselines. The findings highlight permutation-equivariance as a simple, data-centric strategy for robust debleeding and practical multi-channel source separation in music production workflows.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。