无需训练即可提升音频分离效果,通过迭代优化融合实现
Training-Free Multi-Step Audio Source Separation
- 迭代融合输入与上步结果,动态调整混合比例以优化分离
- 多步推理在语音增强和音乐分离任务中均超越单步,性能接近大模型
- 理论证明方法持续改进,适合已有模型快速升级使用
音频源分离旨在将混合信号分解为目标声源。以往系统多采用单步推理,未能充分发挥模型潜力。本文揭示预训练的单步分离模型可无需额外训练即用于多步分离。提出一种简单有效的方法:通过最优融合输入混合信号与前步分离结果进行迭代分离,每步通过最大化某指标确定最佳融合比例。理论上证明该方法始终优于单步推理,并给出基于模型平滑性与指标鲁棒性的误差界。方法与沿噪声到纯净分布线性插值路径上的去噪性质相关,这一特性与去噪扩散桥模型相联系。实证表明,该多步分离方法在语音增强与音乐源分离任务中持续优于单步推理,性能接近训练更大模型、使用更多数据或采用多步训练目标的效果。改进不仅体现在优化指标上,还扩展至几乎所有非优化指标(仅一例外)。最后讨论了方法局限与未来方向。
原文摘要 · Abstract (English)
Audio source separation aims to separate a mixture into target sources. Previous audio source separation systems usually conduct one-step inference, which does not fully explore the separation ability of models. In this work, we reveal that pretrained one-step audio source separation models can be leveraged for multi-step separation without additional training. We propose a simple yet effective inference method that iteratively applies separation by optimally blending the input mixture with the previous step's separation result. At each step, we determine the optimal blending ratio by maximizing a metric. We prove that our method always yield improvement over one-step inference, provide error bounds based on model smoothness and metric robustness, and provide theoretical analysis connecting our method to denoising along linear interpolation paths between noise and clean distributions, a property we link to denoising diffusion bridge models. Our approach effectively delivers improved separation performance as a "free lunch" from existing models. Our empirical results demonstrate that our multi-step separation approach consistently outperforms one-step inference across both speech enhancement and music source separation tasks, and can achieve scaling performance similar to training a larger model, using more data, or in some cases employing a multi-step training objective. These improvements appear not only on the optimization metric during multi-step inference, but also extend to nearly all non-optimized metrics (with one exception). We also discuss limitations of our approach and directions for future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。