arXiv:2511.21342cs.SDcs.AI2025-11中稿 · publication at WAS…被引 2

用扩散模型从混音中分离人声,可调节生成质量与效率。

Generating Separated Singing Vocals Using a Diffusion Model Conditioned on Music Mixtures

  • 以混音为条件,用扩散模型生成独立人声
  • 在补充数据下达到与非生成方法相当的客观评分
  • 支持用户调节采样参数,灵活控制输出质量

音乐混合中分离各声部是音乐分析与实践的重要步骤。传统方法通常使用神经网络对时频表示进行掩码或变换以提取目标声源,而生成式扩散模型因其灵活性与泛化能力,正成为解决此难题的新范式。本文探索使用扩散模型从真实音乐录音中分离人声,该模型在给定混音条件下生成独唱人声。我们的方法优于先前的生成系统,在引入额外训练数据后,其客观评估指标达到与非生成基线相当的水平。扩散采样的迭代特性使用户可灵活调节生成质量与效率之间的权衡,并在需要时对输出进行优化。我们还进行了采样算法的消融研究,揭示了用户可配置参数的影响。

原文摘要 · Abstract (English)

Separating the individual elements in a musical mixture is an essential process for music analysis and practice. While this is generally addressed using neural networks optimized to mask or transform the time-frequency representation of a mixture to extract the target sources, the flexibility and generalization capabilities of generative diffusion models are giving rise to a novel class of solutions for this complicated task. In this work, we explore singing voice separation from real music recordings using a diffusion model which is trained to generate the solo vocals conditioned on the corresponding mixture. Our approach improves upon prior generative systems and achieves competitive objective scores against non-generative baselines when trained with supplementary data. The iterative nature of diffusion sampling enables the user to control the quality-efficiency trade-off, and also refine the output when needed. We present an ablation study of the sampling algorithm, highlighting the effects of the user-configurable parameters.

语音分离扩散模型音乐生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。