arXiv:2511.20470cs.SDcs.AI2025-11中稿 · oral presentation …被引 2

用潜在扩散模型实现高效歌声分离,仅需音轨与混音配对数据训练。

Efficient and Fast Generative-Based Singing Voice Separation using a Latent Diffusion Model

  • 基于潜在扩散模型,在压缩隐空间生成音频并解码,提升推理速度。
  • 仅用开源数据训练,性能超越现有生成式分离方法,媲美非生成模型。
  • 提供噪声鲁棒性分析,适合音乐制作与语音分离研究者使用。

从音乐混音中提取单个声部是音乐制作与练习的重要工具。尽管神经网络通过掩码或变换频谱图来分离源信号已成为主流方法,但音乐信号中的源重叠与相关性带来固有挑战。此外,训练这些系统需获取所有源信号,过程复杂。虽已有生成式方法尝试解决此问题,但分离性能与推理效率仍受限。本文研究扩散模型在生成式歌声分离中的潜力,专注于仅用孤立人声与混音配对数据进行训练。为契合创作流程,采用潜在扩散:系统在紧凑隐空间生成样本,并解码为音频,实现高效优化与快速推理。模型仅使用开放数据训练,性能超越现有生成式分离系统,在信号质量指标与干扰消除方面达到非生成方法水平。我们还对潜在编码器的噪声鲁棒性进行了研究,揭示其在任务中的潜力,并发布模块化工具包以促进后续研究。

原文摘要 · Abstract (English)

Extracting individual elements from music mixtures is a valuable tool for music production and practice. While neural networks optimized to mask or transform mixture spectrograms into the individual source(s) have been the leading approach, the source overlap and correlation in music signals poses an inherent challenge. Also, accessing all sources in the mixture is crucial to train these systems, while complicated. Attempts to address these challenges in a generative fashion exist, however, the separation performance and inference efficiency remain limited. In this work, we study the potential of diffusion models to advance toward bridging this gap, focusing on generative singing voice separation relying only on corresponding pairs of isolated vocals and mixtures for training. To align with creative workflows, we leverage latent diffusion: the system generates samples encoded in a compact latent space, and subsequently decodes these into audio. This enables efficient optimization and faster inference. Our system is trained using only open data. We outperform existing generative separation systems, and level the compared non-generative systems on a list of signal quality measures and on interference removal. We provide a noise robustness study on the latent encoder, providing insights on its potential for the task. We release a modular toolkit for further research on the topic.

歌声分离扩散模型潜在空间音频生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。