arXiv:2512.07226eess.AScs.SD2025-12AAAI被引 4

无需配对数据,用扩散模型分离单通道音频

Unsupervised Single-Channel Audio Separation with Diffusion Source Priors

  • 将分离问题视为概率反问题,仅需独立声源的扩散先验
  • 新设计的逆问题求解器有效缓解梯度冲突,提升分离质量
  • 适合语音、声音事件等真实场景下的无监督音频分离

单通道音频分离旨在从单一混合信号中分离出各个声源。现有方法多依赖合成配对数据的监督学习,但在真实场景中高质量配对数据难以获取,导致模型在未见条件下性能下降,泛化能力受限。为此,本文从无监督视角出发,将问题建模为概率反问题。方法仅需在独立声源上训练的扩散先验,通过迭代引导初始状态向解空间逼近实现分离。关键创新在于设计专用逆问题求解器,缓解扩散先验与重建引导在去噪过程中的梯度冲突,确保各声源分离质量均衡。此外,采用增强混合信号作为去噪初始化,显著提升最终性能。为进一步增强音频建模能力,提出一种基于时频注意力的新网络架构。实验表明,该方法在语音-声音事件、声音事件及语音分离任务中均取得显著提升。

原文摘要 · Abstract (English)

Single-channel audio separation aims to separate individual sources from a single-channel mixture. Most existing methods rely on supervised learning with synthetically generated paired data. However, obtaining high-quality paired data in real-world scenarios is often difficult. This data scarcity can degrade model performance under unseen conditions and limit generalization ability. To this end, in this work, we approach this problem from an unsupervised perspective, framing it as a probabilistic inverse problem. Our method requires only diffusion priors trained on individual sources. Separation is then achieved by iteratively guiding an initial state toward the solution through reconstruction guidance. Importantly, we introduce an advanced inverse problem solver specifically designed for separation, which mitigates gradient conflicts caused by interference between the diffusion prior and reconstruction guidance during inverse denoising. This design ensures high-quality and balanced separation performance across individual sources. Additionally, we find that initializing the denoising process with an augmented mixture instead of pure Gaussian noise provides an informative starting point that significantly improves the final performance. To further enhance audio prior modeling, we design a novel time-frequency attention-based network architecture that demonstrates strong audio modeling capability. Collectively, these improvements lead to significant performance gains, as validated across speech-sound event, sound event, and speech separation tasks.

音频分离扩散模型无监督学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。