arXiv:2509.22063cs.CVcs.SD2025-09IJCV被引 5

用视觉引导生成模型实现跨类别高质量声音分离

High-Quality Sound Separation Across Diverse Categories via Visually-Guided Generative Modeling

  • 基于扩散模型和流匹配的生成式框架,直接从噪声中合成目标声谱
  • 在AVE和MUSIC数据集上优于现有方法,尤其擅长多样声音类别分离
  • 适合关注音频视觉联合建模与高质量声音分离的研究者

我们提出DAVIS,一种基于扩散模型的音视频分离框架,通过生成学习解决音视频声音源分离任务。现有方法通常将声音分离视为基于掩码的回归问题,虽取得显著进展,但在捕捉多样化声音类别所需的复杂数据分布方面存在局限。相比之下,DAVIS利用强大的生成建模范式——去噪扩散概率模型(DDPM)和近期的流匹配(FM),结合专用分离U-Net架构,通过同时条件于混合音频输入和关联视觉信息,直接从噪声分布合成期望的分离声谱。其生成目标的本质使DAVIS特别擅长处理多样化声音类别的高质量分离。我们在标准AVE和MUSIC数据集上对DAVIS及其DDPM与流匹配变体进行了对比评估,结果表明两种变体均超越现有方法,在分离质量上表现优异,验证了该生成框架在音视频源分离任务中的有效性。

原文摘要 · Abstract (English)

We propose DAVIS, a Diffusion-based Audio-VIsual Separation framework that solves the audio-visual sound source separation task through generative learning. Existing methods typically frame sound separation as a mask-based regression problem, achieving significant progress. However, they face limitations in capturing the complex data distribution required for high-quality separation of sounds from diverse categories. In contrast, DAVIS circumvents these issues by leveraging potent generative modeling paradigms, specifically Denoising Diffusion Probabilistic Models (DDPM) and the more recent Flow Matching (FM), integrated within a specialized Separation U-Net architecture. Our framework operates by synthesizing the desired separated sound spectrograms directly from a noise distribution, conditioned concurrently on the mixed audio input and associated visual information. The inherent nature of its generative objective makes DAVIS particularly adept at producing high-quality sound separations for diverse sound categories. We present comparative evaluations of DAVIS, encompassing both its DDPM and Flow Matching variants, against leading methods on the standard AVE and MUSIC datasets. The results affirm that both variants surpass existing approaches in separation quality, highlighting the efficacy of our generative framework for tackling the audio-visual source separation task.

音视频分离扩散模型生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。