arXiv:2509.14379eess.AScs.LG2025-09被引 2

通过建模噪声分布,实现单麦克风下语音分离

Diffusion-Based Unsupervised Audio-Visual Speech Separation in Noisy Environments with Noise Prior

  • 用生成模型同时建模语音和结构化噪声,不依赖混合信号训练
  • 利用视觉线索作为语音先验,提升噪声环境下的分离效果
  • 适合在嘈杂环境中做无监督语音分离的研究者参考

本文针对存在环境噪声时的单麦克风语音分离问题,提出一种生成式无监督方法,直接建模干净语音与结构化噪声成分,仅使用纯净语音和噪声信号进行训练,而非混合信号。该方法基于音视频得分模型,引入视觉线索作为强先验,有效增强语音建模能力。通过显式建模噪声分布,结合反向扩散过程采样后验分布,直接估计并去除噪声成分以恢复清晰语音。实验表明,该方法在复杂声学环境下表现优异,验证了直接建模噪声的有效性。

原文摘要 · Abstract (English)

In this paper, we address the problem of single-microphone speech separation in the presence of ambient noise. We propose a generative unsupervised technique that directly models both clean speech and structured noise components, training exclusively on these individual signals rather than noisy mixtures. Our approach leverages an audio-visual score model that incorporates visual cues to serve as a strong generative speech prior. By explicitly modelling the noise distribution alongside the speech distribution, we enable effective decomposition through the inverse problem paradigm. We perform speech separation by sampling from the posterior distributions via a reverse diffusion process, which directly estimates and removes the modelled noise component to recover clean constituent signals. Experimental results demonstrate promising performance, highlighting the effectiveness of our direct noise modelling approach in challenging acoustic environments.

语音分离扩散模型无监督学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。