arXiv:2510.04157cs.SDeess.AS2025-10

用噪声模型引导扩散过程,提升语音增强对未知噪声的适应能力。

GDiffuSE: Diffusion-based speech enhancement with noise model guidance

  • 通过轻量辅助模型估计噪声分布,引导扩散去噪过程
  • 在不匹配噪声下仍优于现有最优方法,提升显著
  • 适合需要强鲁棒性的语音增强场景

本文提出一种基于去噪扩散概率模型(DDPM)的新型语音增强方法,称为GDiffuSE。与传统直接映射噪声语音到干净语音的方法不同,该方法利用轻量级辅助模型估计噪声分布,并通过引导机制将其融入扩散去噪过程。这一设计提升了对未见噪声类型的鲁棒性,并可借助原本为语音生成训练的大规模DDPM模型实现语音增强。我们在将BBC音效库噪声添加到LibriSpeech语音数据上的测试中评估了该方法,在噪声类型不匹配条件下均表现出持续改进,优于当前最先进的基线方法。示例见项目网页。

原文摘要 · Abstract (English)

This paper introduces a novel speech enhancement (SE) approach based on a denoising diffusion probabilistic model (DDPM), termed Guided diffusion for speech enhancement (GDiffuSE). In contrast to conventional methods that directly map noisy speech to clean speech, our method employs a lightweight helper model to estimate the noise distribution, which is then incorporated into the diffusion denoising process via a guidance mechanism. This design improves robustness by enabling seamless adaptation to unseen noise types and by leveraging large-scale DDPMs originally trained for speech generation in the context of SE. We evaluate our approach on noisy signals obtained by adding noise samples from the BBC sound effects database to LibriSpeech utterances, showing consistent improvements over state-of-the-art baselines under mismatched noise conditions. Examples are available at our project webpage.

语音增强扩散模型噪声建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。