arXiv:2507.02391cs.SDcs.LG2025-07被引 2

用扩散模型直接建模语音增强的后验转移,无需调参且更鲁棒。

Posterior Transition Modeling for Unsupervised Diffusion-Based Speech Enhancement

  • 直接建模噪声语音到干净语音的扩散过程转移分布。
  • 在WSJ0-QUT和VoiceBank-DEMAND上提升信噪比,优于有监督与无监督基线。
  • 适合追求高鲁棒性、不依赖标注数据的语音增强研究者。

我们探索使用扩散模型作为干净语音的表达性生成先验,实现无监督语音增强。现有方法通过近似噪声扰动的似然得分引导反向扩散过程,并与无条件得分通过超参数权衡结合。本文提出两种新算法:第一种以合理方式融合扩散先验与观测模型,避免超参数调优;第二种在噪声语音上定义扩散过程,得到完全可解析的精确似然得分。在WSJ0-QUT和VoiceBank-DEMAND数据集上的实验表明,该方法在增强指标上表现更优,且对域偏移更具鲁棒性,优于既有监督与无监督基线。

原文摘要 · Abstract (English)

We explore unsupervised speech enhancement using diffusion models as expressive generative priors for clean speech. Existing approaches guide the reverse diffusion process using noisy speech through an approximate, noise-perturbed likelihood score, combined with the unconditional score via a trade-off hyperparameter. In this work, we propose two alternative algorithms that directly model the conditional reverse transition distribution of diffusion states. The first method integrates the diffusion prior with the observation model in a principled way, removing the need for hyperparameter tuning. The second defines a diffusion process over the noisy speech itself, yielding a fully tractable and exact likelihood score. Experiments on the WSJ0-QUT and VoiceBank-DEMAND datasets demonstrate improved enhancement metrics and greater robustness to domain shifts compared to both supervised and unsupervised baselines.

语音增强扩散模型无监督学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。