arXiv:2503.06375eess.AS2025-03NAACL被引 6

用扩散模型生成语音增强的先验信息,提升效果与效率

ProSE: Diffusion Priors for Speech Enhancement

  • 在低维隐空间用扩散模型生成语音增强先验
  • 结合变换器回归模型,实现更清晰、低失真的增强结果
  • 计算量更低,适合实时语音增强场景

语音增强(SE)是提升噪声环境下语音清晰度和质量的基础任务。尽管传统深度学习模型常用于此,但近期研究表明,去噪扩散概率模型(DDPM)等生成模型具有潜力。然而,与语音生成不同,语音增强需严格匹配真实语音信号,且多数应用场景要求实时处理,而传统扩散模型推理时需大量迭代且计算开销大。为此,我们提出ProSE(基于扩散先验的语音增强),一种新方法:首先在低维隐空间利用DDPM生成先验分布,再将这些先验信息融入基于变换器的回归模型进行语音增强。由于扩散过程在紧凑的隐空间进行,所需迭代次数少于传统扩散模型,同时回归结构避免了扩散模型生成细节错位导致的失真问题。实验表明,ProSE在多个基准数据集上达到领先性能,且计算成本更低。

原文摘要 · Abstract (English)

Speech enhancement (SE) is the foundational task of enhancing the clarity and quality of speech in the presence of non-stationary additive noise. While deterministic deep learning models have been commonly employed for SE, recent research indicates that generative models, such as denoising diffusion probabilistic models (DDPMs), have shown promise. However, unlike speech generation, SE has a strong constraint in generating results in accordance with the underlying ground-truth signal. Additionally, for a wide variety of applications, SE systems need to be employed in real-time, and traditional diffusion models (DMs) requiring many iterations of a large model during inference are inefficient. To address these issues, we propose ProSE (diffusion-based Priors for SE), a novel methodology based on an alternative framework for applying diffusion models to SE. Specifically, we first apply DDPMs to generate priors in a latent space due to their powerful distribution mapping capabilities. The priors are then integrated into a transformer-based regression model for SE. The priors guide the regression model in the enhancement process. Since the diffusion process is applied to a compact latent space, the diffusion model takes fewer iterations than the traditional DM to obtain accurate estimations. Additionally, using a regression model for SE avoids the distortion issue caused by misaligned details generated by DMs. Our experiments show that ProSE achieves state-of-the-art performance on benchmark datasets with fewer computational costs.

语音增强扩散模型隐空间实时处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。