arXiv:2605.08189eess.AS2026-05中稿 · Interspeech 2026

首个可复现的扩散模型语音增强系统,性能超越现有方法。

DiffVQE: Hybrid Diffusion Voice Quality Enhancement Under Acoustic Echo and Noise

论文配图:DiffVQE: Hybrid Diffusion Voice Quality Enhancement Under Acoustic Echo and Noise
图 1 · 摘自论文原文
  • 采用扩散模型框架实现端到端回声与噪声抑制
  • 在回声和噪声控制上优于DeepVQE,且更轻量高效
  • 基于真实数据集构建,具备可复现性,适合语音处理研究者

免提系统和音箱中的声学回声与背景噪声给语音增强带来挑战。判别式端到端方法在联合回声消除(AEC)与降噪方面表现优异。然而,随着生成式方法的发展,基于扩散的方法在语音增强任务中展现出显著性能。本文首次提出一个可复现的非因果扩散型回声消除模型DiffVQE,其拓扑结构、训练数据和训练框架均公开。使用Interspeech 2025 URGENT挑战赛的数据构建高质量训练集,DiffVQE在回声与噪声抑制性能上均优于微软的判别式DeepVQE模型,同时在计算复杂度和模型尺寸上也更优。

原文摘要 · Abstract (English)

Acoustic echo and background noise pose challenges on speech enhancement in hands-free systems and speakerphones. Discriminatively trained end-to-end methods represent a powerful solution for joint acoustic echo control (AEC) and denoising. However, with the advent of generative methods, diffusion-based approaches have seen remarkable performance in speech enhancement tasks. In this work, to the best of our knowledge, we provide the first (still non-causal) diffusion-based AEC model (DiffVQE) that is reproducible in terms of topology, training data, and training framework. So far, without employing diffusion, Microsoft's discriminative DeepVQE model has been shown to excel any of the ICASSP 2023 AEC Challenge entries achieving remarkable performance. Using data from the Interspeech 2025 URGENT Challenge for a diverse, high-quality training dataset, our DiffVQE excels DeepVQE both in echo and noise control performance, as well as in computational complexity and model size.

语音增强扩散模型回声消除

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。