arXiv:2505.23515eess.AScs.LG2025-05中稿 · Interspeech 2025被引 6

用生成式方法实时增强语音,减少失真并保留语音细节。

DeepFilterGAN: A Full-band Real-time Speech Enhancement System with GAN-based Stochastic Regeneration

  • 基于生成对抗网络的随机重构框架,学习语音完整分布
  • 358万参数,低延迟,支持全频段实时流式处理
  • 在语音质量评测中优于传统模型,适合实际语音应用

本文提出一种基于生成对抗网络的全频段实时语音增强系统。预测模型仅估计目标分布的均值,易导致语音内容过度抑制;而生成模型可学习完整分布,降低输出失真。通过在随机重构框架中结合二者,我们构建了轻量级实时系统,仅含358万参数,延迟极低,适用于实时语音流处理。实验表明,该系统在NISQA-MOS评分上优于第一阶段基线模型。消融实验进一步验证了噪声条件输入对性能的关键作用。我们基于此模型参加了2025年紧急挑战赛,并持续优化。

原文摘要 · Abstract (English)

In this work, we propose a full-band real-time speech enhancement system with GAN-based stochastic regeneration. Predictive models focus on estimating the mean of the target distribution, whereas generative models aim to learn the full distribution. This behavior of predictive models may lead to over-suppression, i.e. the removal of speech content. In the literature, it was shown that combining a predictive model with a generative one within the stochastic regeneration framework can reduce the distortion in the output. We use this framework to obtain a real-time speech enhancement system. With 3.58M parameters and a low latency, our system is designed for real-time streaming with a lightweight architecture. Experiments show that our system improves over the first stage in terms of NISQA-MOS metric. Finally, through an ablation study, we show the importance of noisy conditioning in our system. We participated in 2025 Urgent Challenge with our model and later made further improvements.

语音增强GAN实时系统生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。