arXiv:2506.01460cs.SDeess.AS2025-06中稿 · Interspeech 2025被引 7

用GAN改进扩散模型,1步就能高质量语音去噪

Few-step Adversarial Schrödinger Bridge for Generative Speech Enhancement

  • 将GAN与薛定谔桥结合,加速生成过程
  • 单步推理即超越50步传统方法,在低信噪比下仍稳定
  • 适合实时语音增强场景,兼顾速度与音质

深度生成模型被用于语音增强,以在大规模数据集上生成感知真实的纯净语音。已有扩散模型和可计算的薛定谔桥被提出,用于在干净与噪声语音分布间传输。但这些模型通常依赖迭代反向过程,需超过50次采样步骤。我们发现,当采样步数减少时,基线模型性能显著下降,尤其在低信噪比条件下。为此,我们提出将薛定谔桥与GAN结合,有效缓解此问题,在全频带数据集上实现高质量输出,同时大幅减少采样步数。实验表明,所提模型即使仅用一步推理,也在去噪和去混响任务中优于现有基线。

原文摘要 · Abstract (English)

Deep generative models have recently been employed for speech enhancement to generate perceptually valid clean speech on large-scale datasets. Several diffusion models have been proposed, and more recently, a tractable Schrödinger Bridge has been introduced to transport between the clean and noisy speech distributions. However, these models often suffer from an iterative reverse process and require a large number of sampling steps -- more than 50. Our investigation reveals that the performance of baseline models significantly degrades when the number of sampling steps is reduced, particularly under low-SNR conditions. We propose integrating Schrödinger Bridge with GANs to effectively mitigate this issue, achieving high-quality outputs on full-band datasets while substantially reducing the required sampling steps. Experimental results demonstrate that our proposed model outperforms existing baselines, even with a single inference step, in both denoising and dereverberation tasks.

语音增强生成模型扩散模型快速生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。