用GAN改进扩散模型,1步就能高质量语音去噪
Few-step Adversarial Schrödinger Bridge for Generative Speech Enhancement
- 将GAN与薛定谔桥结合,加速生成过程
- 单步推理即超越50步传统方法,在低信噪比下仍稳定
- 适合实时语音增强场景,兼顾速度与音质
深度生成模型被用于语音增强,以在大规模数据集上生成感知真实的纯净语音。已有扩散模型和可计算的薛定谔桥被提出,用于在干净与噪声语音分布间传输。但这些模型通常依赖迭代反向过程,需超过50次采样步骤。我们发现,当采样步数减少时,基线模型性能显著下降,尤其在低信噪比条件下。为此,我们提出将薛定谔桥与GAN结合,有效缓解此问题,在全频带数据集上实现高质量输出,同时大幅减少采样步数。实验表明,所提模型即使仅用一步推理,也在去噪和去混响任务中优于现有基线。
原文摘要 · Abstract (English)
Deep generative models have recently been employed for speech enhancement to generate perceptually valid clean speech on large-scale datasets. Several diffusion models have been proposed, and more recently, a tractable Schrödinger Bridge has been introduced to transport between the clean and noisy speech distributions. However, these models often suffer from an iterative reverse process and require a large number of sampling steps -- more than 50. Our investigation reveals that the performance of baseline models significantly degrades when the number of sampling steps is reduced, particularly under low-SNR conditions. We propose integrating Schrödinger Bridge with GANs to effectively mitigate this issue, achieving high-quality outputs on full-band datasets while substantially reducing the required sampling steps. Experimental results demonstrate that our proposed model outperforms existing baselines, even with a single inference step, in both denoising and dereverberation tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。