基于漂移模型的语音增强,单步完成去噪,效果超越多步扩散模型。
Speech Enhancement Based on Drifting Models

- 通过演化映射分布匹配纯净语音分布,实现单步推理。
- 在VoiceBank-DEMAND上达到高保真增强,优于多步扩散基线。
- 无需成对数据,可直接学习分布匹配,适合实际应用。
我们提出一种基于漂移模型的语音增强框架(DriftSE),将去噪问题建模为平衡问题。与依赖迭代采样的方法不同,DriftSE通过演化映射函数的前向分布,直接匹配纯净语音分布,实现原生单步推理。该演化由一个学习到的漂移场驱动,该场作为校正向量,引导样本向纯净语音分布的高密度区域移动,从而自然地支持无配对数据训练,仅需匹配分布而非成对样本。我们在两种形式下研究该框架:一种是直接从含噪观测映射,另一种是从高斯先验出发的随机条件生成模型。在VoiceBank-DEMAND基准上的实验表明,DriftSE可在单步内实现高保真增强,超越多步扩散基线,确立了语音增强的新范式。
原文摘要 · Abstract (English)
We propose Speech Enhancement based on Drifting Models (DriftSE), a novel generative framework that formulates denoising as an equilibrium problem. Rather than relying on iterative sampling, DriftSE natively achieves one-step inference by evolving the pushforward distribution of a mapping function to directly match the clean speech distribution. This evolution is driven by a Drifting Field, a learned correction vector that guides samples toward the high-density regions of the clean distribution, which naturally facilitates training on unpaired data by matching distributions rather than paired samples. We investigate the framework under two formulations: a direct mapping from the noisy observation, and a stochastic conditional generative model from a Gaussian prior. Experiments on the VoiceBank-DEMAND benchmark demonstrate that DriftSE achieves high-fidelity enhancement in a single step, outperforming multi-step diffusion baselines and establishing a new paradigm for speech enhancement.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。