用生成先验增强语音增强与分离,提升音质且通用性强。
Predictive-Generative Drift Decomposition for Speech Enhancement and Separation

- 通过随机插值分解动态,将预测结果融入生成采样过程。
- 在语音分离任务中实现+1.0 NISQA的感知质量提升。
- 兼容多种预测器和降噪场景,适合跨任务应用。
我们提出一种即插即用的语音增强与分离框架,通过引入生成式语音先验来增强预测方法。该方法称为语音随机插值先验(SIPS),基于随机插值技术,灵活融合预测与生成建模。具体而言,将插值动态分解为任务相关的漂移项与随机去噪分量,使预测结果可直接融入生成采样过程。由此构建一个数学严谨的框架,结合强预训练预测器与生成模型的表达能力。我们仅使用纯净语音训练得分模型,获得对退化无关的先验,可跨任务复用。推理时,预测器提供确定性漂移以引导采样至任务一致估计,而得分模型保持听觉自然性。与以往依赖特定架构或固定预测器的混合方法不同,SIPS提供统一框架,适用于多种预测器及加性退化任务。我们在SEMamba和FlexIO等近期预测器上验证其有效性,持续提升感知质量,在语音分离任务中实现+1.0 NISQA增益。
原文摘要 · Abstract (English)
We propose a plug-and-play framework for speech enhancement and separation that augments predictive methods with a generative speech prior. Our approach, termed Stochastic Interpolant Prior for Speech (SIPS), builds on stochastic interpolants and leverages their flexibility to bridge predictive and generative modeling. Specifically, we decompose the interpolation dynamics into a task-specific drift and a stochastic denoising component, allowing a predictive estimate to be integrated directly into the generative sampling process. This results in a mathematically grounded framework for combining strong pretrained predictors with the expressive power of generative models. To this end, we train a score model using only clean speech, yielding a degradation-agnostic prior that can be reused across tasks. During inference, the predictor provides a deterministic drift that steers the sampling process toward a task-consistent estimate, while the score model preserves perceptual naturalness. Unlike prior hybrid approaches, which typically rely on architecture-specific conditioning and are tied to particular predictors or degradation settings, SIPS provides a unified framework that generalizes across predictors and additive degradation tasks. We demonstrate its effectiveness for both speech enhancement and speech separation using recent predictors such as SEMamba and FlexIO. The proposed method consistently improves perceptual quality, achieving gains up +1.0 NISQA for speech separation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。