将生成模型与预测模型统一,提升语音增强效率与效果
Rethinking Flow and Diffusion Bridge Models for Speech Enhancement
- 将流模型与扩散桥模型统一为可变均值方差的高斯路径构建
- 优化后模型参数更少、计算量更低,性能超越现有基线
- 适合关注高效语音增强的工程师与研究者
流匹配与扩散桥模型已成为生成式语音增强的主流范式,基于流匹配、得分匹配和薛定谔桥等原理,建模噪声与干净语音之间的随机过程。本文提出一个统一框架,将现有模型解释为在配对数据间构建具有可变均值和方差的高斯概率路径。进一步分析表明,经过数据预测损失优化的生成模型,在每一步采样中理论上等同于执行预测式语音增强。受此启发,我们设计了一种增强型桥模型,融合有效概率路径设计与预测范式的关键元素:改进网络结构、定制化损失函数及优化训练策略。在降噪与去混响任务上的实验表明,该方法在参数更少、计算复杂度更低的前提下,显著优于现有流与扩散基线。结果也揭示了该生成框架固有的预测性质对其可达到性能上限的限制。
原文摘要 · Abstract (English)
Flow matching and diffusion bridge models have emerged as leading paradigms in generative speech enhancement, modeling stochastic processes between paired noisy and clean speech signals based on principles such as flow matching, score matching, and Schrödinger bridge. In this paper, we present a framework that unifies existing flow and diffusion bridge models by interpreting them as constructions of Gaussian probability paths with varying means and variances between paired data. Furthermore, we investigate the underlying consistency between the training/inference procedures of these generative models and conventional predictive models. Our analysis reveals that each sampling step of a well-trained flow or diffusion bridge model optimized with a data prediction loss is theoretically analogous to executing predictive speech enhancement. Motivated by this insight, we introduce an enhanced bridge model that integrates an effective probability path design with key elements from predictive paradigms, including improved network architecture, tailored loss functions, and optimized training strategies. Experiments on denoising and dereverberation tasks demonstrate that the proposed method outperforms existing flow and diffusion baselines with fewer parameters and reduced computational complexity. The results also highlight that the inherently predictive nature of this generative framework imposes limitations on its achievable upper-bound performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。