arXiv:2606.05575cs.SDeess.AS2026-06

一拍即合:用量子桥接加速语音增强,一步搞定

SB-RF: Schrödinger Bridge Rectified Flow for One-Step Robust Speech Enhancement

  • 结合量子桥接与修正流理论,实现单步推理
  • 在低信噪比下表现优于现有生成模型
  • 适合对实时性要求高的语音增强场景

生成模型在语音增强(SE)中表现优异,但通常依赖多步推断,限制了低延迟部署。本文提出SB-RF,一种融合修正流(RF)与量子桥接(SB)理论的一步生成框架。训练时,从SB时间边缘采样中间状态,并使用RF的速度匹配目标训练条件速度场;推理时,从含噪观测出发,仅需一次欧拉更新即可完成增强。实验表明,SB-RF在VoiceBank-DEMAND基准上达到与主流生成方法相当的性能。为进一步评估其在极端条件下的表现,我们在扩展训练数据基础上,针对模拟低信噪比测试集进行了评估,结果表明,SB-RF在该条件下显著优于对比基线,验证了其在真实场景中的应用潜力。

原文摘要 · Abstract (English)

Generative models have shown promising results for speech enhancement (SE), but they often rely on multi-step inference, limiting low-latency deployment. We propose SB-RF, a one-step generative framework that integrates Rectified Flow (RF) with Schrödinger Bridge (SB) theory. During training, SB-RF samples intermediate states from an SB time marginal and trains a conditional velocity field with the RF velocity-matching objective. At inference, SB-RF starts from the noisy observation and applies a single Euler update. Experiments show that SB-RF achieves competitive performance among generative methods on the VoiceBank-DEMAND benchmark. To further assess performance beyond this standard setting, we evaluate SB-RF on a simulated low signal-to-noise ratio test set using an expanded training dataset. Under these conditions, SB-RF achieves superior performance over the compared baselines, supporting its potential for real-world applications.

语音增强生成模型单步推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。