arXiv:2510.16834cs.SDcs.AI2025-10中稿 · Interspeech 2026被引 1

用一步推理实现高质量语音增强,兼顾速度与效果。

Schrödinger Bridge Mamba for One-Step Speech Enhancement

  • 结合斯格明格桥训练与Mamba架构,实现高效语音增强。
  • 单步推理下多项指标超越主流方法,实时因子表现优异。
  • 适合需要低延迟、高保真语音处理的实时场景应用。

我们提出一种名为斯格明格桥Mamba(SBM)的新模型,通过融合斯格明格桥(SB)训练范式与Mamba架构,实现高效的语音增强。在联合去噪与去混响任务中,SBM仅需一步推理即可在多个指标上超越强大多模态生成与判别方法,同时具备具有竞争力的实时因子,满足流式语音处理需求。消融实验表明,相比传统映射方式,SB范式在多种架构下均能稳定提升性能;此外,在SB范式下,Mamba的表现优于多头自注意力(MHSA)和长短期记忆(LSTM)等骨干网络。这些发现凸显了Mamba架构与基于轨迹的SB训练之间的协同效应,为实际语音增强提供了高质量解决方案。演示页面:https://sbmse.github.io

原文摘要 · Abstract (English)

We present Schrödinger Bridge Mamba (SBM), a novel model for efficient speech enhancement by integrating the Schrödinger Bridge (SB) training paradigm and the Mamba architecture. Experiments of joint denoising and dereverberation tasks demonstrate SBM outperforms strong generative and discriminative methods on multiple metrics with only one step of inference while achieving a competitive real-time factor for streaming feasibility. Ablation studies reveal that the SB paradigm consistently yields improved performance across diverse architectures over conventional mapping. Furthermore, Mamba exhibits a stronger performance under the SB paradigm compared to Multi-Head Self-Attention (MHSA) and Long Short-Term Memory (LSTM) backbones. These findings highlight the synergy between the Mamba architecture and the SB trajectory-based training, providing a high-quality solution for real-world speech enhancement. Demo page: https://sbmse.github.io

语音增强Mamba扩散模型实时处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。