arXiv:2411.06217eess.AS2024-11被引 15

用混合卷积与Mamba结构提升语音增强的长序列处理能力

Selective State Space Model for Monaural Speech Enhancement

  • 结合卷积局部建模与Mamba全局建模,设计新骨干网络MambaDC
  • 在多个任务上超越Transformer、Conformer和标准Mamba,性能更优
  • 适合需要高效处理长语音序列的实时语音增强系统

语音用户界面(VUI)通过语音指令实现了人机高效交互。真实场景中声学环境复杂,语音增强对保障VUI鲁棒性至关重要。尽管基于Transformer及其变体(如Conformer)在语音增强中表现优异,但其计算复杂度随序列长度呈二次增长,难以处理长序列。最近提出的状态空间模型Mamba以线性复杂度有效建模长序列,为该问题提供新解。本文提出一种新型混合卷积-Mamba骨干网络MambaDC,融合卷积网络对局部特征的捕捉能力与Mamba对长程全局依赖的建模优势。我们在基础与当前最先进的语音增强框架下,针对两种常用训练目标进行了全面实验。结果表明,MambaDC在所有训练目标上均优于Transformer、Conformer及标准Mamba。基于现有先进框架,使用MambaDC骨干网络的系统性能显著超越现有最先进(SoTA)方法,为语音增强中的高效长程建模奠定了基础。

原文摘要 · Abstract (English)

Voice user interfaces (VUIs) have facilitated the efficient interactions between humans and machines through spoken commands. Since real-word acoustic scenes are complex, speech enhancement plays a critical role for robust VUI. Transformer and its variants, such as Conformer, have demonstrated cutting-edge results in speech enhancement. However, both of them suffers from the quadratic computational complexity with respect to the sequence length, which hampers their ability to handle long sequences. Recently a novel State Space Model called Mamba, which shows strong capability to handle long sequences with linear complexity, offers a solution to address this challenge. In this paper, we propose a novel hybrid convolution-Mamba backbone, denoted as MambaDC, for speech enhancement. Our MambaDC marries the benefits of convolutional networks to model the local interactions and Mamba's ability for modeling long-range global dependencies. We conduct comprehensive experiments within both basic and state-of-the-art (SoTA) speech enhancement frameworks, on two commonly used training targets. The results demonstrate that MambaDC outperforms Transformer, Conformer, and the standard Mamba across all training targets. Built upon the current advanced framework, the use of MambaDC backbone showcases superior results compared to existing \textcolor{black}{SoTA} systems. This sets the stage for efficient long-range global modeling in speech enhancement.

语音增强Mamba长序列建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。