arXiv:2409.10376eess.AScs.SD2024-09中稿 · ICASSP 2025被引 4

用Mamba改进多通道语音增强,同时建模频谱与空间信息。

Leveraging Joint Spectral and Spatial Learning with MAMBA for Multichannel Speech Enhancement

  • 引入Mamba架构,融合全带与窄带频谱及空间特征
  • 在CHiME-3数据集上优于McNet,达当前最优性能
  • 特别擅长捕捉频谱动态变化,适合复杂噪声环境

在多通道语音增强中,有效捕捉不同麦克风间的空间与频谱信息对降噪至关重要。传统方法如CNN或LSTM试图建模全带与子带频谱及空间特征的时序动态,但在动态声学环境下难以充分建模复杂的时序依赖关系。为此,本文在先进模型McNet基础上,引入改进版Mamba(一种状态空间模型),并提出MCMamba。该模型重新设计,整合全带与窄带空间信息,结合子带与全带频谱特征,实现更全面的空间与频谱建模。实验表明,MCMamba显著提升了多通道语音增强中空间与频谱特征的建模能力,在CHiME-3数据集上超越McNet,达到当前最优性能。此外,研究发现Mamba在频谱建模方面表现尤为突出。

原文摘要 · Abstract (English)

In multichannel speech enhancement, effectively capturing spatial and spectral information across different microphones is crucial for noise reduction. Traditional methods, such as CNN or LSTM, attempt to model the temporal dynamics of full-band and sub-band spectral and spatial features. However, these approaches face limitations in fully modeling complex temporal dependencies, especially in dynamic acoustic environments. To overcome these challenges, we modify the current advanced model McNet by introducing an improved version of Mamba, a state-space model, and further propose MCMamba. MCMamba has been completely reengineered to integrate full-band and narrow-band spatial information with sub-band and full-band spectral features, providing a more comprehensive approach to modeling spatial and spectral information. Our experimental results demonstrate that MCMamba significantly improves the modeling of spatial and spectral features in multichannel speech enhancement, outperforming McNet and achieving state-of-the-art performance on the CHiME-3 dataset. Additionally, we find that Mamba performs exceptionally well in modeling spectral information.

语音增强Mamba多通道频谱建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。