arXiv:2409.17898eess.AScs.SD2024-09被引 2

将SEMamba扩展到多麦克风场景,性能优于或媲美现有模型。

MC-SEMamba: A Simple Multi-channel Extension of SEMamba

  • 设计轻量级多通道扩展结构,仅小幅增加参数量。
  • 在CHiME3数据集上表现超越多个基线模型,6麦克风时效果更优。
  • 适合需要高精度语音增强的多麦克风系统部署。

基于Transformer的模型因其在序列建模中的卓越表现,在语音处理研究中日益流行。最近,Mamba作为一种有前景的替代架构出现,因其对长序列的高效建模能力而备受关注。特别是SEMamba已在单通道语音增强中展现出Mamba架构的有效性。本文旨在仅以少量参数增加,将SEMamba适配至多通道应用。所提出的MC-SEMamba系统在CHiME3数据集上的表现与多个先前基线模型相当甚至更优。此外,我们发现麦克风数量从1个增至6个时,MC-SEMamba的语音增强性能显著提升。

原文摘要 · Abstract (English)

Transformer-based models have become increasingly popular and have impacted speech-processing research owing to their exceptional performance in sequence modeling. Recently, a promising model architecture, Mamba, has emerged as a potential alternative to transformer-based models because of its efficient modeling of long sequences. In particular, models like SEMamba have demonstrated the effectiveness of the Mamba architecture in single-channel speech enhancement. This paper aims to adapt SEMamba for multi-channel applications with only a small increase in parameters. The resulting system, MC-SEMamba, achieved results on the CHiME3 dataset that were comparable or even superior to several previous baseline models. Additionally, we found that increasing the number of microphones from 1 to 6 improved the speech enhancement performance of MC-SEMamba.

语音增强Mamba多通道轻量模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。