arXiv:2410.20997cs.SDcs.LG2024-10被引 11

用Mamba提升语音分离效率,计算成本更低

SepMamba: State-space models for speaker separation using Mamba

  • 基于双向Mamba层构建U-Net结构,替代传统Transformer
  • 在WSJ0双说话人数据集上性能优于同类模型,推理更快
  • 适合资源受限场景下的实时语音分离应用

近年来,基于深度学习的单通道语音分离得益于Transformer注意力机制取得显著进展。然而,这些方法计算开销巨大,难以在实际应用中部署。作为计算高效的替代方案,Mamba近期被提出。本文提出SepMamba,一种以双向Mamba层为核心的U-Net架构。实验表明,该方法在WSJ0 2-speaker数据集上表现优于同等规模的主流模型(包括Transformer),同时显著降低计算成本、内存占用和前向传播时间。此外,因果变体的SepMamba也表现出色。本工作为深度语音分离提供了更高效的架构选择。

原文摘要 · Abstract (English)

Deep learning-based single-channel speaker separation has improved significantly in recent years largely due to the introduction of the transformer-based attention mechanism. However, these improvements come at the expense of intense computational demands, precluding their use in many practical applications. As a computationally efficient alternative with similar modeling capabilities, Mamba was recently introduced. We propose SepMamba, a U-Net-based architecture composed primarily of bidirectional Mamba layers. We find that our approach outperforms similarly-sized prominent models - including transformer-based models - on the WSJ0 2-speaker dataset while enjoying a significant reduction in computational cost, memory usage, and forward pass time. We additionally report strong results for causal variants of SepMamba. Our approach provides a computationally favorable alternative to transformer-based architectures for deep speech separation.

语音分离Mamba高效模型深度学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。