arXiv:2412.18217cs.SDcs.LG2024-12被引 8

轻量级Mamba-U-Net模型,高效分离嘈杂混响语音

U-Mamba-Net: A highly efficient Mamba-based U-net style network for noisy and reverberant speech separation

  • 用Mamba作特征筛选,与U-Net交替构建轻量网络
  • 在Libri2mix上性能优于传统模型,计算量仅1.2GFLOPs
  • 适合资源受限场景下的实时语音分离应用

语音分离旨在将多说话人重叠的混合语音分解为各自独立的语音流。尽管已有诸多高效模型涌现,但其规模与计算开销也显著增加,给复现和对比带来巨大负担。本文提出U-Mamba-Net:一种基于Mamba的轻量级U-Net风格模型,用于复杂环境下的语音分离。Mamba作为状态空间序列模型,具备特征选择能力;U-Net结构通过对称的收缩与扩张路径学习多分辨率特征。本工作将Mamba作为特征过滤器,与U-Net交替使用。在Libri2mix数据集上的实验表明,U-Mamba-Net在保持低计算成本(1.2 GFLOPs)的同时实现了优异性能。

原文摘要 · Abstract (English)

The topic of speech separation involves separating mixed speech with multiple overlapping speakers into several streams, with each stream containing speech from only one speaker. Many highly effective models have emerged and proliferated rapidly over time. However, the size and computational load of these models have also increased accordingly. This is a disaster for the community, as researchers need more time and computational resources to reproduce and compare existing models. In this paper, we propose U-mamba-net: a lightweight Mamba-based U-style model for speech separation in complex environments. Mamba is a state space sequence model that incorporates feature selection capabilities. U-style network is a fully convolutional neural network whose symmetric contracting and expansive paths are able to learn multi-resolution features. In our work, Mamba serves as a feature filter, alternating with U-Net. We test the proposed model on Libri2mix. The results show that U-Mamba-Net achieves improved performance with quite low computational cost.

语音分离轻量模型Mamba

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。