arXiv:2508.14556cs.SDcs.AI2025-08被引 1

用Mamba2提升稀疏人声分离效果,更准更稳。

Mamba2 Meets Silence: Robust Vocal Source Separation for Sparse Regions

  • 采用Mamba2捕捉长时序依赖,适合间歇性人声。
  • 在cSDR达11.03 dB,刷新当前最优纪录。
  • 对不同输入长度和人声模式都表现稳定,适合实际场景。

我们提出一种针对精准人声分离的新模型。与依赖Transformer的方法不同,该模型采用近期的态空间模型Mamba2,更有效地捕捉长时间依赖关系。为高效处理长序列输入,结合频带分割策略与双路径架构。实验表明,该方法超越现有最先进模型,在cSDR上达到11.03 dB,创历史新高,并显著提升uSDR。此外,模型在不同输入长度和人声出现模式下均表现出稳定一致的性能。结果验证了Mamba类模型在高分辨率音频处理中的有效性,为音频研究拓展了新方向。

原文摘要 · Abstract (English)

We introduce a new music source separation model tailored for accurate vocal isolation. Unlike Transformer-based approaches, which often fail to capture intermittently occurring vocals, our model leverages Mamba2, a recent state space model, to better capture long-range temporal dependencies. To handle long input sequences efficiently, we combine a band-splitting strategy with a dual-path architecture. Experiments show that our approach outperforms recent state-of-the-art models, achieving a cSDR of 11.03 dB-the best reported to date-and delivering substantial gains in uSDR. Moreover, the model exhibits stable and consistent performance across varying input lengths and vocal occurrence patterns. These results demonstrate the effectiveness of Mamba-based models for high-resolution audio processing and open up new directions for broader applications in audio research.

人声分离Mamba2音频处理长时序建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。