用Mamba2提升稀疏人声分离效果,更准更稳。
Mamba2 Meets Silence: Robust Vocal Source Separation for Sparse Regions
- 采用Mamba2捕捉长时序依赖,适合间歇性人声。
- 在cSDR达11.03 dB,刷新当前最优纪录。
- 对不同输入长度和人声模式都表现稳定,适合实际场景。
我们提出一种针对精准人声分离的新模型。与依赖Transformer的方法不同,该模型采用近期的态空间模型Mamba2,更有效地捕捉长时间依赖关系。为高效处理长序列输入,结合频带分割策略与双路径架构。实验表明,该方法超越现有最先进模型,在cSDR上达到11.03 dB,创历史新高,并显著提升uSDR。此外,模型在不同输入长度和人声出现模式下均表现出稳定一致的性能。结果验证了Mamba类模型在高分辨率音频处理中的有效性,为音频研究拓展了新方向。
原文摘要 · Abstract (English)
We introduce a new music source separation model tailored for accurate vocal isolation. Unlike Transformer-based approaches, which often fail to capture intermittently occurring vocals, our model leverages Mamba2, a recent state space model, to better capture long-range temporal dependencies. To handle long input sequences efficiently, we combine a band-splitting strategy with a dual-path architecture. Experiments show that our approach outperforms recent state-of-the-art models, achieving a cSDR of 11.03 dB-the best reported to date-and delivering substantial gains in uSDR. Moreover, the model exhibits stable and consistent performance across varying input lengths and vocal occurrence patterns. These results demonstrate the effectiveness of Mamba-based models for high-resolution audio processing and open up new directions for broader applications in audio research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。