对比四种模型在单声道语音增强中的表现,Mamba效率最优。
Long-Context Modeling Networks for Monaural Speech Enhancement: A Comparative Study
- 统一框架下比较Transformer、Conformer、Mamba和xLSTM
- Mamba在长语音上训练推理速度最快,性能优于前两者
- xLSTM虽性能好但处理最慢,适合对速度不敏感场景
先进的长上下文建模骨干网络如Transformer、Conformer和Mamba已在语音增强中达到顶尖水平。然而,在统一语音增强框架内对这些骨干网络进行系统性、全面的比较研究仍显不足。此外,一种较新的高效LSTM变体xLSTM在语言建模和通用视觉骨干中已展现出良好效果。本文研究了xLSTM在语音增强中的能力,并在统一框架下对Transformer、Conformer、Mamba和xLSTM骨干网络进行了全面比较与分析,涵盖因果与非因果配置。总体而言,xLSTM和Mamba的性能优于Transformer和Conformer。Mamba在长语音输入下表现出显著更优的训练与推理效率,而xLSTM则存在最慢的处理速度。
原文摘要 · Abstract (English)
Advanced long-context modeling backbone networks, such as Transformer, Conformer, and Mamba, have demonstrated state-of-the-art performance in speech enhancement. However, a systematic and comprehensive comparative study of these backbones within a unified speech enhancement framework remains lacking. In addition, xLSTM, a more recent and efficient variant of LSTM, has shown promising results in language modeling and as a general-purpose vision backbone. In this paper, we investigate the capability of xLSTM in speech enhancement, and conduct a comprehensive comparison and analysis of the Transformer, Conformer, Mamba, and xLSTM backbones within a unified framework, considering both causal and noncausal configurations. Overall, xLSTM and Mamba achieve better performance than Transformer and Conformer. Mamba demonstrates significantly superior training and inference efficiency, particularly for long speech inputs, whereas xLSTM suffers from the slowest processing speed.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。