用轻量Mamba模块增强适配器,高效提升语音音频模型性能
MambAdapter: Lightweight Mamba-Based Adapters for Parameter-Efficient Transfer Learning in Speech and Audio

- 在低秩适配器中注入轻量Mamba模块,共享参数提升效率
- 4个音频分类任务和5种语音识别语言上性能媲美或超越基线
- 适合资源受限场景下语音音频模型的高效微调
基于Transformer的预训练模型已成为语音与音频处理领域域适应的主流策略。为降低微调带来的计算与内存开销,参数高效迁移学习(PETL)方法受到广泛关注。与此同时,近期出现的状态空间模型Mamba为序列建模提供了替代Transformer的有前景方案。本文提出MambAdapter,一种将Mamba融入低秩瓶颈适配器的参数高效迁移学习方法。通过在适配器间共享参数并引入轻量级Mamba模块,有效建模音频特征。实验表明,MambAdapter在四个音频分类任务和五个语音识别语言上表现达到或优于强基线,且在参数预算受限时仍保持优异性能。
原文摘要 · Abstract (English)
Fine-tuning Transformer-based foundation models has become the dominant strategy for domain adaptation in audio and speech processing. To reduce the computational and memory costs of this process, parameter-efficient transfer learning (PETL) methods have been widely explored. Meanwhile, Mamba, a recent state-space model, has emerged as a promising alternative to Transformers for sequence modeling. In this work, we present MambAdapter, a parameter-efficient transfer learning approach that integrates Mamba into low-rank bottleneck adapters. Our design combines parameter sharing across adapters with the injection of a lightweight Mamba module, enabling more effective modeling of audio features. We demonstrate that MambAdapter matches or outperforms strong PETL baselines on four audio classification tasks and five speech recognition languages, even when operating under reduced parameter budgets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。