arXiv:2509.15680cs.SDeess.AS2025-09中稿 · Interspeech 2026

用Mamba-2构建轻量音频语言模型,性能媲美大模型

SAM: A Mamba-2 State-Space Audio-Language Model

  • 采用Mamba-2架构与音频编码器融合,实现高效音频理解
  • 2.7B参数模型在AudioSet上达21.1 mAP,超越多数大模型
  • 首次揭示状态空间模型与音频表示的交互机制,指导设计

我们提出SAM,一种基于Mamba-2骨干网络的音频-语言模型,整合音频编码器与状态空间结构。SAM-2.7B在AudioSet上取得21.1 mAP,在AudioCaps上达到17.6 SPICE,参数更少却表现匹敌甚至超过更大的7B Transformer模型。我们首次系统性地从表示层面分析了状态空间模型(SSMs)与音频编码器输出的交互:(1) 联合微调音频编码器至关重要,表现为准确率提升及不同规模SSM下令牌表示秩与相似性的适应变化;(2) 尽管具有线性扩展性,但SSMs更受益于紧凑且信息丰富的音频令牌表示,而非过长序列;(3) 引入指令跟随监督显著提升推理能力,使MMAU-Sound准确率从22.8提升至56.8。通过全面实验与分析,我们确立了以SSMs作为强健、可扩展音频-语言模型骨干的实用设计原则。

原文摘要 · Abstract (English)

We present SAM, a State-space Audio-language Model that integrates an audio encoder with a Mamba-2 backbone. SAM-2.7B achieves 21.1 mAP on AudioSet and 17.6 SPICE on AudioCaps, matching or surpassing larger 7B transformer-based models with fewer parameters. We further provide the first systematic, representation-level analysis of how SSMs interact with audio encoder outputs: (1) joint audio encoder finetuning is essential, supported by accuracy gains and observed adaptation of token representation rank and similarity across different SSM sizes; (2) despite linear scaling, SSMs benefit more from compact, information-rich audio token representations than from excessively long token sequences; and (3) incorporating instruction-following supervision substantially improves reasoning ability, boosting MMAU-Sound accuracy from 22.8 to 56.8. Through comprehensive experiments and analysis, we establish practical design principles for SSMs as strong, scalable backbones for audio-language models.

音频语言模型Mamba状态空间模型轻量化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。