用自监督音频表征构建通用听觉大模型,性能超越专用模型。
SALMONN-2: Advancing General-Purpose Hearing Abilities with Self-Supervised Representations

- 采用多层特征融合适配器,整合自监督编码器各层信息。
- 在语音、音乐、副语言任务上表现均衡,多项指标达开源模型最佳。
- 证明多模态上下文学习需针对性训练,非自然涌现。
近期音频大语言模型(ALLM)通常基于大量标注数据训练的音频编码器构建。由于自监督学习(SSL)音频编码器能学习通用且可迁移的表征,我们探究其是否可作为ALLM的有效基础。本文提出SALMONN-2,一个基于统一SSL编码器的ALLM。为更好利用SSL编码器学到的分层表征,我们设计多层特征融合(MLF)适配器,在投影至语言模型前聚合所有编码器层的信息。除传统音频理解任务外,我们进一步探索了ALLM中的多模态上下文学习(MICL),并研究如何通过上下文偏置训练获得该能力。实验表明,通用SSL编码器性能可媲美甚至超过专用监督编码器,且在语音、音频、音乐和副语言任务间表现更均衡。SALMONN-2在同类规模开源模型中达到领先水平,在MMAU-Pro、MMAR和MMSU基准上取得最优结果。我们还发现,MICL不会自然涌现,但可通过针对性的上下文偏置训练有效获得。
原文摘要 · Abstract (English)
Recent audio large language models (ALLMs) are typically built upon audio encoders trained with large amounts of supervised data. Since self-supervised learning (SSL) audio encoder models are known to learn general-purpose and transferable representations, we investigate whether general-purpose SSL audio representations can serve as an effective foundation for ALLMs. We present SALMONN-2, an ALLM built upon a unified SSL encoder. To better exploit the hierarchical representations learned by SSL encoders, we propose a multi-layer feature fusion (MLF) adapter that aggregates information from all encoder layers before projecting them into the language model. Beyond conventional audio understanding tasks, we further explore multimodal in-context learning (MICL) in ALLMs and study how this capability can be acquired through contextual biasing training. Experimental results show that a general-purpose SSL encoder achieves performance comparable to, or better than, specialised supervised audio encoders while providing a more balanced capability across speech, audio, music and paralinguistic tasks. SALMONN-2 further achieves state-of-the-art performance among comparable-scale open-weight models on ALLM understanding benchmarks, obtaining the best results on MMAU-Pro, MMAR and MMSU. We also show that MICL does not emerge naturally in ALLMs, but can be effectively acquired through targeted contextual biasing training.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。