arXiv:2509.18816cs.SDcs.CL2025-09被引 5

让大模型更关注音频信息,提升听觉推理能力

Pay More Attention To Audio: Mitigating Imbalance of Cross-Modal Attention in Large Audio Language Models

  • 在注意力计算后动态增强音频令牌权重,无需额外参数
  • 在MMAR基准上使开源模型超越专有Gemini 2.0 Flash
  • 适合需要提升音频理解能力的多模态研究者

大型音频语言模型(LALMs)常因音频与文本注意力失衡,过度侧重文本而忽略声学信息,尤其在Transformer的多模态融合层中表现明显,影响音频推理任务性能。为此,我们提出MATA——一种无需训练的新型方法,通过在自注意力机制中动态增强音频令牌的关注度。MATA在原始注意力得分后介入,仅作用于中间层的最后一个令牌,不引入额外参数或计算开销。在MMAU和MMAR基准上的实验验证了其有效性,性能持续提升。尤其在MMAR上,使开源模型首次超越专有模型Gemini 2.0 Flash。本工作为缓解注意力偏差提供了高效方案,并开辟了增强多模态模型音频处理能力的新方向。

原文摘要 · Abstract (English)

Large Audio-Language Models (LALMs) often suffer from audio-textual attention imbalance, prioritizing text over acoustic information, particularly in the multi-modal fusion layers of the Transformer architecture. This bias hinders their ability to fully utilize acoustic cues, causing suboptimal performance on audio reasoning tasks. To mitigate this, we propose \textbf{MATA}, a novel training-free method that dynamically pushes LALMs to pay \textbf{M}ore \textbf{A}ttention \textbf{T}o \textbf{A}udio tokens within the self-attention mechanism. Specifically, MATA intervenes post raw attention scoring, targeting only the last token in intermediate layers without introducing additional parameters or computational overhead. Experiments on the MMAU and MMAR benchmarks confirm MATA's effectiveness, with consistent performance gains. Notably, on MMAR, MATA enables an open-source model to surpass the proprietary Gemini 2.0 Flash for the first time. Our work provides an efficient solution to mitigate attention bias and opens a new research direction for enhancing the audio-processing capabilities of multi-modal models.

音频理解注意力机制多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。