定位音频模型中的倾听信号,用干预提升音频利用率。
Are Audio-Language Models Listening? Audio-Specialist Heads for Adaptive Audio Steering
- 通过可解释性找到关键音频注意力头,识别模型是否真正听懂音频。
- 在MMAU数据集上使准确率最高提升8.0个百分点,无需更新参数。
- 适合想提升音频模型表现但无法重训练的研究者使用。
多模态大语言模型常表现出文本主导现象,过度依赖语言先验而忽视非文本输入。以大音频-语言模型(LALMs)为例,即使音频包含关键信息,模型也常未充分使用。本文通过机制可解释性,定位出少数音频专用注意力头,其注意力激活可作为模型“倾听”音频的信号。我们发现该信号在音频影响输出时显著增强,从而提供一种标准提示下音频参与度的指标。基于此定位,构建音频-静音调控方向,并在推理阶段对最终表示施加激活干预,放大模型对音频的响应。在MMAU数据集上,该方法使两个基于Qwen的LALMs准确率最高提升8.0个百分点,且无需任何参数更新。
原文摘要 · Abstract (English)
Multimodal large language models can exhibit text dominance, over-relying on linguistic priors instead of grounding predictions in non-text inputs. One example is large audio-language models (LALMs) where decisive audio evidence can be under-utilized even when it contains important information. To address this issue we use mechanistic interpretability to identify a small set of audio-specialist attention heads whose audio attention yields a ``listening'' signal. We show that this signal increases when audio evidence affects the model's output, providing an indicator of audio engagement under standard prompting. Leveraging this localization, we construct an audio--silence steering direction and apply an inference-time activation intervention to the final representation, amplifying the model's audio effect. To demonstrate the utility of this intervention, we show on MMAU that this improves accuracy by up to +8.0 percentage points on two Qwen-based LALMs, without any parameter updates.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。