arXiv:2606.11400cs.SDcs.AI2026-06

通过指令控制注意力,让大模型聚焦音频关键片段。

Steering Where to Listen: Instruction-Based Activation Steering Redirects Temporal Attention in Large Audio-Language Models

论文配图:Steering Where to Listen: Instruction-Based Activation Steering Redirects Temporal Attention in Large Audio-Language Models
图 1 · 摘自论文原文
  • 用不同指令对比激活生成导向向量,动态调整模型听觉注意力位置。
  • 在三事件场景中,定位准确率达60.87%~68.72%,远超传统提示法。
  • 无需训练即可探测模型隐含的时间结构,适合可解释性研究者。

大型音频-语言模型(LALMs)在音频理解上表现优异,但其注意力分布不透明。本文提出基于指令的向量导向方法:固定音频输入,通过对比不同指令下的激活生成导向向量。系统性探查显示,该方法显著重分配时间注意力,集中于声学相关区域。进一步实验表明,此注意力转移具有行为意义:在三事件设定中,读取最大注意力变化位置可无训练恢复目标事件位置,在Qwen2-Audio和Audio Flamingo 3上与真实区间重叠率分别达60.87%和68.72%,远超直接提示法(31.84%、46.75%)和随机基线(27.74%)。结果揭示了指令导向在LALMs中的机制特性,并提供一种无需训练的潜在时间结构探测工具。

原文摘要 · Abstract (English)

Large Audio-Language Models (LALMs) excel at audio understanding but expose little about where in an audio signal they attend. We introduce instruction-based vector steering, which constructs a steering vector by contrasting activations from differently instructed prompts while keeping the audio fixed. Through a systematic probe of LALM attention, we find that - unlike standard prompting or audio-based steering - this intervention significantly redistributes the temporal attention allocated to audio tokens, concentrating it on acoustically relevant regions. We then show that this attention shift is behaviorally meaningful: in a controlled three-event setting, reading out the temporal position of maximal steering-induced attention change recovers the location of a queried sound event without any training, attaining 60.87% and 68.72% overlap with ground-truth intervals on Qwen2-Audio and Audio Flamingo 3, far above direct prompting (31.84%, 46.75%) and random baselines (27.74%). Our results characterize a mechanistic property of instruction-based steering in LALMs and provide a training-free probe for the latent temporal structure these models encode.

音频理解注意力机制可解释性指令调优

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。