arXiv:2602.03677cs.CL2026-02被引 5

揭示指令如何通过注意力机制控制多模态模型的决策过程。

Instruction Anchor: Dissecting the Mechanistic Dynamics of Modality Arbitration

  • 从信息流动视角分析指令如何作为结构锚点引导多模态选择。
  • 仅5%关键注意力头被阻断即显著降低指令跟随能力,但不影响通用视觉语言能力。
  • 精准放大特定注意力头可使失败案例恢复约60%,适用于模型可解释性研究。

模态跟随是根据用户指令选择性利用多模态上下文的能力,对多模态大语言模型(MLLMs)在真实场景中的安全与可靠性至关重要。然而,其内部决策机制仍不明确。本文从信息流角度探究模态跟随机制,发现指令标记充当模态仲裁的结构锚点:浅层注意力层进行无差别信息传递,将多模态线索聚合至指令标记作为潜在缓冲;深层注意力层则选择性增强与指令一致的子空间,依据指令意图完成模态仲裁,且仅有少量注意力头驱动该过程。针对性干预验证了这些头的功能特异性:阻断其中仅5%的头即显著降低模态跟随表现,但保留通用视觉与语言能力;而针对性增强可使失败样本恢复约60%。本工作为模态跟随提供了机制性解释,并为未来提升MLLM在用户指令下整合多模态证据的能力提供指导。

原文摘要 · Abstract (English)

Modality following is the ability to selectively leverage multimodal contexts based on user instructions. It is fundamental to the safety and reliability of multimodal large language models (MLLMs) in real-world deployments. However, the internal mechanisms governing this decision-making process remain largely under-explored. In this work, we investigate the mechanism underlying modality following through an information flow perspective. Our findings reveal that instruction tokens serve as structural anchor for modality arbitration: Shallow attention layers perform undifferentiated information transfer, aggregating multimodal cues to instruction tokens as a latent buffer; in contrast, deep attention layers selectively strengthen the instruction-compliant subspace and resolve modality arbitration according to the instruction-specified intent, with a sparse subset of attention heads driving this process. Targeted attention-head interventions further validate the functional specificity of these heads: blocking only $5\%$ of the identified heads substantially degrades modality following while preserving general visual and language capabilities, whereas targeted amplification can restore failed modality-following samples by up to approximately $60\%$. Together, this work provides a mechanistic account of modality following and informs future efforts to improve how MLLMs integrate and utilize multimodal evidence under user instructions.

多模态注意力机制可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。