arXiv:2505.17097cs.CVcs.CL2025-05AAAI被引 15

通过动态调制注意力提升视觉语言模型的上下文学习能力

Make LVLMs Focus: Context-Aware Attention Modulation for Better Multimodal In-Context Learning

论文配图:Make LVLMs Focus: Context-Aware Attention Modulation for Better Multimodal In-Context Learning
图 1 · 摘自论文原文
  • 提出CAMA方法,根据输入序列动态调整注意力权重
  • 在7个基准上显著提升4种LVLM的上下文学习性能
  • 无需训练、即插即用,适合希望提升推理能力的研究者

多模态上下文学习(ICL)使大型视觉语言模型(LVLMs)能在不更新参数的情况下适应新任务,扩展了其在真实场景中的应用。然而,即使示范样本匹配良好,现有模型表现仍不稳定,表明其未能充分利用上下文信息。现有工作多聚焦于提示工程或后处理校准,而本文研究LVLM内部的注意力机制,发现自注意力存在两大缺陷阻碍有效ICL。为此,提出无需训练、即插即用的上下文感知注意力调制(CAMA)方法,通过双阶段调制增强对语义重要标记(尤其是视觉标记)的关注。在四种LVLM和七个基准上,CAMA均显著优于原始模型与基线,展现良好效果与泛化性,且能激活提示工程的优势,对不同序列配置保持鲁棒。该方法为通过理解注意力动态改进多模态推理开辟新路径。

原文摘要 · Abstract (English)

Multimodal in-context learning (ICL) is becoming a key capability that allows large vision-language models (LVLMs) to adapt to novel tasks without parameter updates, which expands their usefulness in many real-world applications. However, ICL performance remains unstable even when the in-context demonstrations (ICDs) are well matched, showing that LVLMs still struggle to make full use of the provided context. While existing work mainly focuses on prompt engineering or post-hoc logit calibration, we study the attention mechanisms inside LVLMs to address their inherent limitations. We identify two important weaknesses in their self-attention that hinder effective ICL. To address these weaknesses, we propose Context-Aware Modulated Attention (CAMA), a training-free and plug-and-play method that dynamically adjusts attention logits based on the input in-context sequence. CAMA uses a two-stage modulation process that strengthens attention to semantically important tokens, especially visual ones. Across four LVLMs and seven benchmarks, CAMA consistently outperforms vanilla models and baselines, showing clear effectiveness and generalization. It can also activate the intended benefits of prompt engineering methods and remains robust across different sequence configurations. Therefore, CAMA opens up new directions for improving multimodal reasoning through a deeper understanding of attention dynamics.

多模态学习注意力机制上下文学习视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。