提出新方法提升视觉语言模型注意力一致性评估
Listening makes Vision Clear for VLMs

- 从提示词侧分析注意力激活,避免解码偏差影响
- 在多个数据集上显著优于传统基于答案的评估方法
- 适合研究多模态模型对齐机制与评估的学者
现有方法通常通过回答端标记的注意力分布评估视觉-语言一致性,但我们发现最高注意力区域并不总对应目标语义标记。这可能源于解码漂移:先前生成的回答标记所携带的语言先验会累积,与视觉注意力产生错配。此外,结构标记(如模态边界标记)可能覆盖整个上下文,导致对无关区域产生高注意力。为克服这些偏差并实现大模型的一致性评估,我们采用提示词侧语义,提出提示-视觉标记激活图(PV-TAM),并引入滤波器消除模态边界标记带来的系统性偏差。与仅关注重叠掩码的传统方法不同,我们的指标利用注意力峰值分布衡量提示与视觉区域的对齐程度。实验表明,PV-TAM在多个数据集上持续提升基于注意力和IoU的定位指标,优于回答侧基线。
原文摘要 · Abstract (English)
Recent work typically assesses vision--language consistency using attention distributions of answer-side tokens. However, we observe that highest attention regions are not always consistent with the intended semantic token. This probably stems from decoding drift, where language priors from previously generated answer tokens accumulate and mismatch with visual attention. Besides the priors from previous answer tokens, we find that structural tokens, e.g., modality boundary markers, may encompass the entire context and generate high attention to areas unrelated to the target. To avoid these distortions and provide consistency evaluation for large VLMs, we adopt prompt-side semantics and propose Prompt-Vision Token Activation Map (PV-TAM). PV-TAM further incorporates a filter to remove systematic bias induced by modality boundary markers. Unlike traditional methods that evaluate overlap solely through masks while ignoring activation intensity, our metrics leverage the peak distribution of attention to measure the alignment between prompts and visual regions. In experiments, PV-TAM consistently improves both attention-based and IoU-style localization metrics over answer-side baselines on various datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。