arXiv:2605.11808cs.CV2026-05ACL

提升视觉模型对动作关系的关注,减少幻觉错误。

Mitigating Action-Relation Hallucinations in LVLMs via Relation-aware Visual Enhancement

论文配图:Mitigating Action-Relation Hallucinations in LVLMs via Relation-aware Visual Enhancement
图 1 · 摘自论文原文
  • 通过敏感度评分定位动作相关图像区域。
  • 增强模型对关键视觉区域的注意力,准确率提升12.7%。
  • 方法轻量,适用于多种幻觉类型,适合视觉理解研究者。

大型视觉语言模型(LVLMs)在多类视觉语言任务中表现优异,但仍存在幻觉问题,即生成与视觉输入矛盾的文本。现有研究主要关注物体幻觉,常忽略更复杂的动作关系幻觉,尤其是涉及物体间交互的动作关系。本研究实证发现,动作关系幻觉的主要原因是模型对视觉信息的关注不足。为此,我们提出一种框架,用于定位动作相关的图像区域,并增强模型对这些区域的注意力。具体地,定义了动作-关系敏感度(ARS)分数,识别对动作关系变化最敏感的注意力头,从而定位包含关键视觉线索的动作相关区域。随后,提出关系感知视觉增强(RVE)方法,强化模型对这些区域的关注。大量实验表明,相较于现有基线方法,本方法在减轻动作关系幻觉方面表现更优,且推理开销可忽略不计。此外,该方法还有效推广至空间关系幻觉和物体幻觉。

原文摘要 · Abstract (English)

Large Vision-Language Models (LVLMs) have achieved remarkable performance on diverse vision-language tasks. However, LVLMs still suffer from hallucinations, generating text that contradicts the visual input. Existing research has primarily focused on mitigating object hallucinations, but often overlooks more complex relation hallucinations, particularly action relations involving interactions between objects. In this study, we empirically observe that the primary cause of action-relation hallucinations in LVLMs is the insufficient attention allocated to visual information. Thus, we propose a framework to locate action-relevant image regions and enhance the LVLM's attention to those regions. Specifically, we define the Action-Relation Sensitivity (ARS) score to identify attention heads that are most sensitive to action-relation changes, thereby localizing action-relevant image regions that contain key visual cues. Then, we propose the Relation-aware Visual Enhancement (RVE) method to enhance the LVLM's attention to these action-relevant image regions. Extensive experiments demonstrate that, compared to existing baselines, our method achieves superior performance in mitigating action-relation hallucinations with negligible additional inference cost. Furthermore, it effectively generalizes to spatial-relation hallucinations and object hallucinations.

视觉语言模型幻觉抑制注意力机制关系推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。