arXiv:2505.21472cs.CVcs.CL2025-05被引 13

通过自适应注意力校准,减少大模型视觉幻觉问题。

Mitigating Hallucination in Large Vision-Language Models via Adaptive Attention Calibration

  • 基于置信度动态调整注意力分布,纠正视觉感知偏差。
  • 在长文本生成中幻觉率降低,多个基准测试表现更优。
  • 适合需要高视觉一致性的多模态生成任务使用。

大型视觉语言模型(LVLM)在多模态任务中表现优异,但常出现幻觉,即自信地描述图像中不存在的物体或属性。当前无训练干预方法在开放问答和长文本生成场景中难以保持准确性。本文提出置信度感知注意力校准(CAAC)框架,针对两种关键偏差:空间感知偏差(注意力在图像标记上分布不均)和模态偏差(随生成过程注意力从视觉向文本转移)。CAAC采用两步策略:视觉标记校准(VTC)平衡图像标记间的注意力;自适应注意力重缩放(AAR)根据模型置信度强化视觉锚定。该置信度驱动的调整确保生成过程中的视觉一致性。在CHAIR、AMBER和POPE基准测试上的实验表明,CAAC优于基线方法,尤其在长文本生成中有效降低幻觉。

原文摘要 · Abstract (English)

Large vision-language models (LVLMs) achieve impressive performance on multimodal tasks but often suffer from hallucination, and confidently describe objects or attributes not present in the image. Current training-free interventions struggle to maintain accuracy in open-ended and long-form generation scenarios. We introduce the Confidence-Aware Attention Calibration (CAAC) framework to address this challenge by targeting two key biases: spatial perception bias, which distributes attention disproportionately across image tokens, and modality bias, which shifts focus from visual to textual inputs over time. CAAC employs a two-step approach: Visual-Token Calibration (VTC) to balance attention across visual tokens, and Adaptive Attention Re-Scaling (AAR) to reinforce visual grounding guided by the model's confidence. This confidence-driven adjustment ensures consistent visual alignment during generation. Experiments on CHAIR, AMBER, and POPE benchmarks demonstrate that CAAC outperforms baselines, particularly in long-form generations, effectively reducing hallucination.

视觉语言模型幻觉抑制注意力校准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。