用视觉分割量化注意力不确定性,有效减少大模型幻觉
Segmentation-Based Attention Entropy: Detecting and Mitigating Object Hallucinations in Large Vision-Language Models
- 通过语义分割分析视觉注意力分布,衡量其不确定性
- 在多个基准上降低幻觉率,且无需额外训练
- 适合需要高可靠性的机器人视觉任务
大型视觉语言模型(LVLM)在多模态任务中表现强劲,但物体幻觉严重威胁其可靠性。现有研究多聚焦文本模态,认为幻觉源于过强的语言先验和不足的视觉支撑。本文观察到视觉模态内的异常注意力模式也会引发幻觉。为此提出基于分割的注意力熵(SAE),利用语义分割在语义空间中量化视觉注意力的不确定性。基于SAE,设计了幻觉检测的可靠性评分及推理时的注意力调整方法,动态修正视觉注意力以减轻幻觉。在公开基准和四足机器人真实多模态场景中评估,结果表明SAE无需额外训练即可显著降低幻觉,提升LVLM可靠性。
原文摘要 · Abstract (English)
Large Vision-Language Models (LVLMs) achieve strong performance on many multimodal tasks, but object hallucinations severely undermine their reliability. Most existing studies focus on the text modality, attributing hallucinations to overly strong language priors and insufficient visual grounding. In contrast, we observe that abnormal attention patterns within the visual modality can also give rise to hallucinated objects. Building on this observation, we propose Segmentation-based Attention Entropy (SAE), which leverages semantic segmentation to quantify visual attention uncertainty in a semantic space. Based on SAE, we further design a reliability score for hallucination detection and an SAE-guided attention adjustment method that modifies visual attention at inference time to mitigate hallucinations. We evaluate our approach on public benchmarks and in real embodied multimodal scenarios with quadruped robots. Experiments show that SAE reduces hallucinations without additional training, improving LVLM reliability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。