通过引导模型关注图像与标题的关联,减少视觉语言模型的幻觉问题。
CAI: Caption-Sensitive Attention Intervention for Mitigating Object Hallucination in Large Vision-Language Models
- 利用标题查询时的注意力模式增强视觉感知
- 在四个基准上达到当前最佳幻觉抑制效果
- 无需训练,推理开销极小,可直接部署
尽管大型视觉语言模型(LVLMs)在理解视觉信息方面表现出强大能力,但它们经常生成偏离实际视觉内容的内容,导致对象幻觉。现有方法大多依赖昂贵的手动标注和训练成本,或显著增加推理时间。本文观察到,当回答标题相关问题时,LVLMs对视觉信息的注意力显著强于非标题类问题。受此启发,我们提出无训练、即插即用的标题敏感注意力干预(CAI)方法,利用标题查询时的注意力激活模式来提升模型的视觉感知能力。在涵盖判别性和生成性任务的四个基准上的大量实验表明,CAI仅需极小额外推理开销,即可实现最先进的幻觉缓解性能。
原文摘要 · Abstract (English)
Although Large Vision-Language Models (LVLMs) have demonstrated powerful capabilities in interpreting visual information, they frequently produce content that deviates from visual information, leading to object hallucination. To tackle this, recent works mostly depend on expensive manual annotations and training cost, or significantly increase inference time. In this work, we observe that LVLMs' attention to visual information is significantly stronger when answering caption queries compared to non-caption queries. Inspired by this phenomenon, we propose Caption-sensitive Attention Intervention (CAI), a training-free, plug-and-play hallucination mitigation method that leverages the attention activation pattern in response to caption queries to enhance LVLMs' visual perception capability. Extensive experimental results across four benchmarks covering both discriminative and generative tasks, demonstrate that CAI achieves state-of-the-art (SOTA) hallucination mitigating performance only with minimal additional inference cost.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。