背景信息能有效减少视觉语言模型的幻觉生成。
The Role of Background Information in Reducing Object Hallucination in Vision-Language Models: Insights from Cutoff API Prompting
- 通过保留图像背景上下文,抑制对象幻觉。
- 注意力驱动的视觉提示在背景信息缺失时易失败。
- 适合关注模型可靠性与真实场景应用的研究者。
视觉语言模型(VLMs)有时会生成与输入图像矛盾的内容,限制了其在实际应用中的可信度。尽管有研究指出,通过在提示中加入图像内相关区域的视觉提示可抑制幻觉,但该方法在具体区域范围上的有效性尚不明确。本研究分析了注意力驱动视觉提示在对象幻觉中的成功与失败案例,发现保留背景上下文对于缓解对象幻觉至关重要。
原文摘要 · Abstract (English)
Vision-Language Models (VLMs) occasionally generate outputs that contradict input images, constraining their reliability in real-world applications. While visual prompting is reported to suppress hallucinations by augmenting prompts with relevant area inside an image, the effectiveness in terms of the area remains uncertain. This study analyzes success and failure cases of Attention-driven visual prompting in object hallucination, revealing that preserving background context is crucial for mitigating object hallucination.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。