研究视觉语言模型为何因提示而产生幻觉,发现关键注意力头可减少40%以上幻觉。
Mechanisms of Prompt-Induced Hallucination in Vision-Language Models
- 通过控制物体计数实验,分析提示如何诱导模型忽略图像真实内容。
- 发现少数特定注意力头导致幻觉,删除后幻觉降低至少40%。
- 揭示不同模型中幻觉机制的差异,适合研究模型可靠性与可解释性者阅读。
大型视觉语言模型虽强大,却常因过度依赖文本提示而产生幻觉。本文在受控的物体计数任务中研究这一问题:当提示声称图像中有四个睡莲,实际仅有三个时,模型会倾向于遵循提示。在低数量时,模型能纠正错误;但随着物体数量增加,其越来越偏离视觉证据,盲目服从提示。通过对三种VLM进行机制分析,我们识别出一组少量注意力头(PIH-heads),其移除可使提示诱导幻觉(PIH)下降至少40%,且无需额外训练。这些头部在不同模型中以特有方式传递提示信息,其移除后模型更倾向于依据视觉证据做出判断。研究揭示了幻觉行为的内部机制,并指出模型间实现方式存在差异。
原文摘要 · Abstract (English)
Large vision-language models (VLMs) are highly capable, yet often hallucinate by favoring textual prompts over visual evidence. We study this failure mode in a controlled object-counting setting, where the prompt overstates the number of objects in the image (e.g., asking a model to describe four waterlilies when only three are present). At low object counts, models often correct the overestimation, but as the number of objects increases, they increasingly conform to the prompt regardless of the discrepancy. Through mechanistic analysis of three VLMs, we identify a small set of attention heads whose ablation substantially reduces prompt-induced hallucinations (PIH) by at least 40% without additional training. Across models, PIH-heads mediate prompt copying in model-specific ways. We characterize these differences and show that PIH ablation increases correction toward visual evidence. Our findings offer insights into the internal mechanisms driving prompt-induced hallucinations, revealing model-specific differences in how these behaviors are implemented.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。