通过因果分析,找出大模型幻觉生成的隐藏诱因。
Who Brings the Frisbee: Probing Hidden Hallucination Factors in Large Vision-Language Model via Causality Analysis
- 构建因果探针系统,分析图像、文本与模型注意力的关系。
- 发现背景如草地、天空等会显著诱发模型虚构飞盘。
- 可针对性干预模型内部机制,有效降低幻觉率。
大型视觉语言模型(LVLM)在理解视觉与语言输入方面取得显著进展,但其在实际应用中仍面临幻觉问题——即生成不存在的视觉元素,损害用户信任。目前对多模态幻觉背后的驱动机制仍缺乏深入理解。现有研究极少揭示诸如天空、树木或草地等上下文是否诱导模型生成飞盘。本文提出假设:物体、上下文及语义前后景结构等隐藏因素可能引发幻觉。为此,我们设计一种新颖的因果分析方法,构建幻觉探针系统,通过分析图像、文本提示与网络显著性之间的因果关系,系统性地测试阻断这些因素的干预手段。实验结果表明,基于该分析的简单策略能显著减少幻觉;同时,分析还揭示了修改模型内部结构以抑制幻觉输出的潜力。
原文摘要 · Abstract (English)
Recent advancements in large vision-language models (LVLM) have significantly enhanced their ability to comprehend visual inputs alongside natural language. However, a major challenge in their real-world application is hallucination, where LVLMs generate non-existent visual elements, eroding user trust. The underlying mechanism driving this multimodal hallucination is poorly understood. Minimal research has illuminated whether contexts such as sky, tree, or grass field involve the LVLM in hallucinating a frisbee. We hypothesize that hidden factors, such as objects, contexts, and semantic foreground-background structures, induce hallucination. This study proposes a novel causal approach: a hallucination probing system to identify these hidden factors. By analyzing the causality between images, text prompts, and network saliency, we systematically explore interventions to block these factors. Our experimental findings show that a straightforward technique based on our analysis can significantly reduce hallucinations. Additionally, our analyses indicate the potential to edit network internals to minimize hallucinated outputs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。