构建新基准评估视觉语言模型幻觉问题,揭示其隐式推理弱点。
Towards a Systematic Evaluation of Hallucinations in Large-Vision Language Models
- 按显著性与隐性实体分类,设计上下文推理攻击测试幻觉
- 11个主流模型在隐性实体任务中幻觉率超60%
- 适合研究多模态模型可靠性与幻觉抑制的开发者
大型视觉语言模型(LVLMs)在复杂多模态任务中表现卓越,但仍存在幻觉问题,尤其在需要从图像中隐式识别或推断多样视觉实体时。为此,我们提出HALLUCINOGEN——一个新型视觉问答(VQA)基准,采用上下文推理提示作为幻觉攻击手段,评估前沿LVLMs的幻觉程度。该基准首先根据图像中实体的可识别性,将视觉实体分为显著实体(如汽车等明显可见物体)和隐性实体(如从胸部X光片中识别疾病,需领域知识或上下文推理)。随后,针对两类实体设计幻觉攻击,评估模型在定位或推理特定实体时的幻觉情况,要求模型在生成回答前验证查询实体是否存在。对11个主流LVLMs(包括LLaMA-3.2、DeepSeek-V2、Gemini等开源与商业模型)及两种幻觉缓解策略在多个数据集上的广泛评测表明,当前模型仍极易受到幻觉攻击,尤其在处理隐性实体任务时表现脆弱。
原文摘要 · Abstract (English)
Large Vision-Language Models (LVLMs) have demonstrated remarkable performance in complex multimodal tasks. However, these models still suffer from hallucinations, particularly when required to implicitly recognize or infer diverse visual entities from images for complex vision-language tasks. To address this challenge, we propose HALLUCINOGEN, a novel visual question answering (VQA) benchmark that employs contextual reasoning prompts as hallucination attacks to evaluate the extent of hallucination in state-of-the-art LVLMs. Our benchmark provides a comprehensive study of the implicit reasoning capabilities of these models by first categorizing visual entities based on the ease of recognition in an image as either salient (prominent, visibly recognizable objects such as a car) or latent entities (such as identifying a disease from a chest X-ray), which are not readily visible and require domain knowledge or contextual reasoning for accurate inference. Next, we design hallucination attacks for both types of entities to assess hallucinations in LVLMs while performing various vision-language tasks, such as locating or reasoning about specific entities within an image, where models must perform implicit reasoning by verifying the existence of the queried entity within the image before generating responses. Finally, our extensive evaluations of eleven LVLMs, including powerful open-source models (like LLaMA-3.2 and DeepSeek-V2), commercial models like Gemini, and two hallucination mitigation strategies across multiple datasets, demonstrate that current LVLMs remain susceptible to hallucination attacks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。