arXiv:2511.13420cs.CV2025-11中稿 · ACM MM 2026 Datase…

提出新评估框架,检测视觉语言模型在自由想象中的幻觉问题

VOPE: Revisiting Hallucination of Vision-Language Models in Voluntary Imagination Task

  • 设计基于反向核查的评估机制,判断模型对虚构对象是否存在判断是否准确
  • 发现主流模型在自由想象任务中幻觉严重,对虚构对象的识别率极低
  • 揭示现有去幻觉方法在此类任务中效果有限,为未来研究指明方向

当前关于大视觉语言模型幻觉的研究多聚焦于禁止输出图像外内容的事实描述任务,却忽视了故事写作等自由想象任务中的幻觉现象。为此,本文提出自愿想象对象存在性评估(VOPE)——一种基于反向核查的评估基准,用于评测模型在自由想象任务中的对齐行为。具体而言,VOPE通过追问模型对其生成内容中虚构对象的存在性判断来评估其合理性。该评估不惩罚虚构内容本身,而是关注模型对生成对象是否存在判断的准确性。基于此,构建涵盖图像描述、推理和写作任务、具有不同想象强度的数据集。在多个主流视觉语言模型及去幻觉方法上应用VOPE,发现:(1)大多数模型在自由想象中幻觉严重,对虚构对象的存在性判断表现很差;(2)现有去幻觉方法在该任务中效果有限,亟需针对性改进。

原文摘要 · Abstract (English)

Most research on hallucinations in Large Vision-Language Models (LVLMs) focuses on factual description tasks that prohibit any output absent from the image. However, little attention has been paid to hallucinations in voluntary imagination tasks, such as story writing, despite this human-like cognitive ability being essential for real-world generative applications. To address this limitation, we introduce Voluntary-imagined Object Presence Evaluation (VOPE) -- a recheck-based evaluation benchmark for assessing LVLMs' grounding behavior in voluntary imagination tasks. Specifically, VOPE poses recheck-based questions to evaluate how an LVLM interprets the presence of the imagined objects in its own response. Rather than penalizing the imagined content itself, VOPE identifies hallucinations based on the correctness of the model's presence judgments for the generated objects. Built on this idea, we construct a dataset covering captioning, reasoning, and writing tasks with different levels of voluntary imagination. We apply VOPE to several mainstream LVLMs and hallucination mitigation methods, revealing two key findings: (1) most LVLMs hallucinate heavily during voluntary imagination, and their performance in presence evaluation is notably poor on imagined objects; (2) existing hallucination mitigation methods show limited effect in voluntary imagination tasks, making this an important direction for future research.

视觉语言模型幻觉检测自由想象评估基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。