提出分层探测评估框架,量化大模型在图文一致性上的幻觉问题
H-POPE: Hierarchical Polling-based Probing Evaluation of Hallucinations in Large Vision-Language Models
- 采用由粗到细的层级探测设计,分阶段检验物体存在与属性幻觉
- 发现模型在物体存在性上幻觉率超60%,属性幻觉更严重
- 适用于评估视觉语言模型可靠性,尤其关注医疗/安防等高风险场景
通过结合文本与图像,大型视觉语言模型(LVLMs)在多种多模态任务中取得了显著进展。然而,这些模型常出现幻觉现象,例如文本输出与视觉输入不一致。为解决此问题,我们提出了H-POPE——一个由粗到细的基准评估体系,系统性地检测物体存在性和属性层面的幻觉。评估结果表明,模型在物体存在性判断上易产生幻觉,且对细粒度属性的幻觉更为严重。此外,我们进一步探究了模型是否依赖视觉输入生成文本内容。
原文摘要 · Abstract (English)
By leveraging both texts and images, large vision language models (LVLMs) have shown significant progress in various multi-modal tasks. Nevertheless, these models often suffer from hallucinations, e.g., they exhibit inconsistencies between the visual input and the textual output. To address this, we propose H-POPE, a coarse-to-fine-grained benchmark that systematically assesses hallucination in object existence and attributes. Our evaluation shows that models are prone to hallucinations on object existence, and even more so on fine-grained attributes. We further investigate whether these models rely on visual input to formulate the output texts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。