arXiv:2608.30653cs.CVcs.AI2026-08中稿 · CVPR被引 1

构建细粒度多图幻觉评测基准,揭示大模型跨图像推理中的物体幻觉问题

Fine-Grained Multi Image Object Hallucination Benchmark

论文配图:Fine-Grained Multi Image Object Hallucination Benchmark
图 1 · 摘自论文原文
  • 设计四种基础任务与三类推理模式,系统评估多图场景下的物体幻觉
  • 29个模型测试显示,即使顶尖模型如GPT-5在计数与位置任务中仍存在明显幻觉
  • 揭示幻觉源于多图信息整合阶段的表征保持缺陷,非单纯感知错误

多模态大语言模型(MLLMs)在需要跨视觉上下文复杂推理的多图场景中日益应用。然而,现有模型仍受物体幻觉困扰——生成看似合理但事实不一致的物体描述。现有评测基准主要针对单图场景或仅提供高层次的多图评估,无法系统诊断视觉复杂度和推理需求如何触发幻觉。为此,我们提出MIOH,一个细粒度的多图像物体幻觉评测基准,通过三种多图推理模式(综合、对比、选择)在三个受控对抗压力下(视觉上下文规模、感知难度、上下文偏见),系统评估四类基础任务(存在性、计数、属性、位置)中的幻觉现象。对29个模型的评估发现,即使最先进的GPT-5和Gemini-2.5-Pro也表现出不同推理模式与任务下的显著失败模式。研究揭示,幻觉不仅源于感知失误,更根植于多图信息整合阶段的对象表征维持缺陷。MIOH为分析多图物体幻觉提供了可控框架,是开发更可靠多模态AI系统的关键评测工具。

原文摘要 · Abstract (English)

Multimodal Large Language Models (MLLMs) are increasingly deployed in multi-image scenarios requiring complex reasoning across visual contexts. However, current MLLMs remain fundamentally limited by object hallucination-generating plausible yet factually inconsistent descriptions about objects. Existing benchmarks, designed primarily for single-image settings or providing only high-level multi-image assessments, cannot systematically diagnose how visual complexity and reasoning demands trigger hallucination. To address this gap, we introduce MIOH, a fine-grained multi-image object hallucination benchmark that systematically evaluates object hallucination across four foundational tasks (existence, counting, attribute, position) through three multi-image reasoning patterns (comprehensive, comparative, selective) under three controlled adversarial pressures (visual context scale, perceptual difficulty, contextual bias). Through evaluation of 29 models, we reveal that even state-of-the-art systems like GPT-5 and Gemini-2.5-Pro exhibit distinct failure patterns across different reasoning patterns and tasks. Our evaluation reveals that hallucination stems not merely from perceptual failures but from integration-stage limitations when maintaining object representations across multiple images. MIOH provides a controlled framework for analyzing multi-image object hallucination and serves as a critical evaluation tool for developing more reliable multimodal AI systems.

多模态幻觉评测大模型推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。