arXiv:2511.03768cs.LGcs.CV2025-11NeurIPS被引 1

新基准测试发现多模态模型跨场景推理时严重幻觉,连顶尖模型准确率仅35%。

What's in Common? Multimodal Models Hallucinate When Reasoning Across Scenes

  • 设计新基准Common-O,用真实场景图问'有何共同点'来测跨场景推理
  • 顶尖模型在复杂场景上准确率仅1%,远低于感知任务表现
  • 模型易因物体共现模式幻觉,多图训练比单纯扩大规模更有效

多模态语言模型虽能识别海量物体,但在真实场景推理中仍存在严重幻觉,暴露出其在感知基准上表现优异与实际推理能力之间的差距。为此,我们构建了名为Common-O的新基准,包含超过10.5k张未出现在网络训练数据中的真实场景图像,借鉴人类认知测试,通过‘有何共同点’问题探测跨场景推理能力。评估主流多模态模型,包括专门训练链式思维的模型,发现单图物体识别尚可,但跨场景推理极难——最佳模型在Common-O上仅达35%准确率,在更复杂的Common-O Complex上仅1%。有趣的是,当场景中存在相似物体时,模型更易幻觉,暗示其依赖训练中习得的物体共现模式。规模扩大带来小幅提升,而显式使用多图像输入的模型改善更显著,表明规模化多图像训练或具前景。本研究公开基准以推动该领域研究。

原文摘要 · Abstract (English)

Multimodal language models possess a remarkable ability to handle an open-vocabulary's worth of objects. Yet the best models still suffer from hallucinations when reasoning about scenes in the real world, revealing a gap between their seemingly strong performance on existing perception benchmarks that are saturating and their reasoning in the real world. To address this gap, we build a novel benchmark of in-the-wild scenes that we call Common-O. With more than 10.5k examples using exclusively new images not found in web training data to avoid contamination, Common-O goes beyond just perception, inspired by cognitive tests for humans, to probe reasoning across scenes by asking "what's in common?". We evaluate leading multimodal language models, including models specifically trained to perform chain-of-thought reasoning. We find that perceiving objects in single images is tractable for most models, yet reasoning across scenes is very challenging even for the best models, including reasoning models. Despite saturating many leaderboards focusing on perception, the best performing model only achieves 35% on Common-O -- and on Common-O Complex, consisting of more complex scenes, the best model achieves only 1%. Curiously, we find models are more prone to hallucinate when similar objects are present in the scene, suggesting models may be relying on object co-occurrence seen during training. Among the models we evaluated, we found scale can provide modest improvements while models explicitly trained with multi-image inputs show bigger improvements, suggesting scaled multi-image training may offer promise. We make our benchmark publicly available to spur research into the challenge of hallucination when reasoning across scenes.

多模态幻觉推理基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。