提出Geo3R框架,用几何证据解决多模态大模型的空间推理幻觉问题。
Geo3R: Mitigating Spatial Reasoning Hallucination in Multimodal Large Language Models

- 引入几何证据与结构化3D推理,无需训练即可修复空间幻觉
- 在18项任务上显著降低幻觉率,跨多种模型有效
- 针对视角、方向、视点变化等典型场景设计,适合高精度视觉推理应用
尽管多模态大语言模型在视觉理解方面取得显著进展,但在空间关系推理中仍易产生幻觉,常给出与真实3D场景结构矛盾的判断。现有方法虽尝试缓解此问题,但分析表明其在空间推理上的效果有限,因未能弥合2D视觉表示与3D空间现实之间的根本差距。基于此,我们定义由空间结构建模不足引发的幻觉为‘空间推理幻觉’,属于现有缓解方法未覆盖的关系幻觉子类。我们识别出三类典型幻觉场景:透视效应、物体朝向与视角变化。为此,提出Geo3R——一种无需训练、可即插即用的框架,通过引入几何证据和结构化3D推理来减轻空间推理幻觉。在三个基准测试上,覆盖全部三类场景的18项任务,实验表明Geo3R在不增加训练成本的前提下,显著降低各类多模态大模型的空间推理幻觉,优于现有模型与方法。
原文摘要 · Abstract (English)
Despite remarkable progress in visual understanding, Multimodal Large Language Models (MLLMs) remain prone to hallucinations when reasoning about spatial relationships, often producing judgments that contradict the true 3D structure of the scene. Though several existing works have proposed to mitigate hallucinations, our analysis indicates that they show limited effectiveness in spatial reasoning, as they fail to bridge the fundamental gap between 2D visual representations and 3D spatial reality. Based on this finding, we define hallucinations arising from insufficient spatial structure modeling as spatial reasoning hallucination, a subcategory of relation hallucination that existing mitigation methods fail to address. We further identify three typical scenarios where such hallucinations frequently occur: perspective effects, object orientation, and viewpoint changes. To this end, we propose Geo3R, a training-free, plug-and-play framework that incorporates geometric evidence and structured 3D reasoning to mitigate spatial reasoning hallucination. Experiments on three benchmarks, covering 18 tasks across all three scenarios, show that Geo3R substantially reduces spatial reasoning hallucination across diverse MLLMs without additional training, outperforming existing models and methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。