arXiv:2605.24456cs.CV2026-05中稿 · CVPR被引 1

评测大模型在第一人称3D空间关系推理中的表现,揭示其空间认知能力短板。

EgoProx: Evaluating MLLMs on Egocentric 3D Proximity Reasoning Across a Cognitive Hierarchy

论文配图:EgoProx: Evaluating MLLMs on Egocentric 3D Proximity Reasoning Across a Cognitive Hierarchy
图 1 · 摘自论文原文
  • 构建分层认知任务链,覆盖意图、探索、利用和动作链推理
  • 大模型在跨领域任务中提升明显,但空间推理仍不充分
  • 适合研究多模态模型空间理解与具身智能的学者参考

人类在日常生活中持续进行3D空间邻近性推理,即身体与周围物体的关系判断,以指导感知与行为。当前多模态大语言模型(MLLMs)是否具备此类具身3D推理能力尚不明确。为此,我们提出EgoProx基准,用于评估第一人称视角下的3D邻近性推理。任务按认知链条组织,涵盖意图、探索、利用及动作链推理。我们设计了一个基于代理的数据生成引擎,可规模化生成多样且一致的问答对。在EgoProx上对主流MLLMs进行评测,并开展针对数据集与任务的指令微调分析。结果显示显著的跨域性能提升,表明当前模型已具备一定空间知识;然而,它们仍难以有效利用这些知识进行空间推理型视觉问答。

原文摘要 · Abstract (English)

Humans constantly reason about 3D proximity, the relations between their body and surrounding objects, to guide perception and action in daily life. Whether multimodal large language models (MLLMs) can perform such embodied 3D reasoning remains unclear. To this end, we introduce EgoProx, a benchmark for egocentric 3D proximity reasoning. We organize our tasks along a cognitive chain, covering intention, exploration, exploitation, and chain-of-actions reasoning. We also design an agent based data engine that produces diverse and consistent QA pairs at scale. We benchmark prevailing MLLMs on EgoProx and conduct additional analyses with dataset specific and task specific instruction tuning. We observe large cross-domain gains, indicating that current MLLMs contain some spatial knowledge; however, they still struggle to effectively leverage it for spatial reasoning VQA.

3D推理多模态模型具身智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。