arXiv:2512.00736cs.LGcs.AI2025-12被引 3

测试大模型在动态视角下的空间推理能力,发现其远不如人类。

REM: Evaluating LLM Embodied Spatial Reasoning through Multi-Frame Trajectories

  • 用可控制的3D环境生成多帧轨迹,评估模型空间推理能力。
  • 当前最佳模型在中等复杂度任务上已不可靠,人类轻松应对。
  • 适合研究具身智能、视觉语言模型空间理解的研究者。

人类通过导航构建与视角无关的认知地图,从而直观推理物体恒常性与空间关系。我们认为,尽管多模态大语言模型(MLLMs)经过大量视频训练,仍缺乏这一基础的空间推理能力,这对具身应用构成关键限制。为此,我们提出REM(基于具身多帧轨迹的空间推理评估),利用可控3D环境对长时程具身空间推理进行系统评估。REM涵盖物体恒常性/区分、空间关系、数值追踪等核心维度,评估不同动态视角下的表现。评估显示,当前最优模型虽整体表现良好,但在人类轻易处理的中等复杂度任务中已明显不可靠。这些发现揭示了MLLMs从序列视觉输入中构建鲁棒空间表征所面临的挑战。因此,REM提供了针对性指标与诊断工具,以推动未来模型空间理解能力的提升。

原文摘要 · Abstract (English)

Humans build viewpoint-independent cognitive maps through navigation, enabling intuitive reasoning about object permanence and spatial relations. We argue that multimodal large language models (MLLMs), despite extensive video training, lack this fundamental spatial reasoning capability, a critical limitation for embodied applications. To demonstrate these limitations and drive research, we introduce REM (Reasoning over Embodied Multi-Frame Trajectories), a benchmark using controllable 3D environments for long-horizon embodied spatial reasoning. REM systematically evaluates key aspects like object permanence/distinction, spatial relationships, and numerical tracking across dynamic embodied viewpoints. Our evaluation shows that the best-performing current models exhibit promising overall performance, but become increasingly unreliable at even moderate complexity levels easily handled by humans. These findings highlight challenges MLLMs face in developing robust spatial representations from sequential visual input. Consequently, REM provides targeted metrics and diagnostics to foster improved spatial understanding in future models.

空间推理具身智能多模态模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。