arXiv:2512.23328cs.AIcs.CL2025-12被引 2

用魔方测试大模型的空间推理与长期规划能力,发现其在复杂任务中完全失效。

CubeBench: Diagnosing Interactive, Long-Horizon Spatial Reasoning Under Partial Observations

  • 通过分层诊断框架,逐步测试模型从完整信息到部分观测下的空间推理能力。
  • 顶尖大模型在长时序任务上全为0%通过率,暴露长期规划根本缺陷。
  • 适合研究物理世界智能体、具身认知和大模型可解释性的人参考。

大型语言模型虽在数字领域表现优异,但在物理世界部署中仍面临显著挑战,主要源于难以构建并维持稳健的空间心智模型。我们识别出三大核心认知障碍:空间推理、基于心理模拟的长时序状态追踪,以及在部分观测下的主动探索。为隔离并评估这些能力,我们提出CubeBench——一个以魔方为中心的生成式基准。该基准采用三层诊断框架,从具备完整符号信息的状态追踪,逐步过渡到仅依赖部分视觉数据的主动探索。对主流大模型的实验表明,在所有长时序任务中均达到0.00%通过率,暴露出长期规划的根本性失败。我们进一步引入外部求解器工具,构建诊断框架以分离认知瓶颈。通过对失败模式的分析,为开发更具物理基础的智能体提供了关键洞见。

原文摘要 · Abstract (English)

Large Language Model (LLM) agents, while proficient in the digital realm, face a significant gap in physical-world deployment due to the challenge of forming and maintaining a robust spatial mental model. We identify three core cognitive challenges hindering this transition: spatial reasoning, long-horizon state tracking via mental simulation, and active exploration under partial observation. To isolate and evaluate these faculties, we introduce CubeBench, a novel generative benchmark centered on the Rubik's Cube. CubeBench uses a three-tiered diagnostic framework that progressively assesses agent capabilities, from foundational state tracking with full symbolic information to active exploration with only partial visual data. Our experiments on leading LLMs reveal critical limitations, including a uniform 0.00% pass rate on all long-horizon tasks, exposing a fundamental failure in long-term planning. We also propose a diagnostic framework to isolate these cognitive bottlenecks by providing external solver tools. By analyzing the failure modes, we provide key insights to guide the development of more physically-grounded intelligent agents.

空间推理长时序规划大模型评测具身智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。