用积木搭建测试模型物理推理能力,发现大模型在复杂任务中表现骤降。
PhyBlock: A Progressive Benchmark for Physical Understanding and Planning via 3D Block Assembly
- 设计四级递进式积木组装任务,评估空间推理与物理理解。
- 21个主流模型在高阶任务中性能显著下降,错误率超60%。
- 适合研究具身智能、视觉语言模型物理推理的学者参考。
尽管视觉语言模型(VLMs)在具身智能体的推理与规划中展现出潜力,其对物理现象的理解,尤其是在结构化3D环境中的能力仍严重受限。为此,我们提出PhyBlock,一个通过机器人积木组装任务评估VLMs物理理解与规划能力的渐进式基准。PhyBlock包含新型四层认知层级的积木组装任务及针对性的视觉问答(VQA)样本,旨在评估逐步提升的空间推理与基础物理理解,包括物体属性、空间关系和整体场景理解。该基准涵盖2600个任务(400个组装任务,2200个VQA任务),从部分完成度、失败诊断到规划鲁棒性三个维度评估模型表现。我们测试了21个前沿VLMs,揭示其在多步物理规划中的强弱差异。实证结果表明,随着任务复杂度增加,模型性能显著下滑;错误分析显示,空间朝向与依赖关系推理仍存在持续困难。令人意外的是,思维链提示(chain-of-thought prompting)改善有限,说明空间任务高度依赖模型的直觉理解。PhyBlock被定位为统一测试平台,推动具身推理发展,弥合视觉语言理解与真实物理问题求解间的鸿沟。
原文摘要 · Abstract (English)
While vision-language models (VLMs) have demonstrated promising capabilities in reasoning and planning for embodied agents, their ability to comprehend physical phenomena, particularly within structured 3D environments, remains severely limited. To close this gap, we introduce PhyBlock, a progressive benchmark designed to assess VLMs on physical understanding and planning through robotic 3D block assembly tasks. PhyBlock integrates a novel four-level cognitive hierarchy assembly task alongside targeted Visual Question Answering (VQA) samples, collectively aimed at evaluating progressive spatial reasoning and fundamental physical comprehension, including object properties, spatial relationships, and holistic scene understanding. PhyBlock includes 2600 block tasks (400 assembly tasks, 2200 VQA tasks) and evaluates models across three key dimensions: partial completion, failure diagnosis, and planning robustness. We benchmark 21 state-of-the-art VLMs, highlighting their strengths and limitations in physically grounded, multi-step planning. Our empirical findings indicate that the performance of VLMs exhibits pronounced limitations in high-level planning and reasoning capabilities, leading to a notable decline in performance for the growing complexity of the tasks. Error analysis reveals persistent difficulties in spatial orientation and dependency reasoning. Surprisingly, chain-of-thought prompting offers minimal improvements, suggesting spatial tasks heavily rely on intuitive model comprehension. We position PhyBlock as a unified testbed to advance embodied reasoning, bridging vision-language understanding and real-world physical problem-solving.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。