让机器人计划具备空间可执行性,提升真实世界操作能力。
Spatially Grounded Long-Horizon Task Planning in the Wild
- 构建新基准GroundedPlanBench,评估动作规划与空间定位的协同能力
- 引入V2GP框架,用真实机器人视频自动生成带空间坐标的计划数据
- 实验证明当前视觉语言模型在空间执行上仍有明显短板
近年来,机器人操作越来越多地依赖视觉语言模型(VLMs)进行高层推理,如将任务指令分解为自然语言表达的动作序列以指导底层运动执行。然而,现有基准未评估这些计划是否具备空间可执行性,尤其缺乏对机器人应交互位置的明确指定,限制了对真实世界操作能力的评估。为此,我们定义了一种新型的具身规划任务,并提出GroundedPlanBench,一个面向野外环境的长时程、空间具身动作规划的新基准。该基准联合评估层级子动作规划与空间动作定位(何处执行),系统检验生成动作序列的空间可执行性。我们进一步提出视频到空间具身规划(V2GP)框架,利用真实机器人视频演示自动构建高质量训练数据,提升空间具身长时程规划能力。评估显示,当前VLMs在空间具身长时程规划方面仍存在显著瓶颈。实验表明,V2GP能有效提升动作规划与空间定位性能,在本基准及真实机器人实验中均得到验证,推动了空间可执行规划的发展。
原文摘要 · Abstract (English)
Recent advances in robot manipulation increasingly leverage Vision-Language Models (VLMs) for high-level reasoning, such as decomposing task instructions into sequential action plans expressed in natural language that guide downstream low-level motor execution. However, current benchmarks do not assess whether these plans are spatially executable, particularly in specifying the exact spatial locations where the robot should interact to execute the plan, limiting evaluation of real-world manipulation capability. To bridge this gap, we define a novel task of grounded planning and introduce GroundedPlanBench, a newly curated benchmark for spatially grounded long-horizon action planning in the wild. GroundedPlanBench jointly evaluates hierarchical sub-action planning and spatial action grounding (where to act), enabling systematic assessment of whether generated sub-actions are spatially executable for robot manipulation. We further introduce Video-to-Spatially Grounded Planning (V2GP), an automated data generation framework that leverages real-world robot video demonstrations to improve spatially grounded long-horizon planning. Our evaluations reveal that spatially grounded long-horizon planning remains a major bottleneck for current VLMs. Our results demonstrate that V2GP provides a promising approach for improving both action planning and spatial grounding performance, validated on our benchmark as well as through real-world robot manipulation experiments, advancing progress toward spatially actionable planning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。