arXiv:2510.25548cs.RO2025-10被引 4

用视觉语言模型提前识别机器人任务规划中的失败风险

Using VLM Reasoning to Constrain Task and Motion Planning

  • 利用预训练视觉语言模型进行常识空间推理,提前发现任务规划的不可行性
  • 在三个复杂场景中将规划时间大幅缩短,部分情况彻底消除运动规划失败
  • 适合需要高效长时序机器人规划的科研与工程应用

在任务与运动规划中,高层任务规划通过世界抽象实现高效搜索,但其可行性依赖于抽象向连续运动的可细化性。当抽象的可细化性差时,看似合理的任务计划在运动规划阶段可能失败,需重规划,导致整体性能下降。现有方法仅在细化失败后添加约束,浪费大量搜索资源。本文提出VIZ-COAST,利用大型预训练视觉语言模型的常识空间推理能力,预先识别向下细化问题,避免规划中出现失败。在三个挑战性TAMP领域实验表明,该方法能从图像和领域描述中提取有效约束,显著减少规划时间,在某些情况下完全消除向下细化失败,且对更广泛实例具有泛化能力。

原文摘要 · Abstract (English)

In task and motion planning, high-level task planning is done over an abstraction of the world to enable efficient search in long-horizon robotics problems. However, the feasibility of these task-level plans relies on the downward refinability of the abstraction into continuous motion. When a domain's refinability is poor, task-level plans that appear valid may ultimately fail during motion planning, requiring replanning and resulting in slower overall performance. Prior works mitigate this by encoding refinement issues as constraints to prune infeasible task plans. However, these approaches only add constraints upon refinement failure, expending significant search effort on infeasible branches. We propose VIZ-COAST, a method of leveraging the common-sense spatial reasoning of large pretrained Vision-Language Models to identify issues with downward refinement a priori, bypassing the need to fix these failures during planning. Experiments on three challenging TAMP domains show that our approach is able to extract plausible constraints from images and domain descriptions, drastically reducing planning times and, in some cases, eliminating downward refinement failures altogether, generalizing to a diverse range of instances from the broader domain.

机器人规划视觉语言模型任务与运动规划

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。