arXiv:2510.26139cs.RO2025-10被引 2

用视觉大模型提升复杂任务的运动规划成功率

Kinodynamic Task and Motion Planning using VLM-guided and Interleaved Sampling

  • 构建符号与数值状态统一的混合状态树,联合决策任务与动作
  • 实测成功率提升32%至1167%,复杂问题规划更快
  • 视觉大模型引导搜索并回溯错误路径,适合真实机器人场景

任务与运动规划(TAMP)将高层任务规划与底层运动可行性结合,但现有方法在长时序问题中因过度采样而成本高昂。虽然大语言模型提供常识先验,却缺乏三维空间推理能力且无法保证几何或动力学可行性。本文提出一种基于混合状态树的运动学-动力学TAMP规划器,统一表示符号与数值状态,实现任务与运动决策的联合优化。通过现成运动规划器与物理仿真验证动力学约束,视觉大模型则根据状态的视觉渲染结果指导搜索并进行回溯。在模拟环境与真实世界中的实验表明,相比传统及基于LLM的TAMP规划器,平均成功率提升32.14%至1166.67%,复杂问题规划时间显著减少;消融实验进一步验证了视觉大模型回溯机制的优势。

原文摘要 · Abstract (English)

Task and Motion Planning (TAMP) integrates high-level task planning with low-level motion feasibility, but existing methods are costly in long-horizon problems due to excessive motion sampling. While LLMs provide commonsense priors, they lack 3D spatial reasoning and cannot ensure geometric or dynamic feasibility. We propose a kinodynamic TAMP planner based on a hybrid state tree that uniformly represents symbolic and numeric states during planning, enabling task and motion decisions to be jointly decided. Kinodynamic constraints embedded in the TAMP problem are verified by an off-the-shelf motion planner and physics simulator, and a VLM guides exploring a TAMP solution and backtracks the search based on visual rendering of the states. Experiments on the simulated domains and in the real world show 32.14% - 1166.67% increased average success rates compared to traditional and LLM-based TAMP planners and reduced planning time on complex problems, with ablations further highlighting the benefits of VLM backtracking. More details are available at https://graphics.ewha.ac.kr/kinodynamicTAMP/.

任务规划运动规划视觉大模型机器人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。