无需训练即可完成复杂装配任务并自动纠错的机器人规划系统
NovaPlan: Zero-Shot Long-Horizon Manipulation via Closed-Loop Video Language Planning
- 用视觉语言模型闭环规划任务,失败时自动重规划
- 从生成视频中提取物体关键点和手部姿态作为动作参考
- 在遮挡和深度误差下仍能稳定执行,适合真实场景
解决长时程任务需要机器人将高层语义推理与底层物理交互结合。尽管视觉-语言模型(VLM)和视频生成模型能分解任务并预演结果,但常缺乏真实世界执行所需的物理基础。我们提出NovaPlan,一种层次化框架,将闭环VLM与视频规划同几何精准的机器人执行统一,实现零样本长时程操作。高层通过VLM规划器将任务分解为子目标,并在闭环中监控执行,单步失败时可自主重规划。低层则从生成视频中提取任务相关物体关键点与人手姿态作为运动先验,采用切换机制选择更优参考以生成机器人动作,在严重遮挡或深度误差下仍保持稳定执行。我们在三个长时程任务和功能操作基准(FMB)上验证了其有效性。结果表明,NovaPlan可在无任何示范或训练的情况下完成复杂装配任务,并表现出灵巧的错误恢复能力。
原文摘要 · Abstract (English)
Solving long-horizon tasks requires robots to integrate high-level semantic reasoning with low-level physical interaction. While vision-language models (VLMs) and video generation models can decompose tasks and imagine outcomes, they often lack the physical grounding necessary for real-world execution. We introduce NovaPlan, a hierarchical framework that unifies closed-loop VLM and video planning with geometrically grounded robot execution for zero-shot long-horizon manipulation. At the high level, a VLM planner decomposes tasks into sub-goals and monitors robot execution in a closed loop, enabling the system to recover from single-step failures through autonomous re-planning. To compute low-level robot actions, we extract and utilize both task-relevant object keypoints and human hand poses as kinematic priors from the generated videos, and employ a switching mechanism to choose the better one as a reference for robot actions, maintaining stable execution even under heavy occlusion or depth inaccuracy. We demonstrate the effectiveness of NovaPlan on three long-horizon tasks and the Functional Manipulation Benchmark (FMB). Our results show that NovaPlan can perform complex assembly tasks and exhibit dexterous error recovery behaviors without any prior demonstrations or training. Project page: https://nova-plan.github.io/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。