用视觉语言模型生成约束,让机器人在开放世界中理解自然语言指令完成复杂操作。
Open-World Task and Motion Planning via Vision-Language Model Generated Constraints
- 用视觉语言模型生成动作顺序和代码化约束,引导机器人规划
- 在多个长程操作任务中优于纯TAMP或纯VLM方法
- 可直接部署于现成机器人系统,在真实硬件上完成挑战性任务
视觉语言模型(VLM)在常识视觉与语言任务中表现优异,但尚无法直接解决需要精确连续推理的复杂长程机器人操作问题。任务与运动规划(TAMP)系统通过参数化技能的离散-连续混合搜索处理长程推理,但依赖详细环境模型且无法理解新的人类目标,如任意自然语言指令。本文提出将VLM集成到TAMP系统中,通过其生成离散动作顺序约束与代码形式的连续约束,实现开放世界推理。具体而言,使用VLM生成约束以限制动作序列搜索,并用代码扩展传统TAMP操作约束。实验表明,所提方法OWL-TAMP在多个直接以自然语言描述的长程操作任务中优于仅依赖TAMP或仅依赖VLM的基线方法。此外,我们还展示了OWL-TAMP可与现成TAMP系统结合,在真实硬件上完成具有挑战性的操作任务。
原文摘要 · Abstract (English)
Foundation models like Vision-Language Models (VLMs) excel at common sense vision and language tasks such as visual question answering. However, they cannot yet directly solve complex, long-horizon robot manipulation problems requiring precise continuous reasoning. Task and Motion Planning (TAMP) systems can handle long-horizon reasoning through discrete-continuous hybrid search over parameterized skills, but rely on detailed environment models and cannot interpret novel human objectives, such as arbitrary natural language goals. We propose integrating VLMs into TAMP systems by having them generate discrete and continuous language-parameterized constraints that enable open-world reasoning. Specifically, we use VLMs to generate discrete action ordering constraints that constrain TAMP search over action sequences, and continuous constraints in the form of code that augments traditional TAMP manipulation constraints. Experiments show that our approach, OWL-TAMP, outperforms baselines relying solely on TAMP or VLMs across several long-horizon manipulation tasks specified directly in natural language. We additionally demonstrate that OWL-TAMP can be deployed with an off-the-shelf TAMP system to solve challenging manipulation tasks on real-world hardware.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。