用视觉语言模型捕捉物体间交互约束,提升复杂操作任务的规划效率。
Leveraging Inter-object Affordances for Efficient Planning in Contact-rich Tasks

- 基于物体中心的抽象约束,构建统一规划框架
- 规划成功率显著提升,耗时降低1-2个数量级
- 适合需要多物体物理交互的机器人任务
传统任务与运动规划(TAMP)方法主要关注动作序列及几何、运动学约束,但在真实场景中受限于简化物体模型,忽略接触密集任务中的关键物理属性。同时,运动规划常采用非符号推理,导致规划时间剧增、成功率下降。本文提出一种基于视觉语言模型(VLM)的U-TAMP方法,通过生成物体间的交互属性抽象,刻画接触任务中的抓取与支撑等物理约束,从而增强规划域以应对异质形状、尺寸和材质的物体。在模拟厨房桌台整理场景中,相比原始U-TAMP及前沿的基于常识知识的VLM规划器,本方法在成功率上显著提升,并将规划时间缩短一到两个数量级。
原文摘要 · Abstract (English)
Traditional task-and-motion planning (TAMP) approaches primarily focus on defining sequences of actions along with the necessary geometric and kinematic constraints to execute long-horizon tasks. However, their applicability in real-world settings is limited, as they typically assume simplified object models that overlook key physical properties critical for the successful execution of contact-rich tasks. Moreover, they often use sub-symbolic reasoning during motion planning, which drastically increases planning time and decreases overall success rates. We propose a method that leverages a TAMP approach, defining object-centric abstractions of execution constraints, called Unified TAMP (U-TAMP), to execute robotic tasks involving interactions among objects with heterogeneous shapes, sizes, and materials. Using a Vision-Language Model (VLM), we generate abstractions of inter-object affordances for characterizing physical interaction constraints between objects in contact-rich tasks, such as grasp and support constraints. These constraints are used to enrich the U-TAMP planning domain to deal with objects with variable physical properties. We perform experiments in simulated kitchen table organization scenarios and compare our results with those of the original U-TAMP, as well as a state-of-the-art VLM-based planner that leverages common sense knowledge of objects' affordances for plan generation. Our approach achieves significantly higher planning success rates and improves planning times by one to two orders of magnitude compared to other methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。