无需示范,机器人自主学会复杂精细操作任务。
CoDex: Learning Compositional Dexterous Functional Manipulation without Demonstrations

- 用视觉语言模型解析任务语义,生成可执行的抓握约束
- 结合优化与强化学习,实现从仿真到真实世界的策略迁移
- 在6类未见工具上成功完成抓取-移动-触发全流程操作
本文研究组合式灵巧功能操作(CD-FOM):如对植物喷洒喷雾瓶或在木头上使用热熔胶枪等任务,需同时控制物体姿态并触发其内部机制。这类任务对机器人要求高,需融合对物体功能、操作方式和应用区域的语义理解,以及复杂的物理灵巧性以维持抓握稳定、规划运动轨迹并完成触发。我们提出CoDex,一种零示范框架,可自主发现CD-FOM操作策略。CoDex利用视觉语言模型(VLMs)从任务与场景中推断语义约束,指导解析式约束优化,生成少量功能性抓握候选,再通过强化学习高效优化为完整抓取-移动-触发策略,可在仿真与现实间迁移。我们在7自由度机械臂搭配16自由度多指手的系统上,对六种含内部机构的未见物体(包括喷雾瓶、热熔胶枪、气吹罐、手电筒、胡椒研磨器)及其未见目标对象进行评估,验证了其无需人工示范即可自主发现并执行复杂、物理可行的灵巧行为的能力。
原文摘要 · Abstract (English)
In this work, we study Compositional Dexterous Functional Object Manipulation (CD-FOM): tasks such as aiming and actuating a spray bottle on a plant or a glue gun on wood, which require both actuating an object's internal mechanism and controlling its pose to apply the object's function to the environment. These tasks pose significant challenges for robots due to the demanding integration of semantic understanding of the object's function, actuation mode, and application area with intricate physical dexterity to manage grasp stability, movement trajectory, and actuation. We introduce CoDex, a zero-demonstration framework that autonomously discovers CD-FOM manipulation strategies. CoDex uses vision-language models (VLMs) to infer semantic constraints from the task and scene. These constraints guide analytic constrained optimization to generate a short list of functional grasp candidates that can be efficiently refined with reinforcement learning to generate full grasp-move-actuate policies transferable from simulation to the real world. We evaluate CoDex on a 7-DoF robot arm with a 16-DoF multi-fingered hand across six CD-FOM tasks involving previously unseen objects with internal mechanisms, including spray bottles, hot glue guns, air dusters, flashlights, and pepper grinders, and their application to unseen target objects, showcasing its ability to autonomously discover and execute complex, physically viable dexterous behaviors without human demonstrations. More information at https://robin-lab.cs.utexas.edu/CoDex/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。