让机器人听懂新物体指令并灵巧抓取,突破以往方法局限。
AffordDexGrasp: Open-set Language-guided Dexterous Grasp with Generalizable-Instructive Affordance
- 用可泛化的物体功能表示桥接语言与动作的鸿沟
- 在仿真和真实世界中对未见物体抓取准确率超前人方法
- 适合需要灵活理解指令的智能机器人场景
语言引导的机器人灵巧抓取使机器人能根据人类指令操作物体。然而,以往数据驱动方法难以理解意图,并在开放集下对未见类别执行抓取。本文提出新任务:开放集语言引导灵巧抓取,发现核心挑战在于高层语言语义与低层机器人动作间的巨大差距。为此,我们提出AffordDexGrasp框架,通过一种新的可泛化-可指导的功能表示来弥合该差距。该功能表示利用物体局部结构与类别无关的语义属性,实现对未见类别的泛化,有效指导灵巧抓取生成。基于此,框架引入以语言为输入的函数流匹配(AFM)生成功能,以及以功能为输入的抓取流匹配(GFM)生成灵巧抓取。为评估,我们构建了开放集桌面语言引导灵巧抓取数据集。仿真与真实世界大量实验表明,本框架在开放集泛化能力上优于所有先前方法。
原文摘要 · Abstract (English)
Language-guided robot dexterous generation enables robots to grasp and manipulate objects based on human commands. However, previous data-driven methods are hard to understand intention and execute grasping with unseen categories in the open set. In this work, we explore a new task, Open-set Language-guided Dexterous Grasp, and find that the main challenge is the huge gap between high-level human language semantics and low-level robot actions. To solve this problem, we propose an Affordance Dexterous Grasp (AffordDexGrasp) framework, with the insight of bridging the gap with a new generalizable-instructive affordance representation. This affordance can generalize to unseen categories by leveraging the object's local structure and category-agnostic semantic attributes, thereby effectively guiding dexterous grasp generation. Built upon the affordance, our framework introduces Affordance Flow Matching (AFM) for affordance generation with language as input, and Grasp Flow Matching (GFM) for generating dexterous grasp with affordance as input. To evaluate our framework, we build an open-set table-top language-guided dexterous grasp dataset. Extensive experiments in the simulation and real worlds show that our framework surpasses all previous methods in open-set generalization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。