用物体属性实现高效新物抓取,仿真训练真实世界表现优。
Attribute-Based Robotic Grasping with Data-Efficient Adaptation
- 通过图文融合与自监督学习,从基础物体学通用属性表征。
- 在仿真中仅用基础物体训练,实测未知物抓取成功率超81%。
- 两种低数据依赖适配方法,适合资源有限的真实机器人部署。
机器人抓取是核心操作任务,但快速教会机器人在杂乱环境中抓取新物体仍具挑战。本文提出端到端编码器-解码器网络,基于物体属性实现数据高效的快速适应。模型先在多种基本物体上预训练,学习通用属性表征;通过门控注意力融合工作区图像与查询文本嵌入,预测实例抓取能力。利用抓取前后物体的持续性进行自监督训练,仅使用颜色和形状各异的基础物体即可完成。为提升泛化能力,提出对抗适配与单次抓取适配两种方法:前者用未标注图像增强数据调节图像编码器,后者通过一次抓取试验更新整个模型。二者均数据高效,显著提升抓取性能。仿真与真实世界实验表明,对未知物体的抓取成功率超过81%,大幅优于多个基线方法。
原文摘要 · Abstract (English)
Robotic grasping is one of the most fundamental robotic manipulation tasks and has been the subject of extensive research. However, swiftly teaching a robot to grasp a novel target object in clutter remains challenging. This paper attempts to address the challenge by leveraging object attributes that facilitate recognition, grasping, and rapid adaptation to new domains. In this work, we present an end-to-end encoder-decoder network to learn attribute-based robotic grasping with data-efficient adaptation capability. We first pre-train the end-to-end model with a variety of basic objects to learn generic attribute representation for recognition and grasping. Our approach fuses the embeddings of a workspace image and a query text using a gated-attention mechanism and learns to predict instance grasping affordances. To train the joint embedding space of visual and textual attributes, the robot utilizes object persistence before and after grasping. Our model is self-supervised in a simulation that only uses basic objects of various colors and shapes but generalizes to novel objects in new environments. To further facilitate generalization, we propose two adaptation methods, adversarial adaption and one-grasp adaptation. Adversarial adaptation regulates the image encoder using augmented data of unlabeled images, whereas one-grasp adaptation updates the overall end-to-end model using augmented data from one grasp trial. Both adaptation methods are data-efficient and considerably improve instance grasping performance. Experimental results in both simulation and the real world demonstrate that our approach achieves over 81% instance grasping success rate on unknown objects, which outperforms several baselines by large margins.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。