用扩散模型实现跨夹爪的通用抓取,适配不同机械臂设计。
Diffusion for Multi-Embodiment Grasping
- 基于等变扩散模型,编码场景时忽略夹爪形状,解码时融入夹爪几何信息。
- 在多类夹爪上均优于现有方法,包括平行夹爪和人手型夹爪。
- 构建了含杂乱堆叠物体的数据生成框架,提升训练泛化能力。
抓取是机器人学中的基础技能,广泛应用于医疗、工业和家庭场景。然而,当前抓取预测方法通常针对特定夹爪设计,一旦夹爪改变就难以适用。为此,我们探索在不同夹爪设计间迁移抓取策略,实现异构数据源的利用。本文提出一种基于等变扩散模型的方法,能够对包含可抓取物体的场景进行夹爪无关编码,并通过融合夹爪几何信息实现夹爪感知的抓取姿态解码。同时,我们开发了一套数据生成框架,可生成包含不同尺寸物体堆叠的杂乱场景,提升抓取合成方法的训练效果。在多种物体数据集上的实验表明,该方法在从简单平行夹爪到人形手的各类夹爪架构中均表现出优异的泛化性能,优于单夹爪与多夹爪的当前最优方法。
原文摘要 · Abstract (English)
Grasping is a fundamental skill in robotics with diverse applications across medical, industrial, and domestic domains. However, current approaches for predicting valid grasps are often tailored to specific grippers, limiting their applicability when gripper designs change. To address this limitation, we explore the transfer of grasping strategies between various gripper designs, enabling the use of data from diverse sources. In this work, we present an approach based on equivariant diffusion that facilitates gripper-agnostic encoding of scenes containing graspable objects and gripper-aware decoding of grasp poses by integrating gripper geometry into the model. We also develop a dataset generation framework that produces cluttered scenes with variable-sized object heaps, improving the training of grasp synthesis methods. Experimental evaluation on diverse object datasets demonstrates the generalizability of our approach across gripper architectures, ranging from simple parallel-jaw grippers to humanoid hands, outperforming both single-gripper and multi-gripper state-of-the-art methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。