让机器人在移动中精准抓取任意物体,无需人工标注数据。
Self-Supervised Learning of Grasping Arbitrary Objects On-the-Move
- 将抓取分解为两种抓取动作和一种移动动作,降低控制复杂度。
- 在仿真中随机化物体与环境,实现跨场景的抓取泛化。
- 仅用视觉输入,通过三类卷积网络实时预测抓取与运动调整。
移动抓取通过利用机器人的移动能力提升操作效率。本研究旨在使商用现成机器人实现移动抓取,需精确控制时机与姿态。自监督学习可训练出通用策略,根据目标物体形状和姿态调整机器人速度,并确定抓取位置与方向。由于移动抓取复杂,动作原子化与分步学习对避免试错中的数据稀疏至关重要。本研究将移动抓取简化为两个抓取动作原型和一个移动动作原型,可在机械臂自由度有限的情况下操作。引入三个全卷积神经网络(FCN)模型,分别从视觉输入预测静态抓取、动态抓取和残余移动速度误差。采用两阶段抓取学习方法实现FCN模型的无缝训练。消融实验表明,该方法在抓取准确率和拾取放置效率上均达到最优。此外,在仿真中随机化物体形状与环境,有效实现了可泛化的移动抓取。
原文摘要 · Abstract (English)
Mobile grasping enhances manipulation efficiency by utilizing robots' mobility. This study aims to enable a commercial off-the-shelf robot for mobile grasping, requiring precise timing and pose adjustments. Self-supervised learning can develop a generalizable policy to adjust the robot's velocity and determine grasp position and orientation based on the target object's shape and pose. Due to mobile grasping's complexity, action primitivization and step-by-step learning are crucial to avoid data sparsity in learning from trial and error. This study simplifies mobile grasping into two grasp action primitives and a moving action primitive, which can be operated with limited degrees of freedom for the manipulator. This study introduces three fully convolutional neural network (FCN) models to predict static grasp primitive, dynamic grasp primitive, and residual moving velocity error from visual inputs. A two-stage grasp learning approach facilitates seamless FCN model learning. The ablation study demonstrated that the proposed method achieved the highest grasping accuracy and pick-and-place efficiency. Furthermore, randomizing object shapes and environments in the simulation effectively achieved generalizable mobile grasping.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。