用人类抓取数据训练机器人,实现零样本日常物品抓取。
Human Universal Grasping

- 基于人类抓取行为构建流匹配模型,生成自然多样化的抓取姿势。
- 在90个未见物体上比顶尖方法提升23%~34%成功率。
- 适合做具身智能、机器人抓取和人机交互的研究者参考。
人类能轻松抓取各类物品,而多指机器人仍远不及此。我们认为,最自然的机器人抓取数据来源于人类日常抓取行为。本文提出HUG,一个基于流匹配的模型,仅需单张RGB-D图像即可为任意指定物体生成多样化的人类抓取姿态。通过智能眼镜,我们采集了1M-HUG数据集,包含100万帧(27.8小时)的视角数据,覆盖41栋建筑中的6,707种物体实例。为建模自然人类抓取分布,我们的新型流匹配模型融合RGB与深度信息,输出由腕部平移、旋转及MANO手部姿态参数化的抓取动作。预测抓取可零样本适配多种机器人手型,实现在真实家庭场景下的即时抓取。为标准化评估,我们构建了新基准HUG-Bench,包含5种几何类别、不同尺寸的90个未见物体,使用米级精度3D网格。我们在多个立体相机、机器人形态和家庭环境中对HUG-Bench的30个测试物体进行真实世界验证,结果表明其在挑战性物体集上相比现有最优方法分别提升23%和34%。代码、数据、基准、模型权重及交互演示已公开于https://grasping.io/
原文摘要 · Abstract (English)
Humans can grasp objects effortlessly, whereas multi-fingered robots are far from this level of generality. We argue that the most natural source of robot grasping data is from humans, who pick up thousands of objects every day. We present HUG, a flow-matching model that generates diverse human grasps for any user-specified object in a single RGB-D image captured from a stereo camera. Using smart glasses, we first collect 1M-HUGs, an egocentric dataset of human grasps spanning 1M frames (27.8 hrs) and 6,707 object instances across 41 buildings. Next, to model the distribution of natural human grasps, our novel flow-matching model fuses RGB and depth observations to output a grasp parameterized by wrist translation, wrist rotation, and MANO hand pose. Predicted grasps can be retargeted to various robot hands, enabling zero-shot grasping in everyday scenes. To standardize evaluation, we build a new simulated benchmark, HUG-Bench, of 90 unseen objects from five geometric categories and various sizes, with metric-scale 3D meshes. We evaluate HUG in the real world on the 30-object test set of HUG-Bench across multiple stereo cameras, robot embodiments, and household environments. HUG outperforms the state-of-the-art grasping baselines by +23% and +34% on our challenging object set. Code, data, benchmark, checkpoints, and an interactive demo are released on our website: https://grasping.io/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。