arXiv:2410.23701cs.RO2024-10CoRL被引 17

构建350万抓取数据集,提升多指抓取仿真到现实的迁移能力。

Get a Grip: Multi-Finger Grasp Evaluation at Scale Enables Robust Sim-to-Real Transfer

  • 用百万级真实感抓取样本训练视觉评估模型,实现精准抓取选择。
  • 新数据集支持在模拟和真实场景中均超越传统生成与解析方法。
  • 适合机器人抓取、具身智能研究者,尤其关注仿真到现实迁移的团队。

本文探讨了多指抓取算法实现稳健仿真到现实迁移的条件。尽管已有大量数据集可用于大规模学习生成式多指抓取模型,但实际硬件部署时性能仍普遍下降。另一种策略是基于真实传感器输入,使用判别式抓取评估模型进行抓取选择与优化,该方法在视觉引导的平行爪抓取中已达到领先水平,但在多指场景尚未验证。本文发现现有数据集与方法不足以训练多指抓取的判别模型。为实现大规模训练,数据集需包含约百万级别的正负抓取样本,并提供推理时相似的视觉数据。为此,我们发布了一个开源数据集,包含350万次抓取、4300个物体,标注有RGB图像、点云及训练好的NeRF。基于该数据集,我们训练的视觉抓取评估器在多种模拟与真实场景中表现优于分析与生成模型基线。通过大量消融实验表明,性能关键在于评估器质量,且随数据集缩小而下降,证明了新数据集的重要性。

原文摘要 · Abstract (English)

This work explores conditions under which multi-finger grasping algorithms can attain robust sim-to-real transfer. While numerous large datasets facilitate learning generative models for multi-finger grasping at scale, reliable real-world dexterous grasping remains challenging, with most methods degrading when deployed on hardware. An alternate strategy is to use discriminative grasp evaluation models for grasp selection and refinement, conditioned on real-world sensor measurements. This paradigm has produced state-of-the-art results for vision-based parallel-jaw grasping, but remains unproven in the multi-finger setting. In this work, we find that existing datasets and methods have been insufficient for training discriminitive models for multi-finger grasping. To train grasp evaluators at scale, datasets must provide on the order of millions of grasps, including both positive and negative examples, with corresponding visual data resembling measurements at inference time. To that end, we release a new, open-source dataset of 3.5M grasps on 4.3K objects annotated with RGB images, point clouds, and trained NeRFs. Leveraging this dataset, we train vision-based grasp evaluators that outperform both analytic and generative modeling-based baselines on extensive simulated and real-world trials across a diverse range of objects. We show via numerous ablations that the key factor for performance is indeed the evaluator, and that its quality degrades as the dataset shrinks, demonstrating the importance of our new dataset. Project website at: https://sites.google.com/view/get-a-grip-dataset.

机器人抓取仿真到现实视觉评估多指操作

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。