用生成视频模拟人类抓取,零样本实现高效泛化抓取
Grasp as You Dream: Imitating Functional Grasping from Generated Human Demonstrations
- 用视觉生成模型合成人类抓取动作,替代真实数据采集
- 在多个机器人手上实现比现有方法更优的数据效率与泛化能力
- 适合希望减少实测数据、快速部署抓取任务的研究者
构建能在开放世界中完成功能性抓取的通用机器人仍面临巨大挑战,源于物体与任务的多样性。现有方法或局限于狭窄的物品种类/任务范围,或需耗费大量人力收集大规模数据以覆盖真实世界变化。本文提出GraspDreamer方法,利用预训练于互联网规模人类数据的视觉生成模型(如视频生成模型)合成人类示范,实现无需人工数据采集的零样本功能性抓取。其核心思想是:这类模型隐式编码了人类与物理世界交互的通用先验,结合具身特定的动作优化,可仅以极少努力实现功能性抓取。在多个公开基准上使用不同机械臂的大量实验表明,GraspDreamer相比先前方法具有更优的数据效率和泛化性能。真实机器人验证进一步证明了其有效性。此外,我们还展示该方法可自然扩展至下游操作任务,并能生成数据支持视觉-运动策略学习。
原文摘要 · Abstract (English)
Building generalist robots capable of performing functional grasping in everyday, open-world environments remains a significant challenge due to the vast diversity of objects and tasks. Existing methods are either constrained to narrow object/task sets or rely on prohibitively large-scale data collection to capture real-world variability. In this work, we present an alternative approach, GraspDreamer, a method that leverages human demonstrations synthesized by visual generative models (VGMs) (e.g., video generation models) to enable zero-shot functional grasping without labor-intensive data collection. The key idea is that VGMs pre-trained on internet-scale human data implicitly encode generalized priors about how humans interact with the physical world, which can be combined with embodiment-specific action optimization to enable functional grasping with minimal effort. Extensive experiments on the public benchmarks with different robot hands demonstrate the superior data efficiency and generalization performance of GraspDreamer compared to previous methods. Real-world evaluations further validate the effectiveness on real robots. Additionally, we showcase that GraspDreamer can (1) be naturally extended to downstream manipulation tasks, and (2) can generate data to support visuomotor policy learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。