从网络图片学习真实抓握动作,让机器人更灵活地操作物体。
Web2Grasp: Learning Functional Grasps from Web Images of Hand-Object Interactions
- 用网页图像重建3D手物交互,提取自然抓握模式。
- 在仿真中抓取成功率75.8%,能泛化到未见物体。
- 适合想提升机器人灵巧操作能力的研究者。
功能性抓握对多指机器人手有效操作物体至关重要。以往工作主要关注仅持物的力度抓握,或依赖特定物体的域内示范。本文提出利用从网络图像中提取的人类抓握信息,这些图像捕捉了自然且功能性的手物交互(HOI)。通过预训练的3D重建模型,从RGB图像中恢复3D人类手物交互网格。为在噪声数据上训练,我们提出:(1) 以交互为中心的模型,学习手与物体间的功能性交互模式;(2) 基于几何的过滤策略去除不可行抓握,结合物理模拟保留能抗扰动的抓握。在IssacGym仿真中,基于重建的HOI抓握训练的模型,在网络数据集物体上达到75.8%的成功率,并可泛化至未见物体,优于基线方法在抓取成功率与功能质量上的表现。在真实世界实验中,使用LEAP手和Inspire手,在12种物体(包括注射器、喷雾瓶、刀具、镊子等挑战性物体)上实现77.5%的成功率。项目网站:https://web2grasp.github.io/。
原文摘要 · Abstract (English)
Functional grasping is essential for enabling dexterous multi-finger robot hands to manipulate objects effectively. Prior work largely focuses on power grasps, which only involve holding an object, or relies on in-domain demonstrations for specific objects. We propose leveraging human grasp information extracted from web images, which capture natural and functional hand-object interactions (HOI). Using a pretrained 3D reconstruction model, we recover 3D human HOI meshes from RGB images. To train on these noisy HOI data, we propose to use: (1) an interaction-centric model to learn the functional interaction pattern between hand and object, and (2) geometry-based filtering to remove the infeasible grasps and physical simulation to retain grasps who can resist disturbance. In IssacGym simulation, our model trained on reconstructed HOI grasps achieves a 75.8% success rate on objects from the web dataset and generalizes to unseen objects, outperforming baseline methods in both grasp success and functional quality. In real-world experiments with the LEAP hand and Inspire hand, it attains a 77.5% success rate across 12 objects, including challenging ones such as a syringe, spray bottle, knife, and tongs. Project website is at: https://web2grasp.github.io/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。