用抓握数据预训练机器人,让其学会操作带活动部件的工具。
From Grasps to Dexterity: Large-Scale Grasp Pretraining for Dexterous Manipulation

- 分层模仿学习:高层预测手部目标,底层控制器执行动作。
- 35.5万条轨迹预训练,实测任务成功率提升33.3个百分点。
- 适合需要精细触觉交互的复杂操作任务研究者。
大规模灵巧抓握数据集蕴含丰富的手物交互先验,但以往多用于生成抓取或拾放操作。本文探索其在具身工具使用中的功能灵活性潜力——机器人需抓取工具、保持接触并操控其可动部件。我们采用分层模仿学习框架,结合高层手部子目标预测与低层条件化控制器。基于大规模灵巧抓握标注构建了包含35.5万条轨迹的预训练数据集,并用于预训练底层控制器,再在下游任务演示上微调。为评估该方法,提出DexCraft仿真基准,涵盖六种需协调指节运动的工具操作任务。在仿真与真实世界实验中,本方法优于端到端扩散策略基线及从零开始训练的分层策略。真实场景下,任务成功率相较DP3提升33.3个百分点。结果表明,抓握数据不仅能支持抓取生成,还可作为接触密集型灵巧操作的可扩展预训练资源。
原文摘要 · Abstract (English)
Large-scale dexterous grasp datasets encode rich priors over hand-object interaction, but their use has largely been confined to grasp generation and pick-and-place manipulation. We study whether such data can instead support functional dexterity in articulated tool use, where a robot must acquire a tool, maintain contact, and operate its functional moving parts. We adapt a hierarchical imitation learning framework that combines high-level hand sub-goal prediction with a low-level goal-conditioned controller. We construct a 355k-trajectory grasp-pretraining dataset from large-scale dexterous grasp annotations and use it to pretrain the low-level controller. The controller is then fine-tuned on downstream task demonstrations. To evaluate this setting, we introduce DexCraft, a simulation benchmark with six articulated tool-use tasks requiring coordinated finger motion. Across simulation and real-world experiments, our approach outperforms end-to-end diffusion policy baselines and hierarchical policies trained from scratch. In the real world, it improves full-task success by 33.3 percentage points over DP3. These results show that grasp datasets can serve not only as resources for grasp synthesis, but also as scalable pretraining data for contact-rich dexterous manipulation. Videos are shown on https://yingyuan0414.github.io/grasp2dexterity/ .
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。