arXiv:2410.02477cs.ROcs.LG2024-10被引 29

从人类示范中学习多种双手灵巧操作技能,提升机器人泛化能力。

Learning Diverse Bimanual Dexterous Manipulation Skills from Human Demonstrations

  • 基于人类演示构建任务,用师生学习框架统一训练多任务策略。
  • 在TACO数据集上实现74.59%已学任务完成率和51.07%零样本泛化率。
  • 适合研究灵巧操作、人机模仿学习及多任务机器人控制的开发者。

双手灵巧操作是机器人领域关键但研究不足的方向。其高维动作空间与固有任务复杂性给策略学习带来挑战,现有基准任务多样性有限,阻碍通用技能发展。现有方法多依赖强化学习,常受限于针对特定任务设计的复杂奖励函数。本文提出一种新方法,可高效从大量人类示范中学习多样化的双手灵巧操作技能。我们构建了BiDexHD框架,整合现有双手数据集进行任务构造,并采用师生策略学习机制应对所有任务。教师使用通用两阶段奖励函数,在共享行为任务间学习状态驱动策略,学生则将多任务策略蒸馏为视觉输入策略。借助BiDexHD,大规模自动构建任务下的双手灵巧技能学习成为可能,推动通用双手灵巧操作的发展。在包含6类共141个任务的TACO数据集上的实验表明,模型在已训练任务上达成74.59%的任务完成率,在未见任务上达51.07%,验证了方法的有效性与卓越的零样本泛化能力。

原文摘要 · Abstract (English)

Bimanual dexterous manipulation is a critical yet underexplored area in robotics. Its high-dimensional action space and inherent task complexity present significant challenges for policy learning, and the limited task diversity in existing benchmarks hinders general-purpose skill development. Existing approaches largely depend on reinforcement learning, often constrained by intricately designed reward functions tailored to a narrow set of tasks. In this work, we present a novel approach for efficiently learning diverse bimanual dexterous skills from abundant human demonstrations. Specifically, we introduce BiDexHD, a framework that unifies task construction from existing bimanual datasets and employs teacher-student policy learning to address all tasks. The teacher learns state-based policies using a general two-stage reward function across tasks with shared behaviors, while the student distills the learned multi-task policies into a vision-based policy. With BiDexHD, scalable learning of numerous bimanual dexterous skills from auto-constructed tasks becomes feasible, offering promising advances toward universal bimanual dexterous manipulation. Our empirical evaluation on the TACO dataset, spanning 141 tasks across six categories, demonstrates a task fulfillment rate of 74.59% on trained tasks and 51.07% on unseen tasks, showcasing the effectiveness and competitive zero-shot generalization capabilities of BiDexHD. For videos and more information, visit our project page https://sites.google.com/view/bidexhd.

双手操作模仿学习多任务学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。