arXiv:2409.08273cs.ROcs.AI2024-09ICRA被引 56

从真实视频中学习通用机器人操作先验,提升下游任务适应效率与鲁棒性。

Hand-Object Interaction Pretraining from Videos

  • 通过3D空间对齐人类手物交互轨迹,将人动迁移到机器人动作
  • 生成的任务无关基础策略实现样本高效微调,性能优于以往方法
  • 适合需要快速适配新任务的机器人研发人员使用

我们提出一种从真实世界视频中学习通用机器人操作先验的方法,基于3D手物交互轨迹构建框架。通过在共享3D空间中重构人体手部与物体运动,并将人类动作重定向为机器人动作,生成传感器-运动轨迹数据。在此基础上进行生成建模,得到一个任务无关的基础策略,该策略捕捉了通用且灵活的操作先验。实验证明,采用强化学习(RL)和行为克隆(BC)微调该策略,可实现对下游任务的样本高效适应,同时在鲁棒性和泛化能力上优于现有方法。定性实验结果见: https://hgaurav2k.github.io/hop/

原文摘要 · Abstract (English)

We present an approach to learn general robot manipulation priors from 3D hand-object interaction trajectories. We build a framework to use in-the-wild videos to generate sensorimotor robot trajectories. We do so by lifting both the human hand and the manipulated object in a shared 3D space and retargeting human motions to robot actions. Generative modeling on this data gives us a task-agnostic base policy. This policy captures a general yet flexible manipulation prior. We empirically demonstrate that finetuning this policy, with both reinforcement learning (RL) and behavior cloning (BC), enables sample-efficient adaptation to downstream tasks and simultaneously improves robustness and generalizability compared to prior approaches. Qualitative experiments are available at: \url{https://hgaurav2k.github.io/hop/}.

机器人操作视频预训练迁移学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。