arXiv:2504.06084cs.ROcs.CV2025-04被引 14

从第一视角视频学人类精细操作,提升机器人抓取能力

MAPLE: Encoding Dexterous Robotic Manipulation Priors Learned From Egocentric Videos

  • 从第一视角视频中学习物体接触点与手部姿态预测特征
  • 在4个仿真基准和4个新挑战任务中显著提升控制精度
  • 适合研究精细操作机器人或具身智能的开发者

大规模第一视角视频数据集记录了多样场景下的人类活动,为理解需精细操控的物体交互提供了丰富细节。这类复杂精细技能对机器人操作至关重要,但传统数据驱动方法常难以有效建模。为此,我们利用大规模第一视角视频中学到的操作先验,改进精细机器人操作策略的学习。提出MAPLE方法,通过第一视角图像预测物体接触点及接触瞬间的手部姿态,并用这些特征训练下游操作策略。实验表明,MAPLE在4个现有仿真基准及4个新设计的高难度仿真任务(需精细控物与复杂灵巧技能)中均表现优异。真实世界实验进一步验证其有效性,使用17自由度灵巧机械手完成任务,同时在仿真与真实环境中的联合评估在以往工作中较少见。此外,模型在第一视角接触点预测任务上也展现出良好性能,证明其应用价值不仅限于灵巧操作策略学习。

原文摘要 · Abstract (English)

Large-scale egocentric video datasets capture diverse human activities across a wide range of scenarios, offering rich and detailed insights into how humans interact with objects, especially those that require fine-grained dexterous control. Such complex, dexterous skills with precise controls are crucial for many robotic manipulation tasks, yet are often insufficiently addressed by traditional data-driven approaches to robotic manipulation. To address this gap, we leverage manipulation priors learned from large-scale egocentric video datasets to improve policy learning for dexterous robotic manipulation tasks. We present MAPLE, a novel method for dexterous robotic manipulation that learns features to predict object contact points and detailed hand poses at the moment of contact from egocentric images. We then use the learned features to train policies for downstream manipulation tasks. Experimental results demonstrate the effectiveness of MAPLE across 4 existing simulation benchmarks, as well as a newly designed set of 4 challenging simulation tasks requiring fine-grained object control and complex dexterous skills. The benefits of MAPLE are further highlighted in real-world experiments using a 17 DoF dexterous robotic hand, whereas the simultaneous evaluation across both simulation and real-world experiments has remained underexplored in prior work. We additionally showcase the efficacy of our model on an egocentric contact point prediction task, validating its usefulness beyond dexterous manipulation policy learning.

灵巧操作第一视角视频学习机器人控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。