arXiv:2606.08057cs.ROcs.AI2026-06被引 2

仅用一段第一视角视频,让机器人学会复杂抓取操作,无需物体模型。

EgoAERO: Learning Dexterous Manipulation from a Single Egocentric Video without Object Assets

论文配图:EgoAERO: Learning Dexterous Manipulation from a Single Egocentric Video without Object Assets
图 1 · 摘自论文原文
  • 无物体资产条件下,从单段第一视角视频重建手物接触轨迹。
  • 在HOI4D数据集上性能接近有三维模型的基线方法。
  • 适合想用真实人类演示训练机器人抓取能力的研究者。

第一人称RGB-D视频提供了自然的人类灵巧操作示范,但现有数据难以用于机器人学习,因物体姿态、几何形状和接触信息常缺失或需预先扫描物体资产。我们提出EgoAERO,首个无需物体资产即可从单个第一人称RGB-D示范中学习灵巧操作的框架。EgoAERO通过无资产物体跟踪与重建、自适应运动补偿及接触优化,恢复一致的接触手物轨迹,并采用两阶段残差学习将其转化为机器人策略。我们还引入在线质量评估机制,构建了包含430万帧RGB-D图像的大规模第一人称数据集EgoDex-R,用于灵巧操作策略学习。仿真与真实世界实验表明,EgoAERO实现了单示范灵巧操作,且在HOI4D上的下游性能接近基于CAD模型的重建方法。

原文摘要 · Abstract (English)

Egocentric RGB-D videos offer a natural source of human dexterous manipulation demonstrations, but existing data is difficult to use for robot learning because object pose, geometry, and contact information are often missing or require pre-scanned object assets. We present EgoAERO, the first framework that learns dexterous manipulation from a single egocentric RGB-D human demonstration without object assets. EgoAERO reconstructs contact-consistent hand-object trajectories through asset-free object tracking and reconstruction, ego motion compensation, and adaptive contact optimization, then converts them into robot policies using two-stage residual learning. We further introduce an online quality assessment mechanism and construct EgoDex-R, a large-scale egocentric dataset with 4.3M RGB-D frames for dexterous policy learning. Simulation and real-world experiments show that EgoAERO enables single-demonstration dexterous manipulation and achieves downstream performance close to CAD-based reconstructions on HOI4D.

灵巧操作第一视角无资产学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。