arXiv:2606.06627cs.ROcs.AI2026-06被引 1

用日常视频训练机器人抓取,关键在手部姿态与动作适配。

What Matters When Cotraining Robot Manipulation Policies on Everyday Human Videos?

论文配图:What Matters When Cotraining Robot Manipulation Policies on Everyday Human Videos?
图 1 · 摘自论文原文
  • 基于532段真实视频,用高精度手部姿态提升训练效果。
  • 手部准确度高但动作差异大时,成功率仍低,需适配身体形态。
  • 新方法在6个任务中提升29.7%成功率,适合数据少的机器人训练。

用于协同训练机器人操作策略的人类视频数据集多为精心编排的示范视频,动作设计接近机器人行为,且用手部捕捉设备获取3D手姿。而更丰富的数据来源是日常互联网视频,但其能否有效迁移至机器人仍存疑问。本文构建了一个包含532段人类视频、总计28小时高质量三角化手部标注和自然动作的新数据集。研究发现,手部姿态质量影响迁移效果,但即便手部姿态准确,若视觉与策略网络未针对具体身体形态进行专化,仍存在动作差距导致性能受限。通过优化协同训练方案,在六个操作任务中,低机器人数据场景下成功率绝对提升29.7%。

原文摘要 · Abstract (English)

Human video datasets used for cotraining robot manipulation policies largely consist of curated demonstrations where motions are orchestrated to resemble robot behavior and 3D hand poses are captured with specialized hardware. A more plentiful source of data is everyday Internet video, but it is an open question what factors enable transfer from such videos to robots. We investigate this using a new dataset of 532 human videos with 28 hours of high-quality triangulated hand labels and natural motions. We find that hand pose quality affects transfer, but even with accurate hands, the inherent motion gap hinders transfer unless the vision and policy networks specialize to each embodiment. Our cotraining recipe yields consistent improvements, with an absolute success rate gain of $29.7\%$ in the low-robot-data regime across six manipulation tasks.

机器人操控视频迁移手部姿态数据高效

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。