用第一视角人像视频训练机器人动作模型,大幅降低数据采集成本。
EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos
- 从人眼视角视频中学习视觉-语言-动作映射,替代真实机器人采集数据。
- 在双臂操作任务上优于基线模型,仅需少量机器人示范即可适配。
- 适合研究低成本机器人模仿学习的学者与开发者。
真实机器人数据采集推动了机器人操作领域的进步,但依赖硬件设备严重限制了数据规模。本文探索利用第一视角人像视频训练视觉-语言-动作(VLA)模型。相比传统方式,人像视频不仅规模更大,且场景和任务更丰富。基于在人像视频上训练的VLA模型,可预测人类手腕与手部动作,并通过逆运动学与动作重定向将其转换为机器人动作。我们使用少量机器人操作示范对模型进行微调,得到EgoVLA机器人策略。提出仿真基准Ego Humanoid Manipulation Benchmark,设计多样化的双臂操作任务及示范数据。在该基准上微调并评估EgoVLA,结果显示其显著优于基线方法,并验证了人像数据的关键作用。视频演示见官网:https://rchalyang.github.io/EgoVLA。
原文摘要 · Abstract (English)
Real robot data collection for imitation learning has led to significant advancements in robotic manipulation. However, the requirement for robot hardware in the process fundamentally constrains the scale of the data. In this paper, we explore training Vision-Language-Action (VLA) models using egocentric human videos. The benefit of using human videos is not only for their scale but more importantly for the richness of scenes and tasks. With a VLA trained on human video that predicts human wrist and hand actions, we can perform Inverse Kinematics and retargeting to convert the human actions to robot actions. We fine-tune the model using a few robot manipulation demonstrations to obtain the robot policy, namely EgoVLA. We propose a simulation benchmark called Ego Humanoid Manipulation Benchmark, where we design diverse bimanual manipulation tasks with demonstrations. We fine-tune and evaluate EgoVLA with Ego Humanoid Manipulation Benchmark and show significant improvements over baselines and ablate the importance of human data. Videos can be found on our website: https://rchalyang.github.io/EgoVLA
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。