arXiv:2606.08107cs.ROcs.AI2026-06被引 1

用人类视角数据训练机器人,让机械手学会新任务和组合技能。

Ego-Pi: VLA Fine-Tuning for Ego-Centric Human and Robot Data

论文配图:Ego-Pi: VLA Fine-Tuning for Ego-Centric Human and Robot Data
图 1 · 摘自论文原文
  • 基于π₀.₅模型,融合人类与机器人五指手数据进行微调。
  • 仅用人类数据即可让机器人掌握新任务语义并组合已有技能。
  • 适合做具身智能、人机共融系统研究的开发者参考。

机器人领域面临数据稀缺的根本挑战。与语言或视觉研究不同,缺乏互联网规模的机器人操作数据集。一个有前景的方向是利用更易获取、覆盖更广且规模更大的人类视角数据。本文以π₀.₅模型为基础,研究了在配备灵巧五指手的人类与类人机器人之间跨体感学习的关键设计选择。结果表明,仅使用人类数据,机器人即可学会新的任务语义,并将已有技能组合成新颖行为,而无需对应机器人数据。

原文摘要 · Abstract (English)

Robotics faces a fundamental challenge of data scarcity. Unlike language or vision research, there is no internet-scale dataset for robotic manipulation. A promising path forward is to leverage egocentric human data, which can be collected more easily, with greater breadth, and at a larger scale. Towards this end, we investigate key design choices for learning across human and humanoid embodiments equipped with dexterous five-finger hands, using the $π_{0.5}$ model as a foundation. Our results show that human data enables robots to learn new task semantics and compose existing skills into novel behaviors without corresponding robot data. The paper website is here: https://egopipaper.github.io/

机器人多模态具身智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。