让机器人从人类视频中学习动作,规模够大时自动实现跨体态迁移。
Emergence of Human to Robot Transfer in Vision-Language-Action Models
- 用人类视频与机器人数据联合训练,无需手动映射。
- 预训练越多样,机器人在人类动作上表现提升近一倍。
- 适合做通用机器人技能学习的研究者参考。
视觉-语言-动作(VLA)模型具备广泛的开放世界泛化能力,但需要大量多样的数据。利用涵盖丰富真实场景的人类视频数据是理想选择,因其易获取且覆盖广泛。然而,仅用人类视频训练VLA困难,且建立人与机器人之间的映射需人工工程,仍是重大挑战。受大规模语言模型随规模增长涌现学习能力的启发,我们探究类似现象是否在包含人类视频的VLA中出现。提出一种简单共训练方法,发现当模型在足够多场景、任务和机器人形态上预训练后,人类到机器人的技能迁移能力会自然涌现。分析表明,这是由于多样化预训练产生了对身体形态无关的表示。通过一系列实验验证,发现充分多样化的机器人预训练下,该方法在仅见于人类数据的泛化设置上性能几乎翻倍。
原文摘要 · Abstract (English)
Vision-language-action (VLA) models can enable broad open world generalization, but require large and diverse datasets. It is appealing to consider whether some of this data can come from human videos, which cover diverse real-world situations and are easy to obtain. However, it is difficult to train VLAs with human videos alone, and establishing a mapping between humans and robots requires manual engineering and presents a major research challenge. Drawing inspiration from advances in large language models, where the ability to learn from diverse supervision emerges with scale, we ask whether a similar phenomenon holds for VLAs that incorporate human video data. We introduce a simple co-training recipe, and find that human-to-robot transfer emerges once the VLA is pre-trained on sufficient scenes, tasks, and embodiments. Our analysis suggests that this emergent capability arises because diverse pretraining produces embodiment-agnostic representations for human and robot data. We validate these findings through a series of experiments probing human to robot skill transfer and find that with sufficiently diverse robot pre-training our method can nearly double the performance on generalization settings seen only in human data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。