arXiv:2410.11758cs.ROcs.CL2024-10ICLR被引 326

无需机器人动作标签,用网络视频预训练通用动作模型

Latent Action Pretraining from Videos

  • 通过VQ-VAE学习图像帧间的离散潜在动作
  • 在真实任务中超越使用机器人标签的SOTA模型
  • 适合想用海量网络视频训练机器人基础模型的研究者

我们提出无监督的潜在动作预训练方法(LAPA),用于训练视觉-语言-动作(VLA)模型,无需依赖人工遥控器采集的机器人动作标签。现有VLA模型需大量带动作标签的机器人数据,限制了数据规模与来源。本文方法利用互联网级视频(无机器人动作标注)进行预训练:先用基于VQ-VAE的目标学习图像帧间的离散潜在动作;再训练一个潜在VLA模型,从观测和任务描述中预测这些潜在动作;最后在小规模机器人操作数据上微调,将潜在动作映射到实际机器人动作。实验表明,该方法显著优于仅用大规模视频训练的现有技术,在需要语言条件、泛化至未见物体及未见指令的真实世界操作任务中也优于使用机器人标签的SOTA模型。仅使用人类操作视频训练也展现出正向迁移,为构建基于网络规模数据的机器人基础模型提供了可能。

原文摘要 · Abstract (English)

We introduce Latent Action Pretraining for general Action models (LAPA), an unsupervised method for pretraining Vision-Language-Action (VLA) models without ground-truth robot action labels. Existing Vision-Language-Action models require action labels typically collected by human teleoperators during pretraining, which significantly limits possible data sources and scale. In this work, we propose a method to learn from internet-scale videos that do not have robot action labels. We first train an action quantization model leveraging VQ-VAE-based objective to learn discrete latent actions between image frames, then pretrain a latent VLA model to predict these latent actions from observations and task descriptions, and finally finetune the VLA on small-scale robot manipulation data to map from latent to robot actions. Experimental results demonstrate that our method significantly outperforms existing techniques that train robot manipulation policies from large-scale videos. Furthermore, it outperforms the state-of-the-art VLA model trained with robotic action labels on real-world manipulation tasks that require language conditioning, generalization to unseen objects, and semantic generalization to unseen instructions. Training only on human manipulation videos also shows positive transfer, opening up the potential for leveraging web-scale data for robotics foundation model.

动作预训练视觉语言动作无监督学习机器人基础模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。