用无标签视频自监督预训练机器人动作模型,提升泛化与部署效率
Latent Action Pretraining Through World Modeling
- 通过世界建模从无标签视频中学习抽象动作表示
- 在LIBERO基准和真实场景中优于有监督预训练模型
- 模型轻量高效,支持跨任务跨环境迁移
视觉-语言-动作(VLA)模型在遵循语言指令的机器人操作任务中日益流行。当前最先进的VLA模型如OpenVLA和$π_{0}$依赖大规模人工标注的动作数据集,这些数据通过遥操作收集。近期方法如LAPA和villa-X引入了隐式动作表示,可通过建模帧间抽象视觉变化,在无标签数据上实现自监督预训练。尽管效果显著,但这些方法模型规模庞大,难以在实际场景中部署。本文提出LAWM,一种模型无关的框架,通过世界建模从无标签视频中学习隐式动作表示,视频可来自机器人记录或人类日常操作视频。该框架能跨任务、环境和机器人形态迁移知识。在LIBERO基准和真实设置中,其表现超越基于真实机器人动作预训练的模型及其他类似方法,同时具备高效实用的特点,适合真实应用。
原文摘要 · Abstract (English)
Vision-Language-Action (VLA) models have gained popularity for learning robotic manipulation tasks that follow language instructions. State-of-the-art VLAs, such as OpenVLA and $π_{0}$, were trained on large-scale, manually labeled action datasets collected through teleoperation. More recent approaches, including LAPA and villa-X, introduce latent action representations that enable unsupervised pretraining on unlabeled datasets by modeling abstract visual changes between frames. Although these methods have shown strong results, their large model sizes make deployment in real-world settings challenging. In this work, we propose LAWM, a model-agnostic framework to pretrain imitation learning models in a self-supervised way, by learning latent action representations from unlabeled video data through world modeling. These videos can be sourced from robot recordings or videos of humans performing actions with everyday objects. Our framework is able to transfer learned knowledge across tasks, environments, and embodiments. It outperforms models pretrained with ground-truth robot actions and other similar pretraining methods on the LIBERO benchmark and real-world setup, while being efficient and practical for real-world settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。