用视频学通用机器人动作,跨平台部署更高效。
UniVLA: Learning to Act Anywhere with Task-centric Latent Actions

- 从视频中提取任务导向的隐式动作表示,支持多机器人形态和视角
- 仅需不到1/20的预训练算力即超越OpenVLA,下游数据少1/10
- 兼容人类视频等异构数据,适合大规模通用机器人策略训练
通用机器人需在多种环境中有效执行任务。然而,现有方法严重依赖标注动作数据,通常局限于单一物理形态,难以跨不同机器人和环境迁移知识。为此,我们提出UniVLA框架,学习跨形态的视觉-语言-动作(VLA)策略。核心创新在于通过隐式动作模型从视频中提取任务导向的动作表示,从而利用涵盖多种机器人形态和视角的大规模数据。为减少任务无关动态干扰,我们在DINO特征空间中引入语言指令构建隐式动作模型。该策略基于互联网规模视频训练,可通过高效的隐式动作解码部署至各类机器人。在多个操作与导航基准测试及真实机器人部署中,UniVLA取得当前最优性能,其表现优于OpenVLA,且预训练算力不足其1/20,下游数据仅为1/10。随着异构数据(包括人类视频)的引入,性能持续提升,验证了其在可扩展、高效机器人策略学习中的潜力。
原文摘要 · Abstract (English)
A generalist robot should perform effectively across various environments. However, most existing approaches heavily rely on scaling action-annotated data to enhance their capabilities. Consequently, they are often limited to single physical specification and struggle to learn transferable knowledge across different embodiments and environments. To confront these limitations, we propose UniVLA, a new framework for learning cross-embodiment vision-language-action (VLA) policies. Our key innovation is to derive task-centric action representations from videos with a latent action model. This enables us to exploit extensive data across a wide spectrum of embodiments and perspectives. To mitigate the effect of task-irrelevant dynamics, we incorporate language instructions and establish a latent action model within the DINO feature space. Learned from internet-scale videos, the generalist policy can be deployed to various robots through efficient latent action decoding. We obtain state-of-the-art results across multiple manipulation and navigation benchmarks, as well as real-robot deployments. UniVLA achieves superior performance over OpenVLA with less than 1/20 of pretraining compute and 1/10 of downstream data. Continuous performance improvements are observed as heterogeneous data, even including human videos, are incorporated into the training pipeline. The results underscore UniVLA's potential to facilitate scalable and efficient robot policy learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。