arXiv:2511.02776cs.RO2025-11中稿 · ICML被引 16

用统一视觉-运动编码让机器人跨任务跨设备灵活执行复杂操作

XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations

  • 设计统一视觉-运动码表征,联合建模视觉动态与机械动作
  • 在6种机器人上完成1.4万次真实实验,性能超越现有模型
  • 适合需要多机器人通用智能的工业自动化与服务机器人场景

近期大规模机器人数据集和视觉语言模型的发展推动了视觉-语言-动作(VLA)模型的研究。然而,现有VLA模型仍面临两大挑战:从高维观测生成精确低层动作,以及弥合异构数据源间的领域差异,包括不同机器人形态和人类示范。现有方法通常仅编码视觉动态或机器人动作的潜在变量来指导策略学习,未能充分利用大规模异构数据中的多模态互补知识。本文提出XR-1框架,通过双分支向量量化变分自编码器(VQ-VAE)学习一种名为统一视觉-运动码(UVMC)的离散潜在表示,联合编码视觉动态与机器人运动。UVMC作为观测与动作之间的中间表征,并对齐来自异构数据源的多模态动态信息以捕捉互补知识。为有效利用UVMC,我们提出三阶段训练范式:(i) 自监督学习UVMC,(ii) 在大规模跨形态机器人数据集上进行UVMC引导的预训练,(iii) 针对具体任务的微调。我们在六种不同机器人形态上进行了超过14,000次真实世界实验,覆盖120余项多样化操作任务。XR-1在所有任务中持续优于π₀.₅、π₀、RDT、UniVLA和GR00T-N1.5等先进基线模型,展现出对新物体、背景变化、干扰物及光照变化的强泛化能力。

原文摘要 · Abstract (English)

Recent progress in large-scale robotic datasets and vision-language models (VLMs) has advanced research on vision-language-action (VLA) models. However, existing VLA models still face two fundamental challenges: (i) producing precise low-level actions from high-dimensional observations, (ii) bridging domain gaps across heterogeneous data sources, including diverse robot embodiments and human demonstrations. Existing methods often encode latent variables from either visual dynamics or robotic actions to guide policy learning, but they fail to fully exploit the complementary multi-modal knowledge present in large-scale, heterogeneous datasets. In this work, we present X Robotic Model 1 (XR-1), a novel framework for versatile and scalable VLA learning across diverse robots, tasks, and environments. XR-1 introduces the \emph{Unified Vision-Motion Codes (UVMC)}, a discrete latent representation learned via a dual-branch VQ-VAE that jointly encodes visual dynamics and robotic motion. UVMC addresses these challenges by (i) serving as an intermediate representation between the observations and actions, and (ii) aligning multimodal dynamic information from heterogeneous data sources to capture complementary knowledge. To effectively exploit UVMC, we propose a three-stage training paradigm: (i) self-supervised UVMC learning, (ii) UVMC-guided pretraining on large-scale cross-embodiment robotic datasets, and (iii) task-specific post-training. We validate XR-1 through extensive real-world experiments with more than 14,000 rollouts on six different robot embodiments, spanning over 120 diverse manipulation tasks. XR-1 consistently outperforms state-of-the-art baselines such as $π_{0.5}$, $π_0$, RDT, UniVLA, and GR00T-N1.5 while demonstrating strong generalization to novel objects, background variations, distractors, and illumination changes. Our project is at https://xr-1-vla.github.io/.

机器人多模态通用智能动作生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。