arXiv:2602.21736cs.RO2026-02被引 12

用联合对齐的隐式动作表示,从海量人类操作视频中高效预训练机器人视觉语言动作模型。

Joint-Aligned Latent Action: Towards Scalable VLA Pretraining in the Wild

  • 通过隐式动作嵌入实现动作与逆动力学对齐,无需重建完整视觉动态。
  • 在750万段视频(超2000小时)上训练,生成更真实的双手运动轨迹。
  • 适合需要大规模人类示范数据的机器人动作学习研究者使用。

尽管已有进展,视觉-语言-动作模型(VLAs)仍受限于大规模、多样化的机器人数据稀缺。人类操作视频是丰富替代数据源,但现有方法只能在小规模精标注数据和海量但手部追踪标签不可靠的自然视频间二选一。我们提出JALA,一种联合对齐的隐式动作预训练框架。JALA不进行完整的视觉动态重建,而是学习一个与逆动力学和真实动作均对齐的预测性动作嵌入,构建出具备时序感知、以行为为中心的隐空间,可直接用于异构人类数据学习。我们通过UniHand-Mix——一个包含750万段视频(超过2000小时)的混合数据集,融合实验室与自然场景数据,实现了该方法的规模化应用。实验表明,JALA在控制环境和无约束场景下均能生成更逼真的手部动作,在仿真与真实机器人操作任务中显著提升下游性能。结果表明,联合对齐的隐式动作为从人类数据中可扩展地预训练VLA提供了有效路径。

原文摘要 · Abstract (English)

Despite progress, Vision-Language-Action models (VLAs) are limited by a scarcity of large-scale, diverse robot data. While human manipulation videos offer a rich alternative, existing methods are forced to choose between small, precisely-labeled datasets and vast in-the-wild footage with unreliable hand tracking labels. We present JALA, a pretraining framework that learns Jointly-Aligned Latent Actions. JALA bypasses full visual dynamic reconstruction, instead learns a predictive action embedding aligned with both inverse dynamics and real actions. This yields a transition-aware, behavior-centric latent space for learning from heterogeneous human data. We scale this approach with UniHand-Mix, a 7.5M video corpus (>2,000 hours) blending laboratory and in-the-wild footage. Experiments demonstrate that JALA generates more realistic hand motions in both controlled and unconstrained scenarios, significantly improving downstream robot manipulation performance in both simulation and real-world tasks. These results indicate that jointly-aligned latent actions offer a scalable pathway for VLA pretraining from human data.

视觉语言动作机器人学习隐式动作大规模预训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。