arXiv:2509.15212cs.CVcs.RO2025-09被引 24

用人类操作视频训练机器人视觉语言动作模型,提升操控能力。

RynnVLA-001: Using Human Demonstrations to Improve Robot Manipulation

  • 两阶段预训练:先预测未来画面,再联合预测关键点轨迹。
  • 在1200万段第一视角操作视频上训练,显著提升下游任务表现。
  • 适合研究机器人视觉导航与具身智能的学者参考。

本文提出RynnVLA-001,一种基于大规模人类操作视频生成预训练的视觉-语言-动作(VLA)模型。我们设计了一种新型两阶段预训练方法:第一阶段为第一视角视频生成预训练,在1200万段第一视角操作视频上训练图像到视频模型,以初始帧和语言指令为条件预测未来帧;第二阶段为人本轨迹感知建模,联合预测未来关键点轨迹,有效连接视觉帧预测与动作预测。此外,为增强动作表示,提出ActionVAE——一种变分自编码器,将动作序列压缩为紧凑的潜在嵌入,降低VLA输出空间复杂度。在相同下游机器人数据集上微调后,RynnVLA-001性能超越现有最先进基线,证明所提预训练策略能为VLA模型提供更优初始化。

原文摘要 · Abstract (English)

This paper presents RynnVLA-001, a vision-language-action(VLA) model built upon large-scale video generative pretraining from human demonstrations. We propose a novel two-stage pretraining methodology. The first stage, Ego-Centric Video Generative Pretraining, trains an Image-to-Video model on 12M ego-centric manipulation videos to predict future frames conditioned on an initial frame and a language instruction. The second stage, Human-Centric Trajectory-Aware Modeling, extends this by jointly predicting future keypoint trajectories, thereby effectively bridging visual frame prediction with action prediction. Furthermore, to enhance action representation, we propose ActionVAE, a variational autoencoder that compresses sequences of actions into compact latent embeddings, reducing the complexity of the VLA output space. When finetuned on the same downstream robotics datasets, RynnVLA-001 achieves superior performance over state-of-the-art baselines, demonstrating that the proposed pretraining strategy provides a more effective initialization for VLA models.

机器人操控视觉语言动作预训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。