arXiv:2511.17502cs.RO2025-11被引 56

一个能看、会想、能行动的统一智能体,让机器人更懂环境。

RynnVLA-002: A Unified Vision-Language-Action and World Model

  • 视觉语言动作与世界模型一体化设计,双向增强理解与决策。
  • 仿真任务成功率97.4%,真实机器人任务成功率提升50%。
  • 适合研究具身智能、机器人规划与多模态模型融合的学者。

我们提出 RynnVLA-002,一个统一的视觉-语言-动作(VLA)与世界模型。该世界模型利用动作和视觉输入预测未来图像状态,学习环境底层物理规律以优化动作生成;而 VLA 模型则从图像观测中生成后续动作,提升视觉理解并支持世界模型的图像生成。这种统一框架实现了环境动态与动作规划的联合学习。实验表明,RynnVLA-002 超越独立的 VLA 与世界模型,展现相互增强效果。我们在仿真与真实机器人任务中评估了该模型:在无预训练条件下,于 LIBERO 仿真基准上达到 97.4% 的成功率;在真实世界的 LeRobot 实验中,集成的世界模型使整体成功率提升 50%。

原文摘要 · Abstract (English)

We introduce RynnVLA-002, a unified Vision-Language-Action (VLA) and world model. The world model leverages action and visual inputs to predict future image states, learning the underlying physics of the environment to refine action generation. Conversely, the VLA model produces subsequent actions from image observations, enhancing visual understanding and supporting the world model's image generation. The unified framework of RynnVLA-002 enables joint learning of environmental dynamics and action planning. Our experiments show that RynnVLA-002 surpasses individual VLA and world models, demonstrating their mutual enhancement. We evaluate RynnVLA-002 in both simulation and real-world robot tasks. RynnVLA-002 achieves 97.4% success rate on the LIBERO simulation benchmark without pretraining, while in real-world LeRobot experiments, its integrated world model boosts the overall success rate by 50%.

机器人多模态世界模型强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。