arXiv:2512.06963cs.ROcs.AI2025-12NeurIPS被引 60

用视频生成模型让机器人学会泛化操作,想象未来画面提升动作预测能力

VideoVLA: Video Generators Can Be Generalizable Robot Manipulators

  • 将大模型视频生成能力转化为机器人动作与视觉结果联合预测
  • 在新物体和新场景中实现跨形态技能迁移,任务成功率显著提升
  • 适合研究机器人泛化、具身智能与多模态生成的学者关注

机器人操作的泛化能力对开放世界部署和推动通用人工智能至关重要。尽管近期视觉-语言-动作(VLA)模型利用大规模预训练理解模型进行感知与指令遵循,其在新任务、新物体和新环境下的泛化能力仍有限。本文提出VideoVLA,一种将大型视频生成模型转化为可泛化机器人操作者的简单方法。给定语言指令与图像输入,VideoVLA同时预测动作序列及其未来的视觉结果。基于多模态扩散变换器,VideoVLA联合建模视频、语言与动作模态,利用预训练视频生成模型进行联合视觉与动作预测。实验表明,高质量的未来画面想象与可靠的动作用预测及任务成功高度相关,凸显了视觉想象在操作中的重要性。VideoVLA展现出强泛化能力,包括模仿其他形态的技能并处理新物体。这种同时预测动作及其视觉后果的双重策略,开启机器人学习的新范式,为操作系统的泛化能力提供了新路径。

原文摘要 · Abstract (English)

Generalization in robot manipulation is essential for deploying robots in open-world environments and advancing toward artificial general intelligence. While recent Vision-Language-Action (VLA) models leverage large pre-trained understanding models for perception and instruction following, their ability to generalize to novel tasks, objects, and settings remains limited. In this work, we present VideoVLA, a simple approach that explores the potential of transforming large video generation models into robotic VLA manipulators. Given a language instruction and an image, VideoVLA predicts an action sequence as well as the future visual outcomes. Built on a multi-modal Diffusion Transformer, VideoVLA jointly models video, language, and action modalities, using pre-trained video generative models for joint visual and action forecasting. Our experiments show that high-quality imagined futures correlate with reliable action predictions and task success, highlighting the importance of visual imagination in manipulation. VideoVLA demonstrates strong generalization, including imitating other embodiments' skills and handling novel objects. This dual-prediction strategy - forecasting both actions and their visual consequences - explores a paradigm shift in robot learning and unlocks generalization capabilities in manipulation systems.

机器人操作视频生成多模态泛化能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。