arXiv:2604.04502cs.RO2026-04被引 2

用顶级视频模型指导机器人抓取,实现通用操作新路径。

Veo-Act: How Far Can Frontier Video Models Advance Generalizable Robot Manipulation?

  • 用视频模型预测未来动作,逆动力学模型解码成机械臂指令。
  • 零样本下任务级轨迹准确率约70%,但底层控制仍不精确。
  • 分层框架提升指令跟随能力,适合做通用机器人决策的参考。

视频生成模型发展迅速,已展现出对物理动态的深刻理解。本文探究前沿视频模型Veo-3在通用机器人操作中的潜力。我们提出一种零样本方法:Veo-3基于当前机器人观测预测未来图像序列,再由逆动力学模型(IDM)恢复对应动作。IDM仅用随机动作数据训练,无需人工标注或专家示范。核心思想是:若视频模型能生成物理上合理的视觉运动轨迹,那么逆动力学模型可将其转化为可执行的动作。我们在仿真和真实世界中,使用高维灵巧手评估该“Veo-3+IDM”方案。结果表明,得益于视频模型的强大泛化能力,该方法能稳定生成约70%正确级别的任务级轨迹,但低层控制精度不足,难以可靠完成多数任务。为此,我们提出分层框架Veo-Act,以Veo-3作为高层运动规划器,结合视觉-语言-动作(VLA)策略作为低层执行器,显著提升了先进视觉-语言-动作策略的指令遵循性能。总体表明,随着视频生成模型持续进化,其有望成为通用机器人学习的重要组件。

原文摘要 · Abstract (English)

Video generation models have advanced rapidly and are beginning to show a strong understanding of physical dynamics. In this paper, we investigate how far an advanced video generation model such as Veo-3 can support generalizable robotic manipulation. We first study a zero-shot approach in which Veo-3 predicts future image sequences from current robot observations, while an inverse dynamics model IDM recovers the corresponding robot actions. The IDM is trained solely on random-play data, requiring neither human supervision nor expert demonstrations. The key intuition is that, if a video model can generate physically plausible future motions in image space, an IDM can translate those visual trajectories into executable robot actions. We evaluate this "Veo-3+IDM" approach in both simulation and the real world using a high-dimensional dexterous hand. We find that, owing to the strong generalization capability of frontier video models, Veo-3+IDM can consistently generate approximately correct task-level trajectories. However, its low-level control accuracy remains insufficient to solve most tasks reliably. Motivated by this observation, we develop a hierarchical framework, Veo-Act, which uses Veo-3 as a high-level motion planner and a VLA policy as the low-level executor, significantly improving the instruction-following performance of a state-of-the-art vision-language-action policy. Overall, our results suggest that, as video generation models continue to improve, video models can be a valuable component for generalizable robot learning.

机器人视频生成分层控制通用操作

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。