arXiv:2602.10717cs.RO2026-02被引 4

让机器人通过预测视频来精准执行指令操作

Say, Dream, and Act: Learning Video World Models for Instruction-Driven Robot Manipulation

  • 用视频生成模型预测未来动作效果,支持快速决策
  • 生成视频时空一致,任务完成率显著高于基线
  • 适合需要精准视觉-动作对齐的机器人控制场景

机器人操作需要预判环境在动作下的演化,但现有系统大多缺乏这种预测能力,常导致错误和低效。虽然视觉-语言模型(VLMs)能提供高层指导,却无法显式预测未来状态;而现有世界模型或仅能预测短时程,或生成空间不一致的帧。为此,我们提出一种快速、可预测的视频条件动作框架:首先选择并适配一个稳健的视频生成模型以确保可靠未来预测,再通过对抗性蒸馏实现快速、少步数的视频生成,最后训练一个动作模型,利用生成视频与真实观测联合纠正空间误差。大量实验表明,该方法生成的视频在时间上连贯、空间上准确,直接支持精确操作,在具身一致性、空间指代能力和任务完成度方面显著优于现有基线。代码与模型将公开。

原文摘要 · Abstract (English)

Robotic manipulation requires anticipating how the environment evolves in response to actions, yet most existing systems lack this predictive capability, often resulting in errors and inefficiency. While Vision-Language Models (VLMs) provide high-level guidance, they cannot explicitly forecast future states, and existing world models either predict only short horizons or produce spatially inconsistent frames. To address these challenges, we propose a framework for fast and predictive video-conditioned action. Our approach first selects and adapts a robust video generation model to ensure reliable future predictions, then applies adversarial distillation for fast, few-step video generation, and finally trains an action model that leverages both generated videos and real observations to correct spatial errors. Extensive experiments show that our method produces temporally coherent, spatially accurate video predictions that directly support precise manipulation, achieving significant improvements in embodiment consistency, spatial referring ability, and task completion over existing baselines. Codes & Models will be released.

机器人操作视频生成世界模型指令驱动

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。