arXiv:2602.01960cs.LG2026-02被引 2

用世界模型让生成的视频计划变得可执行,解决动作不连贯问题。

Grounding Generated Videos in Feasible Plans via World Models

  • 通过动作条件世界模型,将视频计划投影到动态可行轨迹空间。
  • 在导航与操作任务中,成功修复了违反物理约束的长时序计划。
  • 适合做视觉规划、机器人控制的研究者,尤其关注可执行性场景。

大规模视频生成模型展现出零样本视觉规划的能力,但生成的计划常违背时间一致性与物理约束,导致映射为可执行动作时失败。为此,我们提出基于世界模型的视频计划接地方法(GVP-WM),该方法利用学习到的动作条件世界模型,将视频生成的计划转化为可行的动作序列。测试时,GVP-WM首先从初始与目标观测生成视频计划,再通过视频引导的隐空间对齐,将视频指导投影至动态可行的隐轨迹流形上。具体而言,将接地建模为一种目标条件的隐空间轨迹优化问题,在世界模型动力学下联合优化隐状态与动作,同时保持与视频生成计划的语义一致性。实验表明,GVP-WM可在导航与操作仿真任务中,从零样本图像到视频生成及运动模糊视频中恢复出可行的长时序计划。

原文摘要 · Abstract (English)

Large-scale video generative models have shown emerging capabilities as zero-shot visual planners, yet video-generated plans often violate temporal consistency and physical constraints, leading to failures when mapped to executable actions. To address this, we propose Grounding Video Plans with World Models (GVP-WM), a planning method that grounds video-generated plans into feasible action sequences using a learned action-conditioned world model. At test-time, GVP-WM first generates a video plan from initial and goal observations, then projects the video guidance onto the manifold of dynamically feasible latent trajectories via video-guided latent collocation. In particular, we formulate grounding as a goal-conditioned latent-space trajectory optimization problem that jointly optimizes latent states and actions under world-model dynamics, while preserving semantic alignment with the video-generated plan. Empirically, GVP-WM recovers feasible long-horizon plans from zero-shot image-to-video-generated and motion-blurred videos that violate physical constraints, across navigation and manipulation simulation tasks.

视频生成视觉规划世界模型机器人控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。