arXiv:2601.15284cs.CV2026-01被引 10

让视频模型学会按动作精准预测世界变化,还能在画里导航和操作。

Walk through Paintings: Egocentric World Models from Internet Priors

  • 用轻量级条件层注入动作指令,复用互联网视频模型的世界先验。
  • 在机器人导航与操作任务中,物理一致性提升65%,效果超越现有方法。
  • 适用于不同机器人形态,能泛化到未见环境,如画中世界内的任务。

如果视频生成模型不仅能想象出合理的未来,还能准确反映动作带来的世界变化,会怎样?我们提出无偏见的视角世界模型(EgoWM),一种无需从头训练的简单方法,可将任意预训练视频扩散模型转化为受动作控制的世界模型,实现精确可控的未来预测。通过轻量级条件层注入压缩后的运动指令,复用互联网规模视频模型丰富的世界先验,使模型能忠实跟随动作,同时保持泛化性和真实性。该方法自然适配不同身体形态与动作空间——从3自由度移动机器人到25自由度人形机器人,其中基于本体感觉的关节角度动态预测更具挑战性。模型在导航与操作任务中生成连贯轨迹,仅需少量微调。为独立评估物理正确性,我们引入结构一致性评分(SCS),衡量稳定场景元素是否随动作合理演化。相比之前最优方法(导航世界模型),我们的方法在SCS上提升最高达65%;可无缝应用于三种不同的视频扩散模型架构;有效利用互联网先验,泛化至未见过的环境,包括画内导航与操作。最后,我们展示了EgoWM在机器人规划中的应用潜力。

原文摘要 · Abstract (English)

What if a video generation model could not only imagine a plausible future, but the correct one -- accurately reflecting how the world changes with each action? We answer this by presenting the Egocentric World Model (EgoWM), a simple, architecture-agnostic method that transforms any pre-trained video diffusion model into an action-conditioned world model, enabling precisely controllable future prediction. Rather than training from scratch, we repurpose the rich world priors of Internet-scale video models by injecting appropriately compressed motor commands through lightweight conditioning layers. This allows our model to follow actions faithfully while preserving generalization and realism. Our approach scales naturally across embodiments and action spaces -- from 3-DoF mobile robots to 25-DoF humanoids, where predicting egocentric joint-angle-driven dynamics is substantially more challenging. The model produces coherent rollouts for both navigation and manipulation, requiring only modest fine-tuning. To evaluate physical correctness independent of appearance, we introduce the Structural Consistency Score (SCS), which measures whether stable scene elements evolve consistently with the provided actions. Our method improves SCS by up to 65\% over the prior state of the art, Navigation World Models; applies seamlessly to three different video diffusion model architectures; and effectively utilizes Internet priors to generalize to unseen environments, including navigation and manipulation inside paintings. Finally, we demonstrate the applicability of EgoWM to robotic planning.

世界模型视频生成机器人规划动作控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。