arXiv:2502.07825cs.CVcs.AI2025-02AAAI被引 50

让预训练视频生成模型能模拟动态交互,实现可控动作演化。

Pre-Trained Video Generative Models as World Simulators

  • 用轻量模块让模型响应动作指令,精准对齐视觉变化
  • 引入运动强化损失,提升动作可控性与动态一致性
  • 适用于扩散与Transformer模型,适合机器人和强化学习

大规模互联网数据预训练的视频生成模型在生成逼真合成视频方面表现优异,但通常依赖静态提示(如文本或图像),难以建模交互式动态场景。本文提出动态世界模拟(DWS),将预训练视频生成模型转化为可控制的世界模拟器,能够执行指定的动作轨迹。为实现动作与视觉变化的精确对齐,我们引入一个轻量级、通用的动作条件模块,可无缝集成到任意现有模型中。我们发现,一致的动态过渡建模是构建强大世界模拟器的关键,而非复杂视觉细节。基于此,进一步提出运动强化损失,增强模型捕捉动态变化的能力。实验表明,DWS可广泛应用于扩散模型与自回归Transformer模型,在游戏与机器人领域显著提升动作可控性和动态一致性。此外,为支持下游任务如基于模型的强化学习,我们提出优先想象策略,提升采样效率,性能媲美当前最优方法。

原文摘要 · Abstract (English)

Video generative models pre-trained on large-scale internet datasets have achieved remarkable success, excelling at producing realistic synthetic videos. However, they often generate clips based on static prompts (e.g., text or images), limiting their ability to model interactive and dynamic scenarios. In this paper, we propose Dynamic World Simulation (DWS), a novel approach to transform pre-trained video generative models into controllable world simulators capable of executing specified action trajectories. To achieve precise alignment between conditioned actions and generated visual changes, we introduce a lightweight, universal action-conditioned module that seamlessly integrates into any existing model. Instead of focusing on complex visual details, we demonstrate that consistent dynamic transition modeling is the key to building powerful world simulators. Building upon this insight, we further introduce a motion-reinforced loss that enhances action controllability by compelling the model to capture dynamic changes more effectively. Experiments demonstrate that DWS can be versatilely applied to both diffusion and autoregressive transformer models, achieving significant improvements in generating action-controllable, dynamically consistent videos across games and robotics domains. Moreover, to facilitate the applications of the learned world simulator in downstream tasks such as model-based reinforcement learning, we propose prioritized imagination to improve sample efficiency, demonstrating competitive performance compared with state-of-the-art methods.

视频生成世界模型动作控制强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。