将视频扩散模型转化为可交互的世界模型,提升复杂环境下的决策效率。
Vid2World: Crafting Video Diffusion Models to Interactive World Models
- 通过因果化改造预训练视频扩散模型,实现自回归生成。
- 在机器人操作、3D游戏和开放世界导航中均实现高保真动态预测。
- 支持动作可控性增强,适合需要精准交互的智能体开发场景。
世界模型能从历史观测与动作序列中预测未来状态,在序列决策任务中显著提升数据效率。然而,现有世界模型常需大量领域特定训练,且生成结果保真度低、细节粗糙,难以应用于复杂环境。相比之下,基于大规模互联网数据训练的视频扩散模型展现出生成高质量视频、捕捉多样化真实世界动态的强大能力。本文提出Vid2World,一种通用方法,可将预训练视频扩散模型迁移为可交互的世界模型。该方法系统探索视频扩散模型的因果化改造,重构模型架构与训练目标以支持自回归生成,并引入因果动作引导机制,增强生成过程中的动作可控性。在机器人操作、3D游戏模拟及开放世界导航等多个领域开展的实验表明,该方法为将高性能视频扩散模型转化为可交互世界模型提供了可扩展且高效的技术路径。
原文摘要 · Abstract (English)
World models, which predict future transitions from past observation and action sequences, have shown great promise for improving data efficiency in sequential decision-making. However, existing world models often require extensive domain-specific training and still produce low-fidelity, coarse predictions, limiting their usefulness in complex environments. In contrast, video diffusion models trained on large-scale internet data have demonstrated impressive capabilities in generating high-quality videos that capture diverse real-world dynamics. In this work, we present Vid2World, a general approach for leveraging and transferring pre-trained video diffusion models into interactive world models. To bridge the gap, Vid2World systematically explores video diffusion causalization, reshaping both the architecture and training objective of pre-trained models to enable autoregressive generation. Additionally, it incorporates a causal action guidance mechanism to enhance action controllability in the resulting interactive world models. Extensive experiments across multiple domains, including robot manipulation, 3D game simulation, and open-world navigation, demonstrate that our method offers a scalable and effective pathway for repurposing highly capable video diffusion models into interactive world models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。