让视频生成模型和控制模块一起进化,提升模拟与规划能力。
Co-Evolving Latent Action World Models
- 用预训练世界模型替代原动态模型,联合训练控制模块
- 通过预热阶段对齐表征,避免表示崩溃,实现双向优化
- 适合做通用世界模型的科研人员,尤其关注可控生成与规划
将预训练视频生成模型转化为可控制的世界模型,是构建通用世界模型的重要方向。现有主流方法采用两阶段训练:先独立训练隐空间动作模型(LAM)和世界模型,导致冗余训练且难以协同优化。一个直观但困难的想法是直接用强大世界模型替换LAM中的前向动态模型,并联合训练,但易引发表征崩溃。本文提出CoLA-World,首次成功实现这一协同范式,通过关键的预热阶段,有效对齐从零开始训练的LAM与预训练世界模型的表征。这开启了一个共演化循环:世界模型作为知识导师,提供梯度指导高质量LAM的构建;而LAM则为世界模型提供更精确、灵活的控制接口。实验证明,CoLA-World在视频模拟质量与下游视觉规划任务上达到或超越已有两阶段方法,确立了一种高效稳健的新范式。
原文摘要 · Abstract (English)
Adapting pretrained video generation models into controllable world models via latent actions is a promising step towards creating generalist world models. The dominant paradigm adopts a two-stage approach that trains latent action model (LAM) and the world model separately, resulting in redundant training and limiting their potential for co-adaptation. A conceptually simple and appealing idea is to directly replace the forward dynamic model in LAM with a powerful world model and training them jointly, but it is non-trivial and prone to representational collapse. In this work, we propose CoLA-World, which for the first time successfully realizes this synergistic paradigm, resolving the core challenge in joint learning through a critical warm-up phase that effectively aligns the representations of the from-scratch LAM with the pretrained world model. This unlocks a co-evolution cycle: the world model acts as a knowledgeable tutor, providing gradients to shape a high-quality LAM, while the LAM offers a more precise and adaptable control interface to the world model. Empirically, CoLA-World matches or outperforms prior two-stage methods in both video simulation quality and downstream visual planning, establishing a robust and efficient new paradigm for the field.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。