arXiv:2503.18938cs.AIcs.CV2025-03ICML被引 108

让世界模型学会从视频中自监督提取隐式动作,实现高效适应新环境。

AdaWorld: Learning Adaptable World Models with Latent Actions

论文配图:AdaWorld: Learning Adaptable World Models with Latent Actions
图 1 · 摘自论文原文
  • 通过自监督从视频中提取关键帧间隐式动作
  • 在有限交互下仍能高效学习新动作,模拟质量更优
  • 适合需要快速适应新任务的智能体系统

世界模型旨在学习受动作控制的未来预测,对智能体发展至关重要。但现有模型严重依赖大量带动作标注的数据和高昂训练成本,难以在异构动作环境下通过少量交互快速适应,限制了其跨领域应用。为此,我们提出AdaWorld,一种创新的世界模型学习方法,实现高效适应。核心思想是在预训练阶段融入动作信息:通过自监督方式从视频中提取隐式动作,捕捉帧间关键变化;随后构建基于这些隐式动作的自回归世界模型。该学习范式使模型具备高度可适应性,即使在有限交互和微调条件下也能高效迁移并学习新动作。在多个环境中的综合实验表明,AdaWorld在仿真质量和视觉规划方面均表现更优。

原文摘要 · Abstract (English)

World models aim to learn action-controlled future prediction and have proven essential for the development of intelligent agents. However, most existing world models rely heavily on substantial action-labeled data and costly training, making it challenging to adapt to novel environments with heterogeneous actions through limited interactions. This limitation can hinder their applicability across broader domains. To overcome this limitation, we propose AdaWorld, an innovative world model learning approach that enables efficient adaptation. The key idea is to incorporate action information during the pretraining of world models. This is achieved by extracting latent actions from videos in a self-supervised manner, capturing the most critical transitions between frames. We then develop an autoregressive world model that conditions on these latent actions. This learning paradigm enables highly adaptable world models, facilitating efficient transfer and learning of new actions even with limited interactions and finetuning. Our comprehensive experiments across multiple environments demonstrate that AdaWorld achieves superior performance in both simulation quality and visual planning.

世界模型自监督动作学习泛化能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。