arXiv:2505.05495cs.CVcs.RO2025-05NeurIPS被引 33

构建可记忆环境的3D世界模型,支持长期规划与一致模拟。

Learning 3D Persistent Embodied World Models

  • 用视频扩散模型生成未来观测,融合为持久3D地图。
  • 基于3D地图条件化生成,能同步模拟可见与不可见区域。
  • 适用于复杂场景下的智能体规划与策略学习。

智能体模拟未来行为对世界影响的能力至关重要,有助于其预判动作后果并制定计划。现有基于视频模型的世界建模方法多为短视,缺乏对未观测场景的记忆,难以在复杂环境中实现长期一致的规划。本文提出一种具有显式记忆的持续性具身世界模型,通过视频扩散模型生成未来RGB-D视频,并聚合为持久的3D环境地图。该地图作为条件输入,使视频模型能够忠实模拟已知与未知区域。实验表明,该模型在下游具身任务中显著提升规划与策略学习效果。

原文摘要 · Abstract (English)

The ability to simulate the effects of future actions on the world is a crucial ability of intelligent embodied agents, enabling agents to anticipate the effects of their actions and make plans accordingly. While a large body of existing work has explored how to construct such world models using video models, they are often myopic in nature, without any memory of a scene not captured by currently observed images, preventing agents from making consistent long-horizon plans in complex environments where many parts of the scene are partially observed. We introduce a new persistent embodied world model with an explicit memory of previously generated content, enabling much more consistent long-horizon simulation. During generation time, our video diffusion model predicts RGB-D video of the future observations of the agent. This generation is then aggregated into a persistent 3D map of the environment. By conditioning the video model on this 3D spatial map, we illustrate how this enables video world models to faithfully simulate both seen and unseen parts of the world. Finally, we illustrate the efficacy of such a world model in downstream embodied applications, enabling effective planning and policy learning.

3D世界模型具身智能视频生成长期规划

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。