用物体中心表征构建动态世界模型,提升复杂场景下的预测与规划能力
Dyn-O: Building Structured World Models with Object-Centric Representations
- 基于物体中心表征,分离静态与动态特征以增强建模能力
- 在Procgen游戏上比DreamerV3更准地预测未来状态
- 适合研究具身智能、强化学习与复杂环境建模的学者
世界模型旨在捕捉环境动态,使智能体能够预测和规划未来状态。在多数实际场景中,环境动态高度依赖物体间的交互,这推动了基于物体中心而非整体表征的世界模型发展,以更有效地捕捉环境动态并提升组合泛化能力。然而,现有物体中心世界模型主要局限于视觉结构简单的环境(如基本几何形状)。目前尚不清楚此类模型能否推广到包含多样纹理和杂乱场景的复杂设置。本文填补这一空白,提出Dyn-O——一种基于物体中心表征的增强型结构化世界模型。相比先前工作,Dyn-O在表征学习和动态建模两方面均有提升。在挑战性极高的Procgen游戏中,我们的方法可直接从像素观测中学习物体中心世界模型,在滚动预测精度上优于DreamerV3。此外,通过将物体中心特征解耦为与动态无关和与动态相关两部分,实现了对特征的细粒度操控,生成更多样化的想象轨迹。
原文摘要 · Abstract (English)
World models aim to capture the dynamics of the environment, enabling agents to predict and plan for future states. In most scenarios of interest, the dynamics are highly centered on interactions among objects within the environment. This motivates the development of world models that operate on object-centric rather than monolithic representations, with the goal of more effectively capturing environment dynamics and enhancing compositional generalization. However, the development of object-centric world models has largely been explored in environments with limited visual complexity (such as basic geometries). It remains underexplored whether such models can generalize to more complex settings with diverse textures and cluttered scenes. In this paper, we fill this gap by introducing Dyn-O, an enhanced structured world model built upon object-centric representations. Compared to prior work in object-centric representations, Dyn-O improves in both learning representations and modeling dynamics. On the challenging Procgen games, we find that our method can learn object-centric world models directly from pixel observations, outperforming DreamerV3 in rollout prediction accuracy. Furthermore, by decoupling object-centric features into dynamics-agnostic and dynamics-aware components, we enable finer-grained manipulation of these features and generate more diverse imagined trajectories.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。