用物体中心模型提升强化学习,发现表示漂移会破坏控制性能
When Object-Centric World Models Meet Policy Learning: From Pixels to Policies, and Where It Breaks
- 从像素直接学习物体级潜在表示,实现无监督解耦
- 模型对分布外视觉变化鲁棒,但控制任务表现差于DreamerV3
- 揭示多物体交互时潜在空间漂移是政策不稳定的主因
物体中心世界模型(OCWM)旨在将视觉场景分解为物体级别的表征,提供结构化抽象,以提升强化学习中的组合泛化能力和数据效率。我们假设,显式解耦的物体级表征通过定位任务相关信息,可提升在新特征组合下的策略性能。为此,我们提出DLPWM,一种完全无监督、解耦的物体中心世界模型,直接从像素中学习物体级潜在变量。DLPWM在重建和预测任务上表现优异,对多种分布外(OOD)视觉变化具有鲁棒性。然而,在下游基于模型的控制任务中,基于DLPWM潜在变量训练的策略表现不如DreamerV3。通过潜在轨迹分析,我们识别出多物体交互过程中的表示漂移是导致策略学习不稳定的主因。结果表明,尽管物体中心感知支持稳健的视觉建模,但实现稳定控制仍需缓解潜在漂移。
原文摘要 · Abstract (English)
Object-centric world models (OCWM) aim to decompose visual scenes into object-level representations, providing structured abstractions that could improve compositional generalization and data efficiency in reinforcement learning. We hypothesize that explicitly disentangled object-level representations, by localizing task-relevant information, can enhance policy performance across novel feature combinations. To test this hypothesis, we introduce DLPWM, a fully unsupervised, disentangled object-centric world model that learns object-level latents directly from pixels. DLPWM achieves strong reconstruction and prediction performance, including robustness to several out-of-distribution (OOD) visual variations. However, when used for downstream model-based control, policies trained on DLPWM latents underperform compared to DreamerV3. Through latent-trajectory analyses, we identify representation shift during multi-object interactions as a key driver of unstable policy learning. Our results suggest that, although object-centric perception supports robust visual modeling, achieving stable control requires mitigating latent drift.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。