让智能体显式学习物体间交互,提升强化学习的效率与泛化能力。
Learning Interactive World Model for Object-Centric Reinforcement Learning
- 从像素直接学习物体与交互结构,用模块化表示环境动态。
- 在模拟机器人任务中,样本效率和泛化能力优于基线方法。
- 适合需要鲁棒控制的机器人与具身智能任务研究者。
能够理解物体及其相互作用的智能体可学习更鲁棒、可迁移的策略。然而,大多数物体中心强化学习方法仅对个体物体进行状态分解,而将交互关系隐式处理。本文提出因子化交互物体中心世界模型(FIOC-WM),一种统一框架,能在世界模型中同时学习物体及其交互的结构化表征。FIOC-WM通过解耦且模块化的物体交互表示捕捉环境动态,提升策略学习的样本效率与泛化能力。具体而言,FIOC-WM首先利用预训练视觉编码器从像素中直接学习物体中心潜在表示和交互结构;随后,该世界模型将任务分解为可组合的交互原语,再训练分层策略:高层选择交互类型与顺序,底层执行具体操作。在模拟机器人与具身智能基准测试中,FIOC-WM在政策学习的样本效率与泛化能力上均优于世界模型基线,表明显式、模块化的交互学习对鲁棒控制至关重要。
原文摘要 · Abstract (English)
Agents that understand objects and their interactions can learn policies that are more robust and transferable. However, most object-centric RL methods factor state by individual objects while leaving interactions implicit. We introduce the Factored Interactive Object-Centric World Model (FIOC-WM), a unified framework that learns structured representations of both objects and their interactions within a world model. FIOC-WM captures environment dynamics with disentangled and modular representations of object interactions, improving sample efficiency and generalization for policy learning. Concretely, FIOC-WM first learns object-centric latents and an interaction structure directly from pixels, leveraging pre-trained vision encoders. The learned world model then decomposes tasks into composable interaction primitives, and a hierarchical policy is trained on top: a high level selects the type and order of interactions, while a low level executes them. On simulated robotic and embodied-AI benchmarks, FIOC-WM improves policy-learning sample efficiency and generalization over world-model baselines, indicating that explicit, modular interaction learning is crucial for robust control.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。