用少量标注训练物体中心世界模型,提升强化学习采样效率
Object-Centric World Models from Few-Shot Annotations for Sample-Efficient Reinforcement Learning
- 基于预训练分割网络提取物体级表征,引导模型关注关键物体
- 在Atari 100k上超越基线,在Hollow Knight boss战中达到顶尖采样效率
- 仅需少量标注帧即可建模物体动态与交互,适合复杂视觉任务
尽管基于像素的深度强化学习取得了显著进展,但其样本效率低仍是真实应用的关键瓶颈。模型基础强化学习(MBRL)通过学习世界模型生成模拟经验来缓解该问题,但依赖像素级重建损失的标准方法在复杂动态场景中难以捕捉小而关键的物体。我们提出,物体中心(OC)表征能将模型容量聚焦于语义上有意义的实体,从而提升动态预测和样本效率。本文提出OC-STORM框架,利用预训练分割网络提取的物体表征增强学习的世界模型。通过极少数量的标注帧,OC-STORM能够追踪决策相关物体动态及物体间交互,无需大量标注或特权信息。实验表明,OC-STORM在Atari 100k基准上显著优于STORM基线,并在视觉复杂的《空洞骑士》挑战关卡中实现最先进的样本效率。结果证明,将物体中心先验融入MBRL对复杂视觉领域具有巨大潜力。
原文摘要 · Abstract (English)
While deep reinforcement learning (RL) from pixels has achieved remarkable success, its sample inefficiency remains a critical limitation for real-world applications. Model-based RL (MBRL) addresses this by learning a world model to generate simulated experience, but standard approaches that rely on pixel-level reconstruction losses often fail to capture small, task-critical objects in complex, dynamic scenes. We posit that an object-centric (OC) representation can direct model capacity toward semantically meaningful entities, improving dynamics prediction and sample efficiency. In this work, we introduce OC-STORM, an object-centric MBRL framework that enhances a learned world model with object representations extracted by a pretrained segmentation network. By conditioning on a minimal number of annotated frames, OC-STORM learns to track decision-relevant object dynamics and inter-object interactions without extensive labeling or access to privileged information. Empirical results demonstrate that OC-STORM significantly outperforms the STORM baseline on the Atari 100k benchmark and achieves state-of-the-art sample efficiency on challenging boss fights in the visually complex game Hollow Knight. Our findings underscore the potential of integrating OC priors into MBRL for complex visual domains. Project page: https://oc-storm.weipuzhang.com
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。