arXiv:2606.14418cs.AIcs.LG2026-06

用物体中心的因果模型提升强化学习规划效率

Causal Object-Centric Models for Planning with Monte Carlo Tree Search

论文配图:Causal Object-Centric Models for Planning with Monte Carlo Tree Search
图 1 · 摘自论文原文
  • 将动作绑定到物体槽,通过融合机制预测状态转移
  • 在多个复杂任务中早期训练阶段得分优于基线模型
  • 适合需要高效决策的视觉强化学习场景

我们提出COMET(因果物体中心模型用于高效树搜索),一种基于模型的强化学习算法,在槽结构的隐空间中执行蒙特卡洛树搜索。COMET将冻结的无监督物体中心编码器与基于Transformer的世界模型结合,通过新颖的动作-槽融合机制将动作绑定至物体,用于槽状态转移预测。策略和价值头采用物体因果注意力,通过学习的每槽相关性分数调节令牌间交互,使决策聚焦于任务相关实体。COMET为MuZero类隐式规划引入显式的物体级归纳偏置。在来自物体中心视觉强化学习基准的八个视觉与动态多样性任务(ManiSkill、Robosuite、VizDoom)中,COMET在训练早期阶段实现了比物体中心及整体基线更高的平均归一化得分。

原文摘要 · Abstract (English)

We introduce COMET (Causal Object-centric Model for Efficient Tree search), a model-based reinforcement learning algorithm that performs Monte Carlo Tree Search in a slot-structured latent space. COMET pairs a frozen unsupervised object-centric encoder with a transformer-based world model, in which actions are bound to objects through a novel action-slot fusion mechanism that is used in slot transition prediction. Policy and value heads use object-causal attention, modulating token interactions by learned per-slot relevance scores so that decision-making concentrates on task-relevant entities. COMET adds an explicit object-level inductive bias to MuZero-style latent planning. Across eight visually and dynamically diverse tasks from the Object-Centric Visual RL benchmark, ManiSkill, Robosuite, and VizDoom, COMET achieves a higher mean normalized score during the early stages of training compared to object-centric and monolithic baselines.

强化学习因果建模物体中心规划

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。