arXiv:2505.24784cs.AIcs.LG2025-05被引 5

用物体中心模型让智能体几分钟学会玩游戏,数据效率远超传统方法

AXIOM: Learning to Play Games in Minutes with Expanding Object-Centric Models

  • 将场景建模为物体组合,用分段线性轨迹捕捉物体间稀疏交互
  • 仅用1万步交互就掌握多种游戏,参数量少且无需梯度优化
  • 在线扩展模型结构并定期精简,兼顾高效学习与跨任务泛化

当前深度强化学习方法在多个领域表现卓越,但相比人类学习的数据效率仍显不足,后者依赖关于物体及其相互作用的核心先验。主动推理提供了一个将感官信息与先验知识结合的原理性框架,可学习世界模型并量化信念与预测的不确定性。然而,现有主动推理模型通常针对单一任务设计,缺乏深度强化学习的跨任务灵活性。为此,我们提出一种新架构——AXIOM,通过融入最小但表达力强的一组物体中心动态与交互先验,在低数据环境下加速学习。该方法结合了贝叶斯方法的数据效率与可解释性,以及深度强化学习的跨任务泛化能力。AXIOM将场景表示为物体的组合,其动态建模为分段线性轨迹,以捕捉稀疏的物体-物体交互。生成模型的结构通过在线增长和从单个事件中学习混合模型来扩展,并定期通过贝叶斯模型简化进行优化,促进泛化。AXIOM仅需10,000次交互步骤即可掌握多种游戏,参数量显著低于典型DRL模型,且避免了基于梯度优化的计算开销。

原文摘要 · Abstract (English)

Current deep reinforcement learning (DRL) approaches achieve state-of-the-art performance in various domains, but struggle with data efficiency compared to human learning, which leverages core priors about objects and their interactions. Active inference offers a principled framework for integrating sensory information with prior knowledge to learn a world model and quantify the uncertainty of its own beliefs and predictions. However, active inference models are usually crafted for a single task with bespoke knowledge, so they lack the domain flexibility typical of DRL approaches. To bridge this gap, we propose a novel architecture that integrates a minimal yet expressive set of core priors about object-centric dynamics and interactions to accelerate learning in low-data regimes. The resulting approach, which we call AXIOM, combines the usual data efficiency and interpretability of Bayesian approaches with the across-task generalization usually associated with DRL. AXIOM represents scenes as compositions of objects, whose dynamics are modeled as piecewise linear trajectories that capture sparse object-object interactions. The structure of the generative model is expanded online by growing and learning mixture models from single events and periodically refined through Bayesian model reduction to induce generalization. AXIOM masters various games within only 10,000 interaction steps, with both a small number of parameters compared to DRL, and without the computational expense of gradient-based optimization.

强化学习物体中心主动推理高效学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。