用因果结构提升强化学习的探索效率和控制能力
Towards Empowerment Gain through Causal Structure Learning in Model-Based RL
- 基于因果动态模型实现高效探索与任务学习
- 在6个环境中提升样本效率与最终性能
- 适用于需高效决策的复杂环境建模任务
在基于模型的强化学习(MBRL)中,将因果结构融入动态模型可使智能体获得对环境的结构化理解,从而实现高效决策。赋能作为内在动机,通过最大化未来状态与动作之间的互信息,增强智能体主动控制环境的能力。本文提出一种新框架ECL(Empowerment through Causal Learning),让具备因果动态模型意识的智能体实现赋能驱动的探索,并优化其因果结构以促进任务学习。具体而言,ECL首先基于收集数据训练环境的因果动态模型;随后在因果结构下最大化赋能以进行探索,同时利用探索所得数据更新因果模型,使其比无因果结构的密集动态模型更具可控性。在下游任务学习中引入内在好奇心奖励,平衡因果性并缓解过拟合。ECL方法无关,可兼容多种因果发现方法。在6个环境(包括像素级任务)上评估,结合3种因果发现方法,ECL在因果发现能力、样本效率和最终性能上均优于其他因果MBRL方法。
原文摘要 · Abstract (English)
In Model-Based Reinforcement Learning (MBRL), incorporating causal structures into dynamics models provides agents with a structured understanding of the environments, enabling efficient decision. Empowerment as an intrinsic motivation enhances the ability of agents to actively control their environments by maximizing the mutual information between future states and actions. We posit that empowerment coupled with causal understanding can improve controllability, while enhanced empowerment gain can further facilitate causal reasoning in MBRL. To improve learning efficiency and controllability, we propose a novel framework, Empowerment through Causal Learning (ECL), where an agent with the awareness of causal dynamics models achieves empowerment-driven exploration and optimizes its causal structure for task learning. Specifically, ECL operates by first training a causal dynamics model of the environment based on collected data. We then maximize empowerment under the causal structure for exploration, simultaneously using data gathered through exploration to update causal dynamics model to be more controllable than dense dynamics model without causal structure. In downstream task learning, an intrinsic curiosity reward is included to balance the causality, mitigating overfitting. Importantly, ECL is method-agnostic and is capable of integrating various causal discovery methods. We evaluate ECL combined with 3 causal discovery methods across 6 environments including pixel-based tasks, demonstrating its superior performance compared to other causal MBRL methods, in terms of causal discovery, sample efficiency, and asymptotic performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。