用混合主动推理框架,让智能体在连续任务中学会抽象离散策略。
Learning in Hybrid Active Inference Models
- 高层离散规划器与底层连续控制器分层协同,实现灵活决策。
- 在稀疏奖励的连续山地小车任务中,探索效率提升且成功规划出抽象子目标。
- 适合研究强化学习中抽象层次与高效探索的学者参考。
人工智能中的一个开放问题是如何使系统灵活地学习对解决本质连续问题有用的离散抽象。以往计算神经科学的工作在主动推理框架下探讨了决策过程中离散与连续变量的功能整合(Parr, Friston & de Vries, 2017;Parr & Friston, 2018),但其重点在于类别决策的物理实现,且假设层级混合生成模型已知。因此,该框架如何扩展至学习仍不明确。本文提出一种新型分层混合主动推理智能体:高层离散主动推理规划器位于低层连续主动推理控制器之上。我们利用最近关于递归切换线性动态系统(rSLDS)的研究,通过复杂连续动力学的分段线性分解实现端到端的有意义离散表征学习(Linderman et al., 2016)。rSLDS学习到的表征指导混合决策智能体的结构,使我们能够(1)以类似选项框架的方式指定时序抽象子目标,(2)将探索拓展至离散空间,从而利用信息论探索奖励,(3)在离散规划器中‘缓存’低层问题的近似解。我们将模型应用于稀疏奖励的连续山地小车任务,在增强探索的基础上实现了快速系统识别,并成功完成规划。
原文摘要 · Abstract (English)
An open problem in artificial intelligence is how systems can flexibly learn discrete abstractions that are useful for solving inherently continuous problems. Previous work in computational neuroscience has considered this functional integration of discrete and continuous variables during decision-making under the formalism of active inference (Parr, Friston & de Vries, 2017; Parr & Friston, 2018). However, their focus is on the expressive physical implementation of categorical decisions and the hierarchical mixed generative model is assumed to be known. As a consequence, it is unclear how this framework might be extended to learning. We therefore present a novel hierarchical hybrid active inference agent in which a high-level discrete active inference planner sits above a low-level continuous active inference controller. We make use of recent work in recurrent switching linear dynamical systems (rSLDS) which implement end-to-end learning of meaningful discrete representations via the piecewise linear decomposition of complex continuous dynamics (Linderman et al., 2016). The representations learned by the rSLDS inform the structure of the hybrid decision-making agent and allow us to (1) specify temporally-abstracted sub-goals in a method reminiscent of the options framework, (2) lift the exploration into discrete space allowing us to exploit information-theoretic exploration bonuses and (3) `cache' the approximate solutions to low-level problems in the discrete planner. We apply our model to the sparse Continuous Mountain Car task, demonstrating fast system identification via enhanced exploration and successful planning through the delineation of abstract sub-goals.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。