将主动推理转化为凸马尔可夫决策过程,统一了探索与目标导向行为。
Active Inference as a Convex Markov Decision Process

- 用凸优化视角重构主动推理,将策略学习建模为期望自由能最小化。
- 推导出镜面下降算法,实现政策依赖奖励的在线优化。
- 为现代强化学习提供理论基础,适合研究认知模型与强化学习融合者。
主动推理(AIF)将适应性行为视为期望自由能(EFE)最小化,通过单一变分原理统一了认知探索与实用目标。本文将AIF建模为策略优化问题,证明在闭环控制策略下,EFE最小化可表述为凸马尔可夫决策过程(MDP)。其中,实用项关于预测状态边际呈线性,等价于潜在MDP中的奖励最大化;而认知价值引入非线性成分,使EFE最小化区别于标准强化学习。该视角揭示主动推理的认知驱动力是一种政策依赖的(表演性)奖励。我们分析了有限时域、折扣与平均奖励形式的EFE,推导出一种镜面下降(MD)算法,其在当前状态边际附近对目标函数进行局部线性化,生成兼容演员-评论家方法与动态规划的政策依赖奖励。最后,我们论证将世界模型学习与策略优化耦合,使主动推理具备表演性强化学习结构,为将其融入现代强化学习与优化理论(包括收敛性分析与原则性策略改进保证)提供了路径。
原文摘要 · Abstract (English)
Active Inference (AIF) frames adaptive behavior as the minimization of expected free energy (EFE), combining epistemic and pragmatic objectives within a single variational principle. We frame AIF as policy optimization and show that, for closed-loop control policies, EFE minimization can be formulated as a convex Markov decision process (MDP). In this formulation, the pragmatic terms are linear in the predictive state marginals and therefore equivalent to reward maximization in a latent MDP, while the epistemic value introduces a nonlinear component that distinguishes EFE minimization from standard reinforcement learning. This perspective further reveals the epistemic drive of active inference as a policy-dependent (performative) reward. We analyze finite-horizon, discounted, and average-reward formulations of EFE and derive a mirror descent (MD) algorithm that locally linearizes the objective around the current state marginals, yielding a policy-dependent reward that is compatible with actor-critic methods and dynamic programming. Finally, we argue that coupling world-model learning with policy optimization gives active inference the structure of performative reinforcement learning, providing a route toward grounding active inference within modern reinforcement learning and optimization theory, including convergence analysis and principled policy improvement guarantees.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。