将规划问题统一为变分推断,让智能体自动平衡目标达成与信息获取。
Expected Free Energy-based Planning as Variational Inference
- 用变分自由能框架重写期望自由能,使规划可理论分析
- 在三个复杂环境中验证了信息探索行为和性能提升
- 适合研究主动推理、强化学习理论的学者参考
不确定性下的规划要求智能体在达成目标与收集信息之间取得平衡。主动推理通过期望自由能(EFE)这一代价函数,统一了工具性与认知性目标。然而,现有基于EFE的方法通常依赖专用优化流程,难以扩展或分析。本文表明,基于EFE的规划可被形式化为带有认知先验的生成模型上的变分自由能最小化。核心结果证明:在合理选择先验条件下,变分自由能泛函可分解为期望计划代价(即EFE)加上一个复杂度项。该公式强化了与自由能原理的理论一致性,将规划视为与感知和学习相同的推断过程。我们在三个复杂度递增的环境中验证了该方法:确定性的T形迷宫、随机的Reactivity Maze,以及部分可观测的MiniGrid DoorKey-8x8环境。实验表明,认知先验能诱导出信息探索行为;变分形式下基于策略的推断优于基于计划的方法,在随机转移下表现更优;时间因子分解使算法可扩展至现有表格型主动推理无法处理的环境。
原文摘要 · Abstract (English)
Planning under uncertainty requires agents to balance goal achievement with information gathering. Active inference addresses this through the Expected Free Energy (EFE), a cost function that unifies instrumental and epistemic objectives. However, existing EFE-based methods typically employ specialized optimization procedures that are difficult to extend or analyze. In this paper, we show that EFE-based planning can be formulated as Variational Free Energy minimization on a generative model augmented with epistemic priors. Our main result demonstrates that minimizing a Variational Free Energy functional with appropriately chosen priors yields a decomposition into expected plan costs (the EFE) plus a complexity term. This formulation reinforces theoretical consistency with the Free Energy Principle by casting planning as the same inferential process that governs perception and learning. We validate our approach on three environments of increasing complexity: a deterministic T-maze, a stochastic Reactivity Maze, and a partially observable MiniGrid DoorKey-8x8 environment. The experiments demonstrate that the epistemic priors induce information-seeking behavior, that the variational formulation yields policy-based inference outperforming plan-based methods under stochastic transitions, and that temporal factorization enables scalability to environments where existing tabular active inference methods cannot operate.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。