用贝斯自由能重构主动推断,支持消息传递并控制探索强度。
Expected free energy as an information constraint on the Bethe Lagrangian

- 基于贝斯自由能构建新框架,支持消息传递推断。
- 信息约束调节探索程度,可实现从不探索到最大化探索的连续变化。
- 在三个任务中表现优于传统方法,适合需要可控探索的强化学习场景。
主动推断通过最小化对预测未来的期望自由能来选择行动。然而,对未观测结果的期望使自由能函数失去KL散度结构,阻碍了消息传递的推断处理。本文提出一种基于贝斯自由能函数的替代形式,完全支持消息传递推断。通过引入信息约束(除归一化、边缘化和形式约束外),要求给定动作下未来观测、状态与参数间的互信息至少等于目标先验熵。当对应KKT乘子取特定值时,该约束贝斯拉格朗日的驻点恢复预期自由能解。我们展示随着信息需求变化,求解乘子会经历无活性、内部与饱和三种状态:无活性时代理完全停止探索,饱和时探索达到最大。在三个任务上对比了该约束贝斯代理与EFE及Q-MDP的表现。
原文摘要 · Abstract (English)
Active inference selects actions by minimising an expected free energy functional over predicted futures. However, adding an expectation over yet-unobserved outcomes means the free energy functional no longer has a Kullback-Leibler structure, which hinders message passing treatments of inference procedures. We propose an alternative formulation based on a Bethe free energy functional, fully supporting inference by message passing. The epistemic drive is maintained by imposing an information constraint, next to normalisation, marginalisation and form constraints, insisting that the mutual information between future observations, states and parameters given actions must be at least as large as the entropy of the goal prior. For a specific value of the corresponding Karush-Kuhn-Tucker multiplier, the stationary point of this constrained Bethe Lagrangian recovers the expected free energy solution. We show that, as the information demand is varied, the solved multiplier moves through its inactive, interior, and saturated regimes. In the inactive regime the agent's epistemic drive switches off entirely, while in the saturated regime it is maximal. We compare the performance of the constrained Bethe agent on three tasks against EFE and Q-MDP.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。