arXiv:2607.19518cs.AI2026-07

提出用认知先验框架统一解释复杂推理机制,揭示闭环规划的核心作用。

Sophisticated Policies from Epistemic Priors

  • 基于认知先验的变分自由能框架,实现未来动作依赖未来状态的闭环规划
  • 在随机反应迷宫中验证:仅具信息驱动力或仅具闭环结构均无法成功解题
  • 强调闭环推理本质优于树搜索,适合研究主动推理与强化学习融合的学者

复杂推理是主动推理的一种变体,常与递归信念建模和树搜索相关。我们主张其核心计算角色更简单:在规划时域内,通过允许未来动作依赖未来状态与观测,使主动推理形成闭环。该闭环结构可纳入认知先验变分自由能框架。认知先验提供主动推理目标,而未来状态与动作的联合后验则构成状态依赖的控制结构。我们在一个设计用于分离认知激励与内部时域闭环控制的随机基准测试环境——反应迷宫中评估此分解。对比包含相同状态-动作后验族的三种变分目标:动作-状态分解的主动推理、复杂推理及标准期望自由能规划。结果表明,任一成分单独不足:缺乏认知成分的方法不主动寻求信息;禁止未来动作依赖未来状态的方法无法将信息转化为可靠的目标达成。相反,复杂推理与完整联合认知先验主动推理均通过结合认知驱动力与闭环推理成功解决环境。这说明复杂推理的优势并非源自树搜索本身,而是源于主动推理的闭环形式,且该形式可在认知先验变分推断中实现,前提是后验保持未来动作对未来的状态依赖。

原文摘要 · Abstract (English)

Sophisticated Inference is a variant of active inference often associated with recursive belief modeling and tree search. We argue that its central computational role is simpler: within a planning horizon, it makes active inference closed-loop by allowing future actions to depend on future states and observations. This closed-loop structure can be represented in the epistemic-prior variational free energy framework. Epistemic priors supply the active-inference objective, while a joint posterior over future states and actions supplies the state-contingent control structure. We evaluate this decomposition in the Reactivity Maze, a stochastic benchmark designed to separate epistemic incentive from inner-horizon closed-loop control. The comparison includes three variational objectives with the same state-action posterior family, an action-state factorized active inference objective, Sophisticated Inference, and standard Expected Free Energy planning. The results show that neither ingredient is sufficient on its own. Methods without an epistemic component do not seek information, while methods that prevent future actions from depending on future states cannot turn information into reliable goal-reaching. By contrast, both Sophisticated Inference and full-joint epistemic-prior active inference solve the environment by combining epistemic drive with closed-loop inference. These results show that the advantage associated with Sophisticated Inference need not be specific to tree search itself. It arises from the closed-loop form of active inference, and this form can be represented in epistemic-prior variational inference when the posterior keeps future actions dependent on future states.

主动推理闭环规划认知先验强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。