arXiv:2411.14991cs.AIcs.LG2024-11被引 1

用可解释模型实现主动推理,让智能体基于预测精度自主决策。

Free Energy Projective Simulation (FEPS): Active inference with interpretability

  • 基于自由能原理构建无深度网络的可解释推理模型
  • 在两个生物行为任务中成功消除环境模糊性
  • 适合需要透明决策过程的科研与安全应用

过去十年,自由能原理(FEP)与主动推理(AIF)在连接学习与认知概念模型和感知行动数学模型方面取得诸多进展。本研究提出一种无需深度神经网络的可解释智能体建模方法——自由能投影模拟(FEPS),在FEP与AIF约束下,仅通过内部奖励构建部分可观测环境的表征。基于该世界模型,通过最小化期望自由能推导策略以完成任务。利用模型可解释性,引入技术处理长期目标并减少因隐状态误估计导致的预测误差。在两个受行为生物学启发的强化学习环境(时序反应任务与部分可观测网格导航任务)中测试显示,FEPS智能体仅依赖预测准确性即可完全消解环境歧义,并灵活推断任意目标观测下的最优策略。

原文摘要 · Abstract (English)

In the last decade, the free energy principle (FEP) and active inference (AIF) have achieved many successes connecting conceptual models of learning and cognition to mathematical models of perception and action. This effort is driven by a multidisciplinary interest in understanding aspects of self-organizing complex adaptive systems, including elements of agency. Various reinforcement learning (RL) models performing active inference have been proposed and trained on standard RL tasks using deep neural networks. Recent work has focused on improving such agents' performance in complex environments by incorporating the latest machine learning techniques. In this paper, we take an alternative approach. Within the constraints imposed by the FEP and AIF, we attempt to model agents in an interpretable way without deep neural networks by introducing Free Energy Projective Simulation (FEPS). Using internal rewards only, FEPS agents build a representation of their partially observable environments with which they interact. Following AIF, the policy to achieve a given task is derived from this world model by minimizing the expected free energy. Leveraging the interpretability of the model, techniques are introduced to deal with long-term goals and reduce prediction errors caused by erroneous hidden state estimation. We test the FEPS model on two RL environments inspired from behavioral biology: a timed response task and a navigation task in a partially observable grid. Our results show that FEPS agents fully resolve the ambiguity of both environments by appropriately contextualizing their observations based on prediction accuracy only. In addition, they infer optimal policies flexibly for any target observation in the environment.

主动推理可解释性强化学习自由能原理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。