arXiv:2512.03293cs.AIq-bio.NC2025-12被引 1

通过四类目标设定方式,对比智能体在导航任务中的表现差异。

Prior preferences in active inference agents: soft, hard, and goal shaping

  • 设计四种偏好分布:硬/软目标结合或不结合中间目标引导
  • 引入中间目标的智能体表现最佳,但学习环境动态能力下降
  • 揭示目标设定策略对探索与利用的权衡机制,适合强化学习研究者

主动推理将预期自由能作为规划与决策的目标,以在学习智能体中合理平衡利用与探索。利用驱动(即智能体想要达成的目标)被形式化为变分概率分布与偏好概率分布之间的KL散度,后者表示哪些状态或观测更可能,从而决定智能体在特定环境中的目标。现有文献极少关注偏好分布应如何设定及其对推理与学习的影响。本文研究了四种定义偏好分布的方式:提供硬目标或软目标,以及是否包含目标塑造(即中间目标)。我们在网格世界导航任务中比较了四类智能体的表现。结果表明,目标塑造整体性能最优(促进利用),但牺牲了对环境转移动态的学习能力(抑制探索)。

原文摘要 · Abstract (English)

Active inference proposes expected free energy as an objective for planning and decision-making to adequately balance exploitative and explorative drives in learning agents. The exploitative drive, or what an agent wants to achieve, is formalised as the Kullback-Leibler divergence between a variational probability distribution, updated at each inference step, and a preference probability distribution that indicates what states or observations are more likely for the agent, hence determining the agent's goal in a certain environment. In the literature, the questions of how the preference distribution should be specified and of how a certain specification impacts inference and learning in an active inference agent have been given hardly any attention. In this work, we consider four possible ways of defining the preference distribution, either providing the agents with hard or soft goals and either involving or not goal shaping (i.e., intermediate goals). We compare the performances of four agents, each given one of the possible preference distributions, in a grid world navigation task. Our results show that goal shaping enables the best performance overall (i.e., it promotes exploitation) while sacrificing learning about the environment's transition dynamics (i.e., it hampers exploration).

主动推理目标设定强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。