用期望自由能统一探索与决策,无需手动调参即可自动平衡信息获取与奖励。
Expected Free Energy as Belief-Dependent Utility for rho-POMDPs

- 通过最小化期望自由能,将信息增益作为信念依赖的效用直接建模。
- 在65000状态的结构检测任务中,无需调参即超越纯奖励规划方法。
- 适用于医疗筛查、故障检测等需权衡测试成本与漏检风险的场景。
在部分可观测环境下,智能体需决定何时收集信息及哪些观测值得付出代价。标准POMDP仅通过未来奖励间接评估信息价值,而ρ-POMDP框架则直接以信念依赖效用ρ奖励不确定性降低,但ρ的选择和权重常需人工调优。本文证明:主动推断中最小化期望自由能(EFE)等价于求解一个以预期信息增益为效用的ρ-POMDP,且探索权重固定为1,因变分界同时表达功利性与认知价值(单位:纳特)。该等价关系已在观察-承诺型及因子化观测型POMDP上证明,后者涵盖非破坏性检测与移动传感等交错式观测-行动问题。实验验证:从经典老虎问题到RockSample及新提出的结构检测基准(超65,000状态),未调参的权重表现匹配或优于纯奖励规划,在相同视野下避免了任务特调奖励奖金导致的过度探索,并接近成功-奖励帕累托前沿的最优折衷点。实际应用中,如故障检测与医学筛查,每个测试有成本,每例漏诊有代价,EFE提供了源自理论而非人工调校的信念依赖效用。
原文摘要 · Abstract (English)
An agent acting under partial observability must decide when to gather information and which observations are worth their cost. Standard POMDPs value information only through its eventual effect on reward. The $ρ$-POMDP framework instead rewards uncertainty reduction directly, through a belief-dependent utility $ρ$, but in practice both the choice of $ρ$ and the weight placed on it are tuned by hand for every task. We show that active inference removes this tuning entirely. Minimizing Expected Free Energy (EFE) is exactly equivalent to solving a $ρ$-POMDP whose utility is expected information gain, and the exploration weight is fixed at $w=1$ because the variational bound expresses pragmatic and epistemic value in the same units (nats). We prove this equivalence for observe-then-commit POMDPs and extend it to factored observation POMDPs, a broader class that covers interleaved observe-act problems such as non-destructive testing and mobile sensing, where gathering information leaves the hidden state unchanged. Experiments support the theory. Across environments ranging from the classic Tiger problem to RockSample and a new Structural Inspection benchmark with over 65,000 states, the untuned weight matches or outperforms reward-only planning at the same horizon, avoids the over-exploration of bonuses tuned per task, and sits near the reward-maximizing knee of the success-reward Pareto frontier. The practical payoff is an exploration objective that works out of the box. In applications such as fault detection and medical screening, where every test has a price and every missed fault has a cost, EFE supplies a belief-dependent utility that is derived rather than tuned.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。