提出无需预设探索方向的内在动机机制,自动平衡未知与噪声。
Principled Direction-Free Intrinsic Motivation through Model-Free Epistemic Free-Energy Estimators
- 基于无偏好自由能估计,分离可解释的惊喜与不可约噪声。
- 在状态空间中动态调节探索强度,解决区域已明确时停止探索。
- 适用于高熵与低熵环境,适合不依赖显式预测模型的强化学习场景。
在包含混合不确定性来源的环境中,无监督强化学习需要一种不预设特定惊喜方向的内在动机。传统惊喜最小化仅适用于不稳定环境,而预测误差好奇心会奖励总期望惊喜(含不可约噪声)。若在惊喜最小化与最大化之间切换,会人为引入非平稳性。本文提出一种单一内在奖励,在每个时间窗口内保持平稳,源自无偏好期望自由能目标中的新颖性贡献,以奖励最大化形式表达。核心思想是:参数信息增益(下一状态的预期惊喜减去不可约部分)是高低熵状态空间中合适的内在信号。最大化该值仅追求模型可解释的惊喜。在动态未解区域,该认知项驱动探索;当动态被解析后,认知项消失,而随机性惩罚则偏好低方差转移,且无需拟合显式下一状态预测器。伪计数提供认知价值,基于探测的惩罚捕捉随机方差,短时窗门控保护有信息量的后续状态。窗口冻结所有奖励定义对象,实现平稳贝尔曼算子、显式学习目标边界,并在混合性、平滑性、带宽与容量假设下获得非参数估计器的条件一致收敛结果。从主动推断角度看,代理在保留新颖性时无偏好,全可观测下标准似然歧义消失,引入非标准转移熵惩罚,且在状态空间已解析区域自然涌现惊喜最小化。
原文摘要 · Abstract (English)
Across environments with mixed sources of uncertainty, unsupervised reinforcement learning requires intrinsic motivation that does not precommit to a particular direction of surprise. Surprise minimization is scoped by design to ``unstable'' environments. Prediction-error curiosity rewards total expected surprise, including irreducible noise. Bandit or mixture switching between surprise-minimizing and surprise-maximizing rewards reintroduces non-stationarity by construction. We propose a single intrinsic reward, stationary within each window, derived from the novelty contribution of a preference-free Expected Free Energy objective, expressed in reward-maximization form. Our claim is that parameter information gain, the expected surprise of the next state minus its irreducible part, is the appropriate intrinsic signal in both high-entropy and low-entropy components of the state space. Maximizing it seeks exactly the surprise the model can explain away. In regions of unresolved dynamics, this epistemic term drives exploration. As dynamics become resolved, the epistemic term vanishes, while an aleatoric penalty favors lower-variance transitions, all without fitting an explicit next-state predictor. A pseudocount supplies epistemic value, a probe-based penalty captures aleatoric variance, and a short-horizon gate protects informative successors. A window-based freeze of all reward-defining objects yields a stationary Bellman operator, explicit bounds on learning targets, and a conditional uniform-concentration result for the nonparametric estimators under mixing, smoothness, bandwidth, and capacity assumptions. In active-inference terms, the agent is preference-free where novelty is retained, standard likelihood ambiguity vanishes under full observability, a nonstandard transition-entropy penalty is added, and surprise minimization emerges in resolved regions of the state space.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。