用信息几何统一解释探索激励机制,揭示计数与最大熵探索的本质联系。
An Information-Geometric Approach to Artificial Curiosity
- 基于信息几何的不变性约束,推导出内在奖励应为占据率倒数的严格凹函数。
- 通过信息测地线插值实现探索与利用平衡,仅需一个标量参数即可决定奖励形式。
- 特殊参数值恰好对应计数基和最大熵两种经典探索策略,理论统一性强。
稀疏奖励环境中的学习仍是强化学习的根本挑战。人工好奇心通过内在奖励引导探索,但其精确形式尚未明确。理想情况下,这些奖励应依赖于智能体对环境的信息量,且不依赖于具体表征——这一不变性正是信息几何的核心。我们证明,信息单调性及在智能体-环境交互下的不变性,唯一约束内在奖励必须为占据率倒数的严格凹函数。若要求奖励能实现合理的探索-利用权衡,则需通过占据流形上的信息测地线插值,这实质上将候选奖励限制为仅由一个标量参数决定。值得注意的是,该参数的特殊取值恰好对应于计数基探索和最大熵探索。该框架为内在奖励的设计提供了重要理论约束,并将基础探索方法整合进一个连贯模型。
原文摘要 · Abstract (English)
Learning in environments with sparse rewards remains a fundamental challenge in reinforcement learning. Artificial curiosity addresses this limitation through intrinsic rewards to guide exploration, however, the precise formulation of these rewards has remained elusive. Ideally, such rewards should depend on the agent's information about the environment, remaining agnostic to its representation -- an invariance central to information geometry. Leveraging this, we show that information monotonicity and invariance under the agent-environment interaction uniquely constrains intrinsic rewards to strictly concave functions of the reciprocal occupancy. Requiring these rewards to yield a principled exploration-exploitation trade-off, via information geodesic interpolation on the occupancy manifold, effectively limits the candidates to those determined by a scalar parameter. Remarkably, special values of this parameter are found to correspond to count-based and maximum entropy exploration. This framework provides important constraints to the engineering of intrinsic rewards while integrating foundational exploration methods into a single, cohesive model.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。