用信息增益理论提升强化学习探索效率,让智能体专注填补知识空白。
On Efficient Bayesian Exploration in Model-Based Reinforcement Learning
- 基于信息增益设计探索奖励,聚焦未知而非环境噪声。
- 理论证明奖励随认知确定性提升而归零,确保探索合理终止。
- 结合贝叶斯建模与轨迹采样,适合稀疏奖励或纯探索任务。
本文针对模型化强化学习中的数据高效探索挑战,研究了基于信息论的内在动机方法。重点分析一类针对认知不确定性(epistemic uncertainty)而非环境固有随机性(aleatoric noise)的探索奖励。我们证明,此类奖励天然反映认知信息增益,并在智能体对环境动态和奖励足够确信时收敛至零,从而实现探索与真实知识缺口的对齐。该分析为基于信息增益的方法提供了首个理论保证。为实现实际应用,我们探讨了通过稀疏变分高斯过程、深度核函数及深度集成模型进行可计算近似。进而提出通用框架:基于贝叶斯探索的预测轨迹采样(PTS-BE),融合模型规划与信息论奖励,实现样本高效的深度探索。实验表明,PTS-BE在多种稀疏奖励及纯探索任务环境中显著优于多个基线方法。
原文摘要 · Abstract (English)
In this work, we address the challenge of data-efficient exploration in reinforcement learning by examining existing principled, information-theoretic approaches to intrinsic motivation. Specifically, we focus on a class of exploration bonuses that targets epistemic uncertainty rather than the aleatoric noise inherent in the environment. We prove that these bonuses naturally signal epistemic information gains and converge to zero once the agent becomes sufficiently certain about the environment's dynamics and rewards, thereby aligning exploration with genuine knowledge gaps. Our analysis provides formal guarantees for IG-based approaches, which previously lacked theoretical grounding. To enable practical use, we also discuss tractable approximations via sparse variational Gaussian Processes, Deep Kernels and Deep Ensemble models. We then outline a general framework - Predictive Trajectory Sampling with Bayesian Exploration (PTS-BE) - which integrates model-based planning with information-theoretic bonuses to achieve sample-efficient deep exploration. We empirically demonstrate that PTS-BE substantially outperforms other baselines across a variety of environments characterized by sparse rewards and/or purely exploratory tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。