用认知不确定性指导探索,让强化学习更高效可靠。
EUBRL: Epistemic Uncertainty Directed Bayesian Reinforcement Learning
- 基于认知不确定性动态调整探索策略
- 在稀疏奖励任务中样本效率显著提升
- 适合高不确定性、长周期的复杂决策场景
在已知与未知的边界上,智能体面临探索与利用的权衡。认知不确定性反映了因知识有限而产生的系统性不确定。本文提出一种贝叶斯强化学习算法EUBRL,利用认知不确定性引导探索,从而自适应地降低因估计误差带来的每步损失。我们为一类充分表达能力强的先验建立了近乎极小极大最优的后悔值和样本复杂度保证,适用于无限时域折扣马尔可夫决策过程。实验在具有稀疏奖励、长时域和随机性的任务上验证了EUBRL,结果表明其在样本效率、可扩展性和一致性方面均表现优异。
原文摘要 · Abstract (English)
At the boundary between the known and the unknown, an agent inevitably confronts the dilemma of whether to explore or to exploit. Epistemic uncertainty reflects such boundaries, representing systematic uncertainty due to limited knowledge. In this paper, we propose a Bayesian reinforcement learning (RL) algorithm, $\texttt{EUBRL}$, which leverages epistemic guidance to achieve principled exploration. This guidance adaptively reduces per-step regret arising from estimation errors. We establish nearly minimax-optimal regret and sample complexity guarantees for a class of sufficiently expressive priors in infinite-horizon discounted MDPs. Empirically, we evaluate $\texttt{EUBRL}$ on tasks characterized by sparse rewards, long horizons, and stochasticity. Results demonstrate that $\texttt{EUBRL}$ achieves superior sample efficiency, scalability, and consistency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。