用分位数贝叶斯风险模型平衡探索与鲁棒性,早期抗不确定性,后期主动探索。
Evolving Robustness--Exploration Trade-off in Online Reinforcement Learning via Quantile Bayesian Risk MDPs

- 设计分位数贝叶斯风险MDP,用量化水平调节对不确定性的乐观/悲观态度。
- 理论证明:随着数据积累,乐观/悲观程度递减,实现动态权衡。
- 自适应分位数调度算法,适合探索需求高或成本高的复杂环境。
在线强化学习中,数据稀缺导致认知不确定性,使早期鲁棒性至关重要,但充分探索又需学习真实环境的最优策略。本文通过分位数贝叶斯风险马尔可夫决策过程(BR-MDP),研究这一随时间变化的鲁棒性-探索权衡,其中分位数水平控制后验不确定性如何影响贝尔曼更新。我们通过渐近正态性结果刻画该控制机制:上尾/下尾分位数分别诱导对认知不确定性的乐观/悲观态度,且其强度随数据积累而降低。基于此,提出一种在线贝叶斯风险感知算法,采用自适应分位数调度,在早期强调鲁棒性,逐步鼓励对未访问状态-动作对的探索。理论证明了相对于真实最优值和最优BR-MDP鲁棒值的次线性贝叶斯后悔界。数值实验表明,在探索需求高与探索成本高的环境中均表现优异。
原文摘要 · Abstract (English)
In online reinforcement learning, data scarcity creates epistemic uncertainty that makes robustness important early in learning, whereas sufficient exploration is needed to learn the true-environment optimal policy. We study this time-varying robustness--exploration trade-off through a quantile Bayesian risk-aware Markov decision process (BR-MDP), in which the quantile level controls how posterior uncertainty enters the Bellman backup. We characterize this control through an asymptotic normality result for the difference between the quantile BR-MDP value and the value in the true environment. The result implies that upper/lower-tail quantiles induce optimism/pessimism towards epistemic uncertainty, and the magnitude of the optimism/pessimism decreases as data accumulate. Building on this characterization, we propose an online Bayesian risk-aware algorithm with an adaptive quantile schedule that emphasizes robustness early and gradually encourages exploration of less-visited state--action pairs. We establish sublinear Bayesian regret bounds with respect to both the true optimal value and the optimal BR-MDP robust value. Numerical experiments demonstrate strong performance in both exploration-demanding and exploration-costly environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。