改进离线强化学习中奖励估计的置信界,更精准且无需调参。
Refined PAC-Bayes Bounds for Offline Bandits
- 通过事件空间离散化优化概率参数,提升边界紧致性。
- 提出两种无须调参的边界,分别基于霍夫丁与伯恩斯坦不等式。
- 理论逼近最优,适合关注理论严谨性的研究者。
本文针对随机多臂老虎机中的离线策略学习问题,提出了对经验奖励估计的改进概率界。在Seldin等人(2010)的PAC-Bayes边界基础上,结合Rodríguez等人(2024)提出的新型参数优化方法——该方法通过对可能事件空间进行离散化以优化“以概率成立”的参数——实现了更紧的边界。我们推导出两个无需调参的PAC-Bayes边界:一个基于Hoeffding-Azuma不等式,另一个基于Bernstein不等式。理论证明表明,我们的边界几乎达到最优,其收敛速率等价于在数据观测后才设定该概率参数的情形。
原文摘要 · Abstract (English)
In this paper, we present refined probabilistic bounds on empirical reward estimates for off-policy learning in bandit problems. We build on the PAC-Bayesian bounds from Seldin et al. (2010) and improve on their results using a new parameter optimization approach introduced by Rodríguez et al. (2024). This technique is based on a discretization of the space of possible events to optimize the "in probability" parameter. We provide two parameter-free PAC-Bayes bounds, one based on Hoeffding-Azuma's inequality and the other based on Bernstein's inequality. We prove that our bounds are almost optimal as they recover the same rate as would be obtained by setting the "in probability" parameter after the realization of the data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。