arXiv:2502.17370stat.MLcs.LG2025-02被引 1
改进UCBVI的奖励项与误差分析,提升算法实际表现
A Refined Analysis of UCBVI
- 优化奖励项设计,更精准刻画不确定性
- 理论误差上界常数显著降低,实测性能更好
- 适合关注强化学习理论改进与实用性能的研究者
本文对UCBVI算法(Azar等,2017)进行了精细化分析,改进了其奖励项与后悔率分析。同时,将改进后的版本与原始版本及当前最先进的MVP算法进行对比。实验验证表明,降低理论界中的乘性常数对算法的实测性能有显著正向影响。
原文摘要 · Abstract (English)
In this work, we provide a refined analysis of the UCBVI algorithm (Azar et al., 2017), improving both the bonus terms and the regret analysis. Additionally, we compare our version of UCBVI with both its original version and the state-of-the-art MVP algorithm. Our empirical validation demonstrates that improving the multiplicative constants in the bounds has significant positive effects on the empirical performance of the algorithms.
强化学习后悔分析算法改进
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。