发现贝叶斯深度Q学习中的先验偏差,提出改进方案提升性能
Priors Matter: Addressing Misspecification in Bayesian Deep Q-Learning
- 挑战传统高斯似然和先验假设,发现其常不成立
- 实证显示后验温度降低反使性能提升,存在冷后验效应
- 提供可落地的先验优化方法,推动更鲁棒的贝叶斯强化学习
强化学习中的不确定性量化可显著提升探索效率与鲁棒性。尽管近来基于近似贝叶斯的方法在无模型算法中流行,但研究多集中于后验近似的精度,而忽视了支撑后验的先验与似然假设本身的准确性。本文揭示了贝叶斯深度Q学习中存在冷后验效应:与理论相反,降低后验温度反而提升性能。通过统计检验,我们发现常见的高斯似然假设在实践中常被违反。因此,我们主张未来研究应优先关注设计更合适的似然函数与先验分布,并提出了简单可行的先验改进方案,在深度Q学习中实现更优的贝叶斯算法表现。
原文摘要 · Abstract (English)
Uncertainty quantification in reinforcement learning can greatly improve exploration and robustness. Approximate Bayesian approaches have recently been popularized to quantify uncertainty in model-free algorithms. However, so far the focus has been on improving the accuracy of the posterior approximation, instead of studying the accuracy of the prior and likelihood assumptions underlying the posterior. In this work, we demonstrate that there is a cold posterior effect in Bayesian deep Q-learning, where contrary to theory, performance increases when reducing the temperature of the posterior. To identify and overcome likely causes, we challenge common assumptions made on the likelihood and priors in Bayesian model-free algorithms. We empirically study prior distributions and show through statistical tests that the common Gaussian likelihood assumption is frequently violated. We argue that developing more suitable likelihoods and priors should be a key focus in future Bayesian reinforcement learning research and we offer simple, implementable solutions for better priors in deep Q-learning that lead to more performant Bayesian algorithms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。