基于贝叶斯方法学习随机最短路径的最优策略,无需不切实际假设。
Bayesian learning for the stochastic shortest path problem

- 直接用贝尔曼方程构建$Q^*$后验分布,避免传统近似。
- 在深海基准测试中验证,比其他方法更高效且准确量化不确定性。
- 适合研究不确定性建模与数据效率的强化学习学者。
序列决策问题常被建模为马尔可夫决策过程(MDP)。本文聚焦于随机最短路径(SSP)问题,即具有吸收终止状态的无限时域无折扣MDP。我们提出一种贝叶斯框架,通过与任务交互学习最优决策策略。具体地,我们直接学习最优动作值函数$Q^*$,而不依赖不现实的建模假设或启发式近似。方法基于贝尔曼最优性方程构造$Q^*$的后验信念。对于确定性奖励,后验分布具有流形密度;为简化推断,放松似然以引入勒贝格密度,但带来不可识别性问题——松弛后的后验可能在非正当决策规则上具有显著质量,而精确后验不会。我们还计算了$Q^*$的表格参数化、高斯似然松弛和高斯先验下的精确最优动作选择后验概率,对基准研究有实用价值。在深度海平面变体上的数值实验验证了结论。结果表明,该框架能忠实量化不确定性,且相比其他基于时序差分的贝叶斯方法更具数据效率。最后给出未来工作建议。
原文摘要 · Abstract (English)
Sequential decision-making problems are often modelled as a Markov decision process (MDP). We focus on the stochastic shortest path (SSP) problem, which is an infinite-horizon undiscounted MDP with absorbing terminal states. We develop a Bayesian framework to learn the optimal decision strategy through interactions with the decision-making task. Specifically, we learn the optimal action-value function $Q^*$, but unlike many existing Bayesian approaches, we do not rely on unrealistic modelling assumptions and ad-hoc approximations. Our approach is to directly construct the posterior beliefs for $Q^*$ through Bellman's optimality equations. For deterministic rewards, we characterise the posterior as a distribution with a manifold density. To facilitate simpler inference, we relax the likelihood so that a Lebesgue density exists. The flip side is to create unidentifiability issues. Specifically, the relaxed posterior can have significant mass on improper decision rules, while the exact posterior will not. We also calculate the exact posterior probabilities for optimal action selections for the tabular parametrisation of $Q^*$, a Gaussian likelihood relaxation and a Gaussian prior, which is useful in benchmarking studies. Numerical studies on variants of the Deep Sea benchmark verify our findings. We demonstrate that our framework faithfully quantifies uncertainty and, compared to other temporal-difference-based Bayesian methodologies, is more data efficient. We conclude with recommendations for future work.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。