arXiv:2509.14077cs.LG2025-09被引 1

用贝叶斯方法让强化学习更谨慎,避免因数据少而高估风险。

Online Bayesian Risk-Averse Reinforcement Learning

  • 基于后验采样构建在线学习算法,动态适应不确定性。
  • 理论证明后悔率亚线性下降,数据越多越稳健。
  • 适合数据稀缺场景下的安全决策问题,如医疗、金融。

本文研究强化学习中的贝叶斯风险规避框架。为应对数据不足导致的认知不确定性,采用贝叶斯风险马尔可夫决策过程(BRMDP)建模未知模型参数的不确定性。推导出贝叶斯风险价值函数与真实分布下原价值函数之差的渐近正态性,结果表明贝叶斯风险规避方法倾向于悲观低估原始价值函数,且该偏差随风险规避程度增强而增大,随数据增多而减小。利用此自适应特性,我们在在线强化学习及在线上下文多臂老虎机(CMAB,RL特例)中设计两种后验采样算法。在标准后悔定义下,均建立亚线性后悔界;对CMAB还建立了以贝叶斯风险后悔定义的亚线性边界。数值实验验证了算法有效性和理论性质。

原文摘要 · Abstract (English)

In this paper, we study the Bayesian risk-averse formulation in reinforcement learning (RL). To address the epistemic uncertainty due to a lack of data, we adopt the Bayesian Risk Markov Decision Process (BRMDP) to account for the parameter uncertainty of the unknown underlying model. We derive the asymptotic normality that characterizes the difference between the Bayesian risk value function and the original value function under the true unknown distribution. The results indicate that the Bayesian risk-averse approach tends to pessimistically underestimate the original value function. This discrepancy increases with stronger risk aversion and decreases as more data become available. We then utilize this adaptive property in the setting of online RL as well as online contextual multi-arm bandits (CMAB), a special case of online RL. We provide two procedures using posterior sampling for both the general RL problem and the CMAB problem. We establish a sub-linear regret bound, with the regret defined as the conventional regret for both the RL and CMAB settings. Additionally, we establish a sub-linear regret bound for the CMAB setting with the regret defined as the Bayesian risk regret. Finally, we conduct numerical experiments to demonstrate the effectiveness of the proposed algorithm in addressing epistemic uncertainty and verifying the theoretical properties.

强化学习贝叶斯方法风险规避在线学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。