arXiv:2512.24580q-fin.RMcs.LG2025-12

提出一种兼顾风险敏感与模型不确定性的强化学习新框架

Robust Bayesian Dynamic Programming for On-policy Risk-sensitive Reinforcement Learning

  • 用内外双重风险度量统一处理状态随机性与转移不确定性
  • 算法在训练环境中收敛至近优策略,具备理论保证的样本与计算复杂度
  • 适用于金融对冲等需兼顾风险控制与鲁棒性的高要求场景

我们提出一种新型风险敏感强化学习框架,以应对转移过程中的不确定性。定义了两种相互关联的风险度量:内层风险度量用于处理状态和成本的随机性,外层风险度量则捕捉转移动态的不确定性。该框架通过允许内、外风险度量采用一般一致风险测度,统一并推广了现有大多数强化学习框架。在此基础上,构建了风险敏感鲁棒马尔可夫决策过程(RSRMDP),推导其贝尔曼方程,并在给定后验分布下提供误差分析。进一步设计了一种贝叶斯动态规划(Bayesian DP)算法,交替进行后验更新与值迭代。该方法使用结合蒙特卡洛采样与凸优化的风险贝尔曼算子估计器,证明了强一致性。此外,我们分析了在狄利克雷后验和条件风险价值(CVaR)下的样本复杂度与计算复杂度。两个数值实验验证了算法出色的收敛性,并直观展示了其在风险敏感性与鲁棒性方面的优势。实证还表明该算法在期权对冲任务中表现优异。

原文摘要 · Abstract (English)

We propose a novel framework for risk-sensitive reinforcement learning (RSRL) that incorporates robustness against transition uncertainty. We define two distinct yet coupled risk measures: an inner risk measure addressing state and cost randomness and an outer risk measure capturing transition dynamics uncertainty. Our framework unifies and generalizes most existing RL frameworks by permitting general coherent risk measures for both inner and outer risk measures. Within this framework, we construct a risk-sensitive robust Markov decision process (RSRMDP), derive its Bellman equation, and provide error analysis under a given posterior distribution. We further develop a Bayesian Dynamic Programming (Bayesian DP) algorithm that alternates between posterior updates and value iteration. The approach employs an estimator for the risk-based Bellman operator that combines Monte Carlo sampling with convex optimization, for which we prove strong consistency guarantees. Furthermore, we demonstrate that the algorithm converges to a near-optimal policy in the training environment and analyze both the sample complexity and the computational complexity under the Dirichlet posterior and CVaR. Finally, we validate our approach through two numerical experiments. The results exhibit excellent convergence properties while providing intuitive demonstrations of its advantages in both risk-sensitivity and robustness. Empirically, we further demonstrate the advantages of the proposed algorithm through an application on option hedging.

强化学习风险敏感贝叶斯方法鲁棒控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。