arXiv:2501.02087cs.LGstat.ML2025-01ICML被引 9

用更通用的风险度量提升强化学习决策能力,避免过度保守。

Beyond CVaR: Leveraging Static Spectral Risk Measures for Enhanced Decision-Making in Distributional Reinforcement Learning

  • 用静态谱风险度量替代固定风险度量,优化决策过程
  • 实验显示新模型在多种场景下优于现有风险中性与敏感模型
  • 方法可解释性强,适合金融、医疗等高风险领域应用

在金融、医疗和机器人等领域,管理最坏情况至关重要,失败可能导致灾难性后果。分布强化学习(DRL)为引入风险敏感性提供了自然框架。然而,现有方法存在两大局限:(1) 每个决策步骤使用固定风险度量,常导致策略过于保守;(2) 学习策略的解释性和理论性质不清晰。尽管优化静态风险度量可解决这些问题,其在DRL中的应用仍局限于简单的静态CVaR。本文提出一种具有收敛保证的新DRL算法,可优化更广泛的静态谱风险度量(SRM)。同时,通过利用DRL中的回报分布和静态一致风险度量的分解,给出了学习策略的清晰解释。大量实验表明,该模型所学策略与SRM目标一致,并在多种设置下超越现有的风险中性与风险敏感型DRL模型。

原文摘要 · Abstract (English)

In domains such as finance, healthcare, and robotics, managing worst-case scenarios is critical, as failure to do so can lead to catastrophic outcomes. Distributional Reinforcement Learning (DRL) provides a natural framework to incorporate risk sensitivity into decision-making processes. However, existing approaches face two key limitations: (1) the use of fixed risk measures at each decision step often results in overly conservative policies, and (2) the interpretation and theoretical properties of the learned policies remain unclear. While optimizing a static risk measure addresses these issues, its use in the DRL framework has been limited to the simple static CVaR risk measure. In this paper, we present a novel DRL algorithm with convergence guarantees that optimizes for a broader class of static Spectral Risk Measures (SRM). Additionally, we provide a clear interpretation of the learned policy by leveraging the distribution of returns in DRL and the decomposition of static coherent risk measures. Extensive experiments demonstrate that our model learns policies aligned with the SRM objective, and outperforms existing risk-neutral and risk-sensitive DRL models in various settings.

强化学习风险控制分布学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。