arXiv:2504.11412cs.LGcs.AI2025-04

系统研究九种变异性度量,提升风险敏感强化学习的稳定性与性能。

Measures of Variability for Risk-averse Policy Gradient

  • 提出四种新度量的策略梯度公式,改进已有指标梯度估计。
  • 实验证明CVaR偏差和吉尼偏差在不同场景下表现稳定且收益高。
  • 适合高风险决策领域研究者参考,推动风险度量发展。

风险敏感强化学习(RARL)在不确定性环境下的决策中至关重要,尤其适用于高风险场景。然而,现有工作多集中于风险度量(如条件风险价值,CVaR),而对变异性度量的研究仍不充分。本文系统研究了九种常见变异性度量:方差、吉尼离差、均值离差、均值-中位数离差、标准差、分位数间距、CVaR离差、半方差和半标准差,其中四种此前未在RARL中被研究。我们推导了这些度量的策略梯度公式,改进了吉尼离差的梯度估计,分析了其梯度性质,并将其集成至REINFORCE与PPO框架中以惩罚回报的离散性。实验表明,基于方差的度量会导致策略更新不稳定;而CVaR离差和吉尼离差在不同随机性和评估域下表现一致,实现高回报并有效学习风险规避策略。均值离差和半标准差也表现出良好适应性。本工作为RARL中的变异性度量提供了全面综述,为风险感知决策提供实践指导,并引导未来风险度量与算法研究。

原文摘要 · Abstract (English)

Risk-averse reinforcement learning (RARL) is critical for decision-making under uncertainty, which is especially valuable in high-stake applications. However, most existing works focus on risk measures, e.g., conditional value-at-risk (CVaR), while measures of variability remain underexplored. In this paper, we comprehensively study nine common measures of variability, namely Variance, Gini Deviation, Mean Deviation, Mean-Median Deviation, Standard Deviation, Inter-Quantile Range, CVaR Deviation, Semi_Variance, and Semi_Standard Deviation. Among them, four metrics have not been previously studied in RARL. We derive policy gradient formulas for these unstudied metrics, improve gradient estimation for Gini Deviation, analyze their gradient properties, and incorporate them with the REINFORCE and PPO frameworks to penalize the dispersion of returns. Our empirical study reveals that variance-based metrics lead to unstable policy updates. In contrast, CVaR Deviation and Gini Deviation show consistent performance across different randomness and evaluation domains, achieving high returns while effectively learning risk-averse policies. Mean Deviation and Semi_Standard Deviation are also competitive across different scenarios. This work provides a comprehensive overview of variability measures in RARL, offering practical insights for risk-aware decision-making and guiding future research on risk metrics and RARL algorithms.

强化学习风险敏感策略梯度变异性度量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。