arXiv:2507.03900cs.LGstat.ML2025-07被引 3

提出新框架优化静态谱风险度量,提升在线与离线强化学习的风险敏感性。

Risk-sensitive Actor-Critic with Static Spectral Risk Measures for Online and Offline Reinforcement Learning

  • 基于静态谱风险度量设计新型优化框架,灵活调整风险偏好。
  • 在多领域实验中,性能优于现有风险敏感方法,且稳定可靠。
  • 适用于在线与离线强化学习,理论保证收敛性,适合高风险场景应用。

分布强化学习(DRL)通过在价值函数中引入期望以外的风险度量,为基于价值和演员-评论家方法提供了自然的风险敏感机制。尽管该方法因简单性被广泛应用于多种在线与离线强化学习算法,但风险度量的直接集成常导致次优策略,尤其在需严格控制最坏结果的场景中危害显著。为此,本文提出一种新框架,用于优化静态谱风险度量(SRM),这一可灵活调整的风险度量族能统一处理如条件风险价值(CVaR)和均值-CVaR等目标,并支持定制风险偏好。所提方法适用于在线与离线强化学习算法。我们在有限状态-动作空间下建立了理论收敛性证明,并通过大量实验证明,该方法在多个领域中持续优于现有风险敏感方法,无论在线或离线环境。

原文摘要 · Abstract (English)

The development of Distributional Reinforcement Learning (DRL) has introduced a natural way to incorporate risk sensitivity into value-based and actor-critic methods by employing risk measures other than expectation in the value function. While this approach is widely adopted in many online and offline RL algorithms due to its simplicity, the naive integration of risk measures often results in suboptimal policies. This limitation can be particularly harmful in scenarios where the need for effective risk-sensitive policies is critical and worst-case outcomes carry severe consequences. To address this challenge, we propose a novel framework for optimizing static Spectral Risk Measures (SRM), a flexible family of risk measures that generalizes objectives such as CVaR and Mean-CVaR, and enables the tailoring of risk preferences. Our method is applicable to both online and offline RL algorithms. We establish theoretical guarantees by proving convergence in the finite state-action setting. Moreover, through extensive empirical evaluations, we demonstrate that our algorithms consistently outperform existing risk-sensitive methods in both online and offline environments across diverse domains.

强化学习风险敏感分布强化谱风险

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。