arXiv:2602.07202cs.LG2026-02AAAI

提出稳定算法,让智能体在复杂任务中更安全地应对风险

Risk-Sensitive Exponential Actor Critic

  • 基于熵风险度量设计新型策略梯度方法
  • 在MuJoCo连续控制任务中可靠学习风险敏感策略
  • 适合需要安全决策的现实场景应用

无模型深度强化学习在诸多挑战性任务中取得巨大成功,但在真实应用中仍存在安全顾虑,亟需具备风险意识的智能体。当前用于训练此类智能体的常见效用函数是熵风险度量,但现有策略梯度方法在此度量下更新方差高且数值不稳定。因此,已有风险敏感的无模型方法仅限于简单任务和表格型设置。本文为熵风险度量上的策略梯度方法提供了完整的理论依据,涵盖随机与确定性策略下的在线与离线情形。受理论启发,我们提出风险敏感指数演员-评论家(rsEAC),一种离线无模型方法,通过新颖机制避免显式表示指数价值函数及其梯度,并针对熵风险度量优化策略。实验表明,rsEAC相比现有方法具有更稳定的数值更新,在MuJoCo中复杂的高风险连续任务上可可靠学习风险敏感策略。

原文摘要 · Abstract (English)

Model-free deep reinforcement learning (RL) algorithms have achieved tremendous success on a range of challenging tasks. However, safety concerns remain when these methods are deployed on real-world applications, necessitating risk-aware agents. A common utility for learning such risk-aware agents is the entropic risk measure, but current policy gradient methods optimizing this measure must perform high-variance and numerically unstable updates. As a result, existing risk-sensitive model-free approaches are limited to simple tasks and tabular settings. In this paper, we provide a comprehensive theoretical justification for policy gradient methods on the entropic risk measure, including on- and off-policy gradient theorems for the stochastic and deterministic policy settings. Motivated by theory, we propose risk-sensitive exponential actor-critic (rsEAC), an off-policy model-free approach that incorporates novel procedures to avoid the explicit representation of exponential value functions and their gradients, and optimizes its policy w.r.t the entropic risk measure. We show that rsEAC produces more numerically stable updates compared to existing approaches and reliably learns risk-sensitive policies in challenging risky variants of continuous tasks in MuJoCo.

强化学习风险敏感演员评论家安全决策

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。