arXiv:2605.07857cs.LG2026-05被引 2

提出无需扰动转移的动态风险优化算法,可学习规避风险的策略。

Actor-Critic Algorithm for Dynamic Expectile and CVaR

论文配图:Actor-Critic Algorithm for Dynamic Expectile and CVaR
图 1 · 摘自论文原文
  • 基于softmax参数化设计无扰动的策略梯度
  • 利用可诱导性实现动态期望损失与条件风险值的模型无关学习
  • 适合需要稳定风险控制的强化学习场景

在随机策略下优化动态风险在策略更新和价值学习方面均具挑战性:前者通常需转移扰动,后者可能依赖模型。为此,我们提出一种在softmax策略参数化下无需转移扰动的代理策略梯度。进一步,通过利用可诱导性,发展了动态期望损失与条件风险值的模型无关价值学习方法。受期望SARSA与期望策略梯度启发,构建了一个模型无关的离线策略演员-评论家算法。在具备可验证风险规避行为的领域中,实验表明该算法能学习风险规避策略,并持续优于现有方法。

原文摘要 · Abstract (English)

Optimizing dynamic risk with stochastic policies is challenging in both policy updates and value learning. The former typically requires transition perturbation, while the latter may rely on model-based approaches. To address these challenges, we propose a surrogate policy gradient without transition perturbation under softmax policy parameterization. We further develop model-free value learning methods for dynamic expectile and conditional value-at-risk by leveraging elicitability. Finally, inspired by Expected SARSA and Expected Policy Gradient, a model-free off-policy actor-critic algorithm is constructed. Empirical results in domains with verifiable risk-averse behavior show that our algorithm can learn risk-averse policy and consistently outperforms other existing methods.

强化学习风险控制策略梯度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。