通过量化梯度提升风险规避策略优化的样本效率
Boosting CVaR Policy Optimization with Quantile Gradients
- 引入期望分位数项改进CVaR优化,利用全部采样轨迹
- 在马尔可夫策略类中显著优于传统CVaR-PG方法
- 适合需要高鲁棒性决策的强化学习应用
使用策略梯度优化条件风险价值(CVaR)面临样本效率低的问题,因其仅关注尾部表现而忽略多数采样轨迹。本文通过引入期望分位数项来增强CVaR,使分位数优化具备动态规划形式,从而利用所有采样数据,提升样本效率。该改进不改变原始CVaR目标,因CVaR即尾部分位数的期望。在具有可验证风险规避行为的领域中,实验表明该算法在马尔可夫策略类中显著优于CVaR-PG,并持续超越其他现有方法。
原文摘要 · Abstract (English)
Optimizing Conditional Value-at-risk (CVaR) using policy gradient (a.k.a CVaR-PG) faces significant challenges of sample inefficiency. This inefficiency stems from the fact that it focuses on tail-end performance and overlooks many sampled trajectories. We address this problem by augmenting CVaR with an expected quantile term. Quantile optimization admits a dynamic programming formulation that leverages all sampled data, thus improves sample efficiency. This does not alter the CVaR objective since CVaR corresponds to the expectation of quantile over the tail. Empirical results in domains with verifiable risk-averse behavior show that our algorithm within the Markovian policy class substantially improves upon CVaR-PG and consistently outperforms other existing methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。