将行为经济学中的前景理论引入强化学习,提升决策鲁棒性。
Policy Gradients for Cumulative Prospect Theory in Reinforcement Learning
- 基于顺序统计量设计新型梯度估计器,支持前景理论目标优化。
- 证明算法渐近收敛至非凸目标的一阶驻点,具备理论保障。
- 适用于风险敏感场景,适合关注决策稳健性的研究者。
我们推导了有限时域强化学习中累积前景理论(CPT)目标的策略梯度定理,该定理推广了标准策略梯度,并将基于扭曲的风险目标作为特例包含在内。受行为经济学启发,CPT结合了以参考点为中心的非对称效用变换与概率扭曲。基于该定理,我们设计了一种一阶策略梯度算法,采用基于顺序统计量的蒙特卡洛梯度估计器。我们建立了该估计器的统计性质,并证明所提出的算法能渐近收敛至(通常非凸的)CPT目标的一阶驻点。仿真结果展示了由CPT诱导的定性行为,并将我们的方法与现有的零阶方法进行了对比。
原文摘要 · Abstract (English)
We derive a policy gradient theorem for Cumulative Prospect Theory (CPT) objectives in finite-horizon Reinforcement Learning (RL), generalizing the standard policy gradient theorem and encompassing distortion-based risk objectives as special cases. Motivated by behavioral economics, CPT combines an asymmetric utility transformation around a reference point with probability distortion. Building on our theorem, we design a first-order policy gradient algorithm for CPT-RL using a Monte Carlo gradient estimator based on order statistics. We establish statistical guarantees for the estimator and prove asymptotic convergence of the resulting algorithm to first-order stationary points of the (generally non-convex) CPT objective. Simulations illustrate qualitative behaviors induced by CPT and compare our first-order approach to existing zeroth-order methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。