arXiv:2605.08053cs.LG2026-05被引 1

为指数效用强化学习设计收敛算法,解决风险规避下的决策优化问题。

Reinforcement Learning for Exponential Utility: Algorithms and Convergence in Discounted MDPs

  • 基于贝尔曼方程构造两类Q值扩展,确保算子收缩性。
  • 提出双时标与单时标算法,分别实现几乎必然收敛与有限时间分析。
  • 首次在向量情形下完成非标准度量下的收敛证明,适合风险敏感研究者。

在折扣马尔可夫决策过程(MDP)中,针对固定风险规避下的指数效用优化,现有强化学习方法缺乏系统性的基于价值的算法。本文基于 extcite{porteus1975optimality} 研究的指数效用贝尔曼方程,推导出两种类Q值扩展,并证明其对应的算子分别在 $L_ ty$ 和 sup-log/Thompson 度量下为压缩映射。我们刻画了其不动点,并证明由贪婪策略诱导的平稳策略在平稳策略集中对指数效用目标是最优的。由此导出两个无模型算法:一个双时标Q学习风格算法,通过时标分离证明几乎必然收敛并给出有限时间速率;另一个单时标算法由次线性幂律算子控制。由于后者在标准度量下不具全局压缩性,我们通过局部利普希茨性、单调性、齐次性及狄尼导数等精细论证证明其收敛性,并提供标量情形下的有限时间分析,揭示向量情形下获取收敛速率的挑战。本工作为指数效用目标下的价值型强化学习奠定了基础。

原文摘要 · Abstract (English)

Reinforcement learning (RL) for exponential-utility optimization in discounted Markov decision processes (MDPs) lacks principled value-based algorithms. We address this gap in the fixed risk-aversion setting. Building on the Bellman-type equation for exponential utility studied in \cite{porteus1975optimality}, we derive two Q-value-style extensions and show that the associated operators are contractions in the $L_\infty$ and sup-log/Thompson metrics, respectively. We characterize their fixed points and prove that the induced greedy stationary policy is optimal for the exponential-utility objective among stationary policies. These structural results lead to two model-free algorithms: a two-timescale Q-learning--style algorithm, for which we establish almost-sure convergence and provide finite-time convergence rates via timescale separation, and a one-timescale algorithm governed by a sublinear power-law operator. Since the latter does not admit a global contraction in standard metrics, we prove its convergence using delicate arguments based on local Lipschitzness, monotonicity, homogeneity, and Dini derivatives, and provide a scalar finite-time analysis that highlights the challenges in obtaining convergence rates in the vector case. Our work provides a foundation for value-based RL under exponential-utility objectives.

强化学习指数效用收敛分析风险规避

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。