arXiv:2608.01917cs.LG2026-08

首次给出折扣指数效用强化学习的有限时间收敛速率,无需调参。

Finite-Time Analysis of Discounted Exponential-Utility Reinforcement Learning

论文配图:Finite-Time Analysis of Discounted Exponential-Utility Reinforcement Learning
图 1 · 摘自论文原文
  • 提出无参数步长的单/双时标算法,解决非线性效用优化难题。
  • 在异步马尔可夫采样下实现1/√n的有限时间收敛率。
  • 适用于风险敏感决策场景,适合对理论严谨性有要求的研究者。

折扣指数效用为风险敏感的序列决策提供了合理的准则,但其非线性结构使强化学习面临挑战。近期工作引入贝尔曼相容的代理函数和两种无模型固定点算法,用于优化平稳策略下的该效用,但其主要收敛结果为渐近性质。本文在异步马尔可夫采样下,建立了前述两种算法的有限时间收敛速率,为$ ilde{O}(1/ ext{sqrt}{n})$,其中$n$为迭代次数,$ ilde{O}$忽略对数项。关键创新在于采用无参数步长设计。对于更简单的单时标方法,其更新方程与底层幂律算子的收缩几何不匹配;通过利用算子的有界性、单调性和齐次性,我们推导出相对误差动态的局部伪收缩性质。结合Moreau包络的李雅普诺夫函数与Polyak-Ruppert平均,实现了无参数步长下的收敛速率。对于双时标方法,核心挑战在于控制快速时标上的追踪误差。这些结果首次为无模型折扣指数效用强化学习提供了有限时间保证。

原文摘要 · Abstract (English)

Discounted exponential utility provides a principled criterion for risk-sensitive sequential decision-making, but its nonlinear structure complicates reinforcement learning. A recent work \citep{thoppe2026reinforcement} addressed this difficulty by introducing a Bellman-compatible surrogate and two model-free fixed-point algorithms for optimizing it over stationary policies. However, their main convergence results are asymptotic. In this work, we establish finite-time rates of $\tilde{O} (1/\sqrt{n})$ for the aforementioned two algorithms under asynchronous Markovian sampling, where $n$ is the iteration index and $\tilde{O}$ hides logarithmic expressions. Importantly, we employ parameter-free choices for the stepsize parameter to derive these rate results. For the algorithmically simpler one-timescale method, the main challenge is that its update equation is not directly aligned with the contraction geometry of its underlying power-law operator. We overcome this mismatch by exploiting the boundedness, monotonicity, and homogeneity of the operator to obtain a local pseudo-contraction property for the relative-error dynamics. We then use a Moreau-envelope-based Lyapunov function and Polyak--Ruppert averaging to obtain the stated convergence rate with parameter-free stepsizes. For the two-timescale method, the main challenge is to control a tracking error on the faster timescale. These results provide the first finite-time guarantees for model-free discounted exponential-utility reinforcement learning.

强化学习收敛分析风险敏感无参数

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。