arXiv:2603.26547cs.LG2026-03被引 1

为随机博弈的软最大策略梯度提供稳定分析,优化学习率以降低后悔值。

A Lyapunov Analysis of Softmax Policy Gradient for Stochastic Bandits

  • 基于李雅普诺夫方法分析离散时间下的软最大策略梯度
  • 在最优学习率下,后悔值达到 $O(k \log k \log n / η)$
  • 适用于需理论保障的强化学习初学者与算法设计者

我们将Lattimore(2026)对连续时间k臂随机博弈策略梯度的分析方法,适配到标准离散时间框架。与连续时间情形类似,证明了当学习率 $η= O(Δ_{ ext{min}}^2/(Δ_{ ext{max}} \log(n)))$ 时,后悔值为 $O(k \log(k) \log(n) / η)$,其中 $n$ 为决策周期数,$Δ_{ ext{min}}$ 和 $Δ_{ ext{max}}$ 分别为最小与最大奖励差距。

原文摘要 · Abstract (English)

We adapt the analysis of policy gradient for continuous time $k$-armed stochastic bandits by Lattimore (2026) to the standard discrete time setup. As in continuous time, we prove that with learning rate $η= O(Δ_{\min}^2/(Δ_{\max} \log(n)))$ the regret is $O(k \log(k) \log(n) / η)$ where $n$ is the horizon and $Δ_{\min}$ and $Δ_{\max}$ are the minimum and maximum gaps.

强化学习策略梯度后悔分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。