任意固定学习率下,梯度带访问算法几乎必然收敛到全局最优策略。
Small steps no more: Global convergence of stochastic gradient bandits for arbitrary learning rates
- 使用任意常数学习率,突破传统平滑性和噪声控制假设限制。
- 证明算法在探索与利用间保持平衡,实现全局收敛。
- 适合关注强化学习中简单优化方法理论性质的研究者。
我们对随机梯度带访问算法提供了新理解,证明其在使用任意常数学习率时,几乎必然收敛至全局最优策略。该结果表明,即使标准的光滑性与噪声控制假设不成立,该算法仍能恰当地平衡探索与利用。证明基于关于动作采样速率及累积进展与噪声关系的新发现,拓展了当前对简单随机梯度方法在带访问设置下行为的理解。
原文摘要 · Abstract (English)
We provide a new understanding of the stochastic gradient bandit algorithm by showing that it converges to a globally optimal policy almost surely using \emph{any} constant learning rate. This result demonstrates that the stochastic gradient algorithm continues to balance exploration and exploitation appropriately even in scenarios where standard smoothness and noise control assumptions break down. The proofs are based on novel findings about action sampling rates and the relationship between cumulative progress and noise, and extend the current understanding of how simple stochastic gradient methods behave in bandit settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。