arXiv:2603.10219stat.MLcs.AI2026-03被引 2
为随机老虎机设计的策略梯度方法,首次给出连续时间扩散分析。
A Diffusion Analysis of Policy Gradient for Stochastic Bandits
- 用连续时间扩散模型分析策略梯度在多臂老虎机上的行为。
- 证明学习率η=O(Δ²/log n)时,累积损失为O(k log k log n / η)。
- 构造反例:若η不满足O(Δ²),即使只有对数级臂数也会产生线性损失。
我们研究了k臂随机老虎机中策略梯度的连续时间扩散近似。证明当学习率η = O(Δ²/log(n))时,累积损失(regret)为O(k log(k) log(n) / η),其中n为决策周期长度,Δ为最小奖励差距。此外,我们构造了一个仅含对数级臂数的实例,证明除非η = O(Δ²),否则损失将呈线性增长。
原文摘要 · Abstract (English)
We study a continuous-time diffusion approximation of policy gradient for $k$-armed stochastic bandits. We prove that with a learning rate $η= O(Δ^2/\log(n))$ the regret is $O(k \log(k) \log(n) / η)$ where $n$ is the horizon and $Δ$ the minimum gap. Moreover, we construct an instance with only logarithmically many arms for which the regret is linear unless $η= O(Δ^2)$.
强化学习策略梯度多臂老虎机扩散模型
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。