arXiv:2503.04855stat.MLcs.LG2025-03被引 3

揭示UCB算法在双臂赌博机中的自适应采样机制

A characterization of sample adaptivity in UCB data

  • 通过扰动分析建立新型联合中心极限定理
  • 发现采样次数与奖励均值的分布随臂间差距平滑变化
  • 为样本偏差提供首项主导的启发式解释,适合强化学习研究者

我们刻画了在随机双臂赌博机环境下,基于UCB算法时各臂的采样次数与样本平均奖励的联合中心极限定理。该结果具有多重含义:(1) 推导出一种非标准的采样次数和伪遗憾的中心极限定理,其形式在臂间差距较大时呈现标准形式,在差距较小时表现为慢收敛形式,实现平滑过渡;(2) 从采样次数与样本均值间的相关性出发,启发式地推导出样本偏差的首项主导项。本研究的分析框架基于一种新颖的扰动分析方法,本身也具有广泛的研究价值。

原文摘要 · Abstract (English)

We characterize a joint CLT of the number of pulls and the sample mean reward of the arms in a stochastic two-armed bandit environment under UCB algorithms. Several implications of this result are in place: (1) a nonstandard CLT of the number of pulls hence pseudo-regret that smoothly interpolates between a standard form in the large arm gap regime and a slow-concentration form in the small arm gap regime, and (2) a heuristic derivation of the sample bias up to its leading order from the correlation between the number of pulls and sample means. Our analysis framework is based on a novel perturbation analysis, which is of broader interest on its own.

强化学习统计推断概率分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。