乐观性让贝叶斯老虎机推断更稳定,实现可靠统计推断。
Optimism Stabilizes Thompson Sampling for Adaptive Inference
- 引入乐观机制,使各臂抽样次数集中于确定尺度。
- 在任意多最优臂情况下,次优臂抽样数渐近对数增长。
- 适合关注自适应采样下统计推断有效性的研究者。
Thompson采样(TS)广泛用于随机多臂老虎机问题,但其在自适应数据收集下的推断性质复杂。经典大样本理论可能失效,因为每臂样本量是随机的,且与奖励通过动作选择规则耦合。本文研究了具有独立子高斯奖励噪声的K臂老虎机中基于高斯随机索引的TS的自适应推断,识别出‘乐观性’是恢复‘稳定性’的关键机制,即每臂抽样次数集中在确定尺度上。该稳定性确保即使在自适应采样下仍能进行渐近有效的Wald推断。首先,我们证明方差放大版的TS对任意K≥2均稳定,包括多个最优臂的挑战情形,实现了最优臂间渐近均匀分配,次优臂抽样数满足精确的对数渐近特性。这解决了Halder等人提出的K臂扩展问题,采用新的胜者映射和李雅普诺夫漂移技术控制多最优臂间的分配。其次,分析了一种保持高斯索引方差不变但向索引中心添加显式均值奖励的乐观修改,也得到类似稳定性结论。总之,合理实施的乐观性可稳定TS,并在多臂老虎机中实现渐近有效的Wald推断,仅带来轻微额外遗憾代价。
原文摘要 · Abstract (English)
Thompson sampling (TS) is widely used for stochastic multi-armed bandits, yet its inferential properties under adaptive data collection are subtle. Classical asymptotic theory for sample means can fail because arm-specific sample sizes are random and coupled with the rewards through the action-selection rule. We study adaptive inference for Thompson sampling with Gaussian randomized indices in $K$-armed stochastic bandits with independent sub-Gaussian reward noises, and identify \emph{optimism} as a key mechanism for restoring \emph{stability}, meaning that each arm's pull count concentrates around a deterministic scale. This stability yields asymptotically valid Wald inference despite adaptive sampling. First, we prove that variance-inflated TS is stable for any $K \ge 2$, including the challenging regime where multiple arms are optimal, with asymptotically uniform allocation over optimal arms and sharp logarithmic pull-count asymptotics for suboptimal arms. This resolves the $K$-armed extension question raised by \citet{halder2025stable}, using new winner-map and Lyapunov-drift techniques to control allocation among multiple optimal arms. Second, we analyze an alternative optimistic modification that keeps the Gaussian index variance unchanged but adds an explicit mean bonus to the index center, and establish a similar stability conclusion. In summary, suitably implemented optimism stabilizes Thompson sampling and enables asymptotically valid Wald inference in multi-armed bandits, while incurring only a mild additional regret cost.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。