arXiv:2601.21131math.STcs.IT2026-01被引 6

揭示汤普森采样中动作选择的动态规律,区分稳定与不稳定情形。

Thompson sampling: Precise arm-pull dynamics and adaptive inference

  • 提出逆过程方法分析次优动作的采样频率,用斯特尔吉斯积分建模。
  • 发现最优动作采样数收敛于随机分布,非稳定时呈现非高斯极限。
  • 为非稳定情况下的置信区间构建提供理论支持,拓展推断应用范围。

自适应采样常导致复杂依赖关系,使传统推断失效。近期研究发现,多臂老虎机中的UCB类算法具有渐近确定性动作选择频次,使推断如同独立同分布场景。本文研究另一类经典汤普森采样算法的精确动作选择动态:我们证明,动作选择频次渐近确定当且仅当该动作为次优或唯一最优;否则其分布收敛至一个随机微分方程(SDE)的唯一不变律。这一二元现象揭示了现有(不)稳定性结果的统一原则:若动作与统计噪声的交互在渐近下可忽略,则其稳定。作为应用,我们证明归一化动作均值也遵循相同二元性:稳定动作服从高斯极限,不稳定动作则趋于半通用非高斯极限。这不仅允许在非正态下构造置信区间,还揭示了超越稳定区间的可处理推断程序潜力。证明基于两种新方法:对次优动作采用‘逆过程’法,通过斯特尔吉斯积分刻画动作选择频次的逆过程;对最优动作,通过重参数化降低自然SDE的奇异性,并利用抛物型霍尔曼条件与斯特朗-瓦拉达汉支撑定理证明另一SDE的不变律唯一性。

原文摘要 · Abstract (English)

Adaptive sampling schemes are well known to create complex dependence that may invalidate conventional inference methods. A recent line of work shows that this need not be the case for UCB-type algorithms in multi-armed bandits. A central emerging theme is a `stability' property with asymptotically deterministic arm-pull counts in these algorithms, making inference as easy as in the i.i.d. setting. In this paper, we study the precise arm-pull dynamics in another canonical class of Thompson-sampling type algorithms. We show that the phenomenology is qualitatively different: the arm-pull count is asymptotically deterministic if and only if the arm is suboptimal or is the unique optimal arm; otherwise it converges in distribution to the unique invariant law of an SDE. This dichotomy uncovers a unifying principle behind many existing (in)stability results: an arm is stable if and only if its interaction with statistical noise is asymptotically negligible. As an application, we show that normalized arm means obey the same dichotomy, with Gaussian limits for stable arms and a semi-universal, non-Gaussian limit for unstable arms. This not only enables the construction of confidence intervals for the unknown mean rewards despite non-normality, but also reveals the potential of developing tractable inference procedures beyond the stable regime. The proofs rely on two new approaches. For suboptimal arms, we develop an `inverse process' approach that characterizes the inverse of the arm-pull count process via a Stieltjes integral. For optimal arms, we adopt a reparametrization of the arm-pull and noise processes that reduces the singularity in the natural SDE to proving the uniqueness of the invariant law of another SDE. We prove the latter by a set of analytic tools, including the parabolic Hörmander condition and the Stroock-Varadhan support theorem.

强化学习贝叶斯推断随机过程动态分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。