改进强化学习采样方法,让结果可推断且误差可控
Stable Thompson Sampling: Valid Inference via Variance Inflation
- 用对数因子放大后验方差,使自适应采样数据仍能生成正态估计
- 新方法使参数估计渐近正态,置信区间有效,而后悔值仅多一个对数项
- 适合需要统计推断的强化学习场景,如医疗试验、推荐系统
研究通过汤普森采样类算法收集数据时的统计推断问题。尽管汤普森采样(TS)在渐近最优性和实际效果上表现优异,但其自适应采样机制给参数置信区间的构建带来挑战。本文提出并分析了一种名为稳定汤普森采样(Stable Thompson Sampling)的变体,通过引入对数因子放大后验方差。该修改使得即使在非独立同分布的数据下,各臂均值估计仍趋于渐近正态。关键优势在于:这一统计性能提升仅以对数级别的后悔值增加为代价,相比标准TS可忽略。结果揭示了合理权衡:以极小的后悔代价换取有效的统计推断,为自适应决策算法提供了可信的分析基础。
原文摘要 · Abstract (English)
We consider the problem of statistical inference when the data is collected via a Thompson Sampling-type algorithm. While Thompson Sampling (TS) is known to be both asymptotically optimal and empirically effective, its adaptive sampling scheme poses challenges for constructing confidence intervals for model parameters. We propose and analyze a variant of TS, called Stable Thompson Sampling, in which the posterior variance is inflated by a logarithmic factor. We show that this modification leads to asymptotically normal estimates of the arm means, despite the non-i.i.d. nature of the data. Importantly, this statistical benefit comes at a modest cost: the variance inflation increases regret by only a logarithmic factor compared to standard TS. Our results reveal a principled trade-off: by paying a small price in regret, one can enable valid statistical inference for adaptive decision-making algorithms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。