arXiv:2606.00913stat.MLcs.LG2026-06

为自适应采样数据构建置信区间,解决带通货算法的统计推断难题。

Bandit Simulation for Average Reward Inference

论文配图:Bandit Simulation for Average Reward Inference
图 1 · 摘自论文原文
  • 用观测数据拟合环境模拟器,评估任意策略的平均奖励
  • 在标准方法失效时仍保持名义覆盖率,置信区间有效
  • 无需重要性加权,弱探索假设下即可应用,适合真实场景

多臂赌博机算法广泛应用于在线平台、临床试验和社会科学实验,但对其性能进行有效的统计推断仍是开放挑战。部署后,自然问题是能否构建其平均奖励的置信区间,并评估是否显著优于基准策略。单次部署的总奖励是随机的,同一群体重复部署通常产生不同的奖励轨迹,因奖励具有随机性。标准统计推断方法不适用,因赌博机算法引入复杂依赖关系,违反经典方法所需的独立同分布假设。现有自适应数据收集的推断方法仅适用于不依赖数据采集算法的估计量(如固定动作下的平均奖励)。本文提出带通货模拟推断(BSI)框架:从观测数据(无论在线或离线)拟合赌博机环境的模拟器,进而估计任意评估策略下的平均奖励,包括自适应黑盒算法。BSI将模拟器参数估计的不确定性正式传播到置信区间构造中。此外,为保证有效性,只需对行为策略施加弱探索假设,避免使用重要性加权。我们证明了BSI可生成渐近有效的置信区间,并通过实验证明,在标准离线策略评估方法失效的情况下,其仍能保持名义覆盖度。

原文摘要 · Abstract (English)

Multi-arm bandit algorithms are increasingly used in online platforms, clinical trials, and social science experiments, but valid statistical inference on their performance remains an open challenge. After deploying bandits, a natural question is whether one can construct a confidence interval for its mean reward and assess whether it reliably outperforms a baseline policy. The total reward achieved in any single bandit deployment is random, and deploying a bandit twice on the same population typically yields different reward trajectories due to stochastic rewards. Standard statistical inference methods cannot be used because bandit algorithms introduce complex dependencies in the collected data, which violate the i.i.d. assumption underlying many classical approaches. Moreover, existing inference methods for adaptively collected data only apply to estimands that do not depend on the data-collection algorithm (such as the mean reward under a fixed action). We propose Bandit Simulation for Inference (BSI), a framework that fits a simulator of the bandit environment from observed data--either on-policy or off-policy--and uses it to estimate the mean reward under any evaluation policy, including adaptive blackbox algorithms. BSI formally propagates uncertainty in the estimated simulator parameters into the confidence interval construction. Furthermore, for BSI to be valid, it requires only weak exploration assumptions on the behavior policy and avoids importance weighting. We prove that BSI yields asymptotically valid confidence intervals, and demonstrate empirically that it maintains nominal coverage in settings where standard off-policy evaluation methods fail.

统计推断多臂赌博机置信区间强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。