arXiv:2608.01069cs.LGmath.PR2026-08

揭示索引型强化学习算法如何导致后续推断偏差

Characterizing Bias in Post-Bandit Inference under Index Algorithms

  • 提出有效探索率概念,量化算法引发的偏差来源
  • 发现非最优臂的标准化偏差衰减速度仅为1/√logT
  • 揭示探索性与误差之间的权衡,适合算法设计者参考

Bandit算法生成数据用于后续推断,但自适应采样会引入样本均值偏差。本文分析稳定索引算法(如UCB1及其推广)下的偏差,推导出样本均值偏差和期望Z统计量的精确主导项表达式。分析揭示偏差源于一个依赖索引函数的关键量——有效探索率。例如,在UCB1下,有效探索率为√logT量级,任意非唯一最优臂的标准化偏差以极慢的1/√logT速率衰减。同时表明,索引函数的选择同时影响遗憾和偏差,存在遗憾-偏差权衡:更注重探索的算法可降低偏差但增加遗憾。本文通过新颖的采样动态经验流近似方法,实现了对偏差的精确刻画,该方法本身也具有独立研究价值。

原文摘要 · Abstract (English)

Bandit algorithms generate data for downstream inference, but adaptive sampling biases post-bandit sample means. We analyze this bias for stable index algorithms, including UCB1 and its generalizations, and derive sharp leading-order expressions for the sample-mean bias and expected $Z$-statistic. Our characterization reveals the algorithmic origin of bias through a key index-function-dependent quantity, which we term effective exploration rate. For example, under UCB1, the effective exploration rate is of order $\sqrt{\log T}$, and the standardized bias of any arm (that is not uniquely optimal) decays at the extremely slow rate $1/\sqrt{\log T}$. We also show how the choice of the index function affects both regret and bias, which reveals a regret-bias trade-off: more exploratory algorithm reduces bias but increases regret. Our sharp characterization for bias uses a novel empirical fluid approximation of the algorithm's sampling dynamics, which may be of independent interest.

强化学习偏差分析索引算法后悔值

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。