揭示索引型强化学习算法如何导致后续推断偏差
Characterizing Bias in Post-Bandit Inference under Index Algorithms
- 提出有效探索率概念,量化算法引发的偏差来源
- 发现非最优臂的标准化偏差衰减速度仅为1/√logT
- 揭示探索性与误差之间的权衡,适合算法设计者参考
Bandit算法生成数据用于后续推断,但自适应采样会引入样本均值偏差。本文分析稳定索引算法(如UCB1及其推广)下的偏差,推导出样本均值偏差和期望Z统计量的精确主导项表达式。分析揭示偏差源于一个依赖索引函数的关键量——有效探索率。例如,在UCB1下,有效探索率为√logT量级,任意非唯一最优臂的标准化偏差以极慢的1/√logT速率衰减。同时表明,索引函数的选择同时影响遗憾和偏差,存在遗憾-偏差权衡:更注重探索的算法可降低偏差但增加遗憾。本文通过新颖的采样动态经验流近似方法,实现了对偏差的精确刻画,该方法本身也具有独立研究价值。
原文摘要 · Abstract (English)
Bandit algorithms generate data for downstream inference, but adaptive sampling biases post-bandit sample means. We analyze this bias for stable index algorithms, including UCB1 and its generalizations, and derive sharp leading-order expressions for the sample-mean bias and expected $Z$-statistic. Our characterization reveals the algorithmic origin of bias through a key index-function-dependent quantity, which we term effective exploration rate. For example, under UCB1, the effective exploration rate is of order $\sqrt{\log T}$, and the standardized bias of any arm (that is not uniquely optimal) decays at the extremely slow rate $1/\sqrt{\log T}$. We also show how the choice of the index function affects both regret and bias, which reveals a regret-bias trade-off: more exploratory algorithm reduces bias but increases regret. Our sharp characterization for bias uses a novel empirical fluid approximation of the algorithm's sampling dynamics, which may be of independent interest.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。