提出新方法解决自适应实验中模型错设导致的推断失效问题
Statistical Inference for Misspecified Contextual Bandits
- 基于逆概率加权的Z估计框架,适用于多种目标参数
- 发现标准算法在模型错设时会失稳,导致非正态分布
- 实证验证覆盖可靠,适合需要严谨推断的在线实验场景
上下文多臂老虎机算法通过实时个性化决策推动现代实验发展,但其自适应特性给统计推断带来挑战。本文不假设结果模型正确,研究了在工作模型错设下的推断问题。我们发现,如LinUCB等标准算法在模型错设下可能无法稳定,导致估计量非正态分布,推断无效。这是实际重要问题,因现实中的在线代理常使用复杂系统近似模型以平衡收益、计算效率与鲁棒性。为此,我们提出了针对广义边际矩目标(包括投影参数、带噪声上下文的结构参数、离策略值)的逆概率加权Z估计框架,并识别出关键稳定性条件——缩放逆倾向收敛。在此条件下,IPW-Z估计量具有一致性和渐近正态性,且可一致估计协方差矩阵。我们进一步为多臂老虎机算法和光滑上下文分配策略等政策类给出了满足该条件的充分条件。模拟与HeartSteps V1真实数据校准应用表明,该方法在多个目标上均具有可靠覆盖率和竞争力。总体而言,研究强调了在自适应设计中考虑稳定性对有效事后推断的重要性。
原文摘要 · Abstract (English)
Contextual bandit algorithms have transformed modern experimentation by enabling real-time adaptation for personalized treatment. Yet these advantages create challenges for statistical inference due to adaptivity. We study inference with contextual-bandit data without assuming a well-specified outcome model. In this setting, we show a previously overlooked issue: standard algorithms such as LinUCB may fail to stabilize under misspecified working models, leading to non-Gaussian estimator behavior and invalid inference. This issue is practically important, as misspecified working models -- such as approximations of complex dynamical systems -- are often employed by online agents in real-world adaptive experiments to balance reward, computational tractability, and robustness. We develop an inverse-probability-weighted Z-estimation framework for a broad class of marginal moment targets, including projection parameters, structural parameters with noisy contexts, and off-policy values. We identify a stability condition tailored to this framework, scaled inverse-propensity convergence, under which the IPW-Z estimator is consistent and asymptotically normal with a consistent sandwich variance estimator. We further establish sufficient conditions for scaled inverse-propensity convergence for several policy classes, including multi-armed bandit algorithms and smooth contextual allocation policies. Simulations and a HeartSteps V1 real-data-calibrated application show reliable coverage and competitive performance across multiple targets. Overall, our results highlight the importance of stability-aware adaptive design for valid post-experiment inference.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。