arXiv:2509.06287math.STcs.AI2025-09被引 4

提出可保证在模型错误时仍收敛的上下文老虎机算法,支持稳定推断。

Statistical Inference for Misspecified Contextual Bandits

  • 设计一类在模型错误下仍能收敛的算法,解决自适应实验中推断失效问题。
  • 提出基于逆概率加权的Z估计器,实现渐近正态性与一致方差估计。
  • 适用于真实场景中模型近似复杂系统时的稳健统计推断,适合实验设计者。

上下文老虎机算法通过实时自适应实现个性化干预和数据高效利用,但其自适应性给统计推断带来挑战。有效推断的关键在于策略收敛性,即给定上下文时动作选择概率趋于稳定。本文指出:广泛使用的算法(如LinUCB)在奖励模型错误设定下可能不收敛,导致推断基础崩塌。这一问题在实践中尤为突出,因线性近似复杂动态系统常被用于平衡偏差与方差。为此,我们提出并分析一类在模型错误下仍保证收敛的算法族。基于该收敛性,构建基于逆概率加权Z估计器(IPW-Z)的一般推断框架,并证明其渐近正态性及一致方差估计。模拟研究表明,该方法提供稳健且数据高效的置信区间,优于仅适用于离线策略评估的现有方法。结果强调,在实际实验中设计具有收敛保障的自适应算法对稳定性和有效推断至关重要。

原文摘要 · Abstract (English)

Contextual bandit algorithms have transformed modern experimentation by enabling real-time adaptation for personalized treatment and efficient use of data. Yet these advantages create challenges for statistical inference due to adaptivity. A fundamental property that supports valid inference is policy convergence, meaning that action-selection probabilities converge in probability given the context. Convergence ensures replicability of adaptive experiments and stability of online algorithms. In this paper, we highlight a previously overlooked issue: widely used algorithms such as LinUCB may fail to converge when the reward model is misspecified, and such non-convergence creates fundamental obstacles for statistical inference. This issue is practically important, as misspecified models -- such as linear approximations of complex dynamic system -- are often employed in real-world adaptive experiments to balance bias and variance. Motivated by this insight, we propose and analyze a broad class of algorithms that are guaranteed to converge even under model misspecification. Building on this guarantee, we develop a general inference framework based on an inverse-probability-weighted Z-estimator (IPW-Z) and establish its asymptotic normality with a consistent variance estimator. Simulation studies confirm that the proposed method provides robust and data-efficient confidence intervals, and can outperform existing approaches that exist only in the special case of offline policy evaluation. Taken together, our results underscore the importance of designing adaptive algorithms with built-in convergence guarantees to enable stable experimentation and valid statistical inference in practice.

上下文老虎机统计推断模型误设

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。