提出新算法实现上下文老虎机中可解释的高效学习与推断
Kernel Single-Index Bandits: Estimation, Inference, and Learning
- 用核方法结合斯坦估计和逆倾向加权,兼顾灵活性与可解释性
- 证明了自适应采样下估计量渐近正态,给出有效置信区间
- 适用于需同时学习与统计推断的场景,如个性化推荐系统
我们研究具有有限动作的上下文老虎机问题,其中每个臂的回报遵循带有臂特定指数参数和未知非参数链接函数的单指数模型。考虑臂对应稳定决策选项、协变量在带子策略下自适应演化的设定,该设置带来显著统计挑战:采样分布依赖于分配规则,观测值随时间相关,逆倾向加权导致方差膨胀。我们提出一种基于核的ε-贪婪算法,结合斯坦估计器估计指数参数与逆倾向加权核岭回归估计回报函数。该方法实现灵活的半参数学习并保持可解释性。分析发展了处理自适应数据的新工具:在自适应采样下建立了单指数估计量的渐近正态性,获得有效置信区域;推导出再生核希尔伯特空间(RKHS)估计量的方向性功能中心极限定理,提供渐近有效的逐点置信区间。分析依赖于逆加权格拉姆矩阵的集中界与鞅中心极限定理。进一步获得有限时间悔悟保证,包括在常见链接利普希茨条件下达到 ilde{O}(√{T})的速率,表明半参数结构可在不牺牲统计效率的前提下被利用。这些结果为单指数上下文老虎机中的联合学习与推断提供了统一框架。
原文摘要 · Abstract (English)
We study contextual bandits with finitely many actions in which the reward of each arm follows a single-index model with an arm-specific index parameter and an unknown nonparametric link function. We consider a regime in which arms correspond to stable decision options and covariates evolve adaptively under the bandit policy. This setting creates significant statistical challenges: the sampling distribution depends on the allocation rule, observations are dependent over time, and inverse-propensity weighting induces variance inflation. We propose a kernelized $\varepsilon$-greedy algorithm that combines Stein-based estimation of the index parameters with inverse-propensity-weighted kernel ridge regression for the reward functions. This approach enables flexible semiparametric learning while retaining interpretability. Our analysis develops new tools for inference with adaptively collected data. We establish asymptotic normality for the single-index estimator under adaptive sampling, yielding valid confidence regions, and derive a directional functional central limit theorem for the RKHS estimator, which provides asymptotically valid pointwise confidence intervals. The analysis relies on concentration bounds for inverse-weighted Gram matrices together with martingale central limit theorems. We further obtain finite-time regret guarantees, including $\tilde{O}(\sqrt{T})$ rates under common-link Lipschitz conditions, showing that semiparametric structure can be exploited without sacrificing statistical efficiency. These results provide a unified framework for simultaneous learning and inference in single-index contextual bandits.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。