提出新算法,让离线数据评估推荐策略更可靠且有理论保障。
PAC Off-Policy Prediction of Contextual Bandits
- 基于概率近似正确框架,构建条件覆盖的预测区间
- 在有限样本下保证覆盖率,比传统方法更稳健
- 适合安全关键场景,如医疗推荐或金融决策
本文研究上下文随机博弈中的离线策略评估问题,旨在利用不同甚至未知行为策略收集的数据,量化目标策略的性能。近期基于可证实预测的方法已能构建可靠的预测区间,在有限样本下保证边际覆盖率,特别适用于安全敏感场景。为实现对给定离线数据集的条件覆盖率,我们提出一种新型算法,构建概率近似正确(PAC)预测区间。该方法建立在PAC有效的可证实预测框架之上,并通过建立覆盖率的PAC型界来强化理论保证。我们分析了所提方法的有限样本与渐近性质,并在模拟实验中与现有方法进行了实证比较。
原文摘要 · Abstract (English)
This paper investigates off-policy evaluation in contextual bandits, aiming to quantify the performance of a target policy using data collected under a different and potentially unknown behavior policy. Recently, methods based on conformal prediction have been developed to construct reliable prediction intervals that guarantee marginal coverage in finite samples, making them particularly suited for safety-critical applications. To further achieve coverage conditional on a given offline data set, we propose a novel algorithm that constructs probably approximately correct prediction intervals. Our method builds upon a PAC-valid conformal prediction framework, and we strengthen its theoretical guarantees by establishing PAC-type bounds on coverage. We analyze both finite-sample and asymptotic properties of the proposed method, and compare its empirical performance with existing methods in simulations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。