提出新方法提升上下文老虎机离线评估精度,尤其在行为策略不准确时表现更优。
Kernel weighted importance sampling for off-policy evaluation in contextual bandits

- 融合加权重要性采样与原始重要性采样的优势,引入核函数优化权重
- 在行为策略错误设定下,显著优于传统加权重要性采样等基线方法
- 适合需要高可靠性离线评估的推荐系统、医疗决策等场景
本文提出一种新的离线评估估计器(Kernel-WIS),仅使用离线数据对上下文老虎机问题进行离线策略评估。该方法在行为策略不准确时仍能保持良好性能,实验证明其在多种设置下均优于强基线(包括加权重要性采样)。理论上,Kernel-WIS 具有渐近一致性。其优势源于将加权重要性采样的有界性与原始重要性采样的线性特性相结合,通过核函数实现更稳健的权重调整。
原文摘要 · Abstract (English)
This article presents a novel estimator for performing off-policy evaluation using only offline data for contextual bandits. The proposed estimator, Kernel-WIS is demonstrated to be asymptotically consistent and to empirically outperform strong baselines (including weighted importance sampling), particularly under behaviour policy miss-specification. The benefit of Kernel-WIS is derived from combining the bounded property of weighted importance sampling with the linearity of vanilla importance sampling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。