arXiv:2508.07914stat.MLcs.IR2025-08被引 1

融合多个离线评估方法,提升推荐系统效果估计的准确性。

Meta Off-Policy Estimation

  • 用元分析框架整合多个离线评估器及其置信区间
  • 在真实和模拟数据上均实现更高统计效率
  • 适合需要可靠评估结果的算法研究者和工程师

离线策略评估(OPE)方法可无偏地评估推荐系统的在线表现,直接从离线数据中估计目标策略的在线奖励,并具备统计保证。尽管已有多种竞争性估计算法,其中双重稳健方法通过结合值函数与策略基估计算法成为主流。本文提出一种新视角:将一组OPE估计算法及其置信区间整合为一个更精确的估计量。该方法采用相关固定效应元分析框架,显式处理因共享数据导致的估计算法间依赖关系,获得目标策略价值的最佳线性无偏估计(BLUE),并提供反映估算器相关性的保守置信区间。在模拟和真实数据上的实验表明,该方法在统计效率上优于现有单一估计算法。

原文摘要 · Abstract (English)

Off-policy estimation (OPE) methods enable unbiased offline evaluation of recommender systems, directly estimating the online reward some target policy would have obtained, from offline data and with statistical guarantees. The theoretical elegance of the framework combined with practical successes have led to a surge of interest, with many competing estimators now available to practitioners and researchers. Among these, Doubly Robust methods provide a prominent strategy to combine value- and policy-based estimators. In this work, we take an alternative perspective to combine a set of OPE estimators and their associated confidence intervals into a single, more accurate estimate. Our approach leverages a correlated fixed-effects meta-analysis framework, explicitly accounting for dependencies among estimators that arise due to shared data. This yields a best linear unbiased estimate (BLUE) of the target policy's value, along with an appropriately conservative confidence interval that reflects inter-estimator correlation. We validate our method on both simulated and real-world data, demonstrating improved statistical efficiency over existing individual estimators.

离线评估推荐系统元分析统计推断

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。