优化日志策略以提升离线评估准确性,平衡奖励覆盖与方差风险。
Logging Policy Design for Off-Policy Evaluation

- 提出统一框架,根据目标策略已知程度设计最优日志策略。
- 揭示奖励-覆盖权衡:高奖励集中降低方差但可能遗漏信号。
- 为推荐系统选型提供可操作指导,适合需安全实验的工业场景。
离线策略评估(OPE)利用不同日志策略收集的数据来估计目标策略(如推荐系统)的价值,实现高风险实验而无需上线部署。然而实际精度高度依赖日志策略的设计。本文研究如何为给定目标策略设计最小化OPE误差的日志策略。我们揭示了一个根本性的奖励-覆盖权衡:将概率集中在高回报动作上可减少方差,但可能遗漏目标策略会采取的动作信号。为此,我们提出了一个统一的日志策略设计框架,并在三种典型信息情形下推导出最优策略:(i) 目标策略和奖励分布已知,(ii) 完全未知,(iii) 通过先验或日志时的噪声估计部分已知。结果为公司选择多个候选推荐系统提供了可行动指南。我们强调了数据收集中处理选择的重要性,并在该目标为主要诉求时给出理论最优方法。此外,还提炼出在实际约束下无法实现理论最优时的实用设计原则。
原文摘要 · Abstract (English)
Off-policy evaluation (OPE) estimates the value of a target treatment policy (e.g., a recommender system) using data collected by a different logging policy. It enables high-stakes experimentation without live deployment, yet in practice accuracy depends heavily on the logging policy used to collect data for computing the estimate. We study how to design logging policies that minimize OPE error for given target policies. We characterize a fundamental reward-coverage tradeoff: concentrating probability mass on high-reward actions reduces variance but risks missing signal on actions the target policy may take. We propose a unifying framework for logging policy design and derive optimal policies in canonical informational regimes where the target policy and reward distribution are (i) known, (ii) unknown, and (iii) partially known through priors or noisy estimates at logging time. Our results provide actionable guidance for firms choosing among multiple candidate recommendation systems. We demonstrate the importance of treatment selection when gathering data for OPE, and describe theoretically optimal approaches when this is a firm's primary objective. We also distill practical design principles for selecting logging policies when operational constraints prevent implementing the theoretical optimum.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。