arXiv:2608.12489cs.LGstat.ME2026-08

评估预算分配策略效果时,哪些方法可靠?研究给出实证基准和避坑指南。

When Can You Trust Offline Evaluation of Equal-Cost Top-k Allocation? A Controlled, Reproducible Benchmark and Practitioner's Guide

  • 通过真实数据对比六种评估方法,发现重叠度取决于日志与目标策略的动作匹配度。
  • 交叉拟合无法消除模型复用偏差,需改变评估目标才能避免。
  • 倾向性估计误差是最大干扰因素,影响尤其严重,需谨慎处理。

组织在预算约束下决定对谁采取行动,希望在部署前评估目标策略的收益。离策略评估从记录数据中提供这一能力,但可部署的策略是确定性的 top-k 策略:它取消了动作上的平均,使弱重叠直接影响估计结果。我们在五个数据集和两个已知效应的实验中基准测试了六种估计器,并以非模拟的成对参考验证其机制。首先,弱重叠由日志器与目标策略的动作一致性决定,而非仅由日志锐度决定:支持范围取决于日志器对目标动作的生成概率。即使使用目标自身得分构建的日志器,其锐化也无法改善重叠;动作层面的不一致会直接导致重叠崩溃。有效样本量能跨日志环境排序风险,但在单一日志内排名能力弱,且其阈值不可迁移。其次,优化器诅咒无法通过交叉拟合结果扰动项解决:当策略在用于评估的数据上拟合时,仅交叉拟合扰动项仍保留复用偏差并使其恶化。诚实的策略级划分通过靶向学习过程的值来规避复用,这改变了评估目标,而非对全样本策略进行去偏。第三,倾向性估计误差是测量到的最大退化:离袋估计对 IPS 的损害超过其他所有扰动,对双重稳健估计影响很小,甚至可能反转重叠诊断本身。日志数据为合成数据,倾向性下限设为 0.02,因此所有失败均发生在有界权重下;该下限也使两种调参混合估计器退化为未调参基线,最终仅剩四种实际不同的估计器,所有精确值表面均为合成或半合成。我们公开该基准,仅提供公共数据。

原文摘要 · Abstract (English)

Organizations decide whom to treat under a budget and want to know what a targeting rule would have earned before deploying it. Off-policy evaluation promises this from logged data, but the deployable rule is a deterministic top-k policy: it removes all averaging over actions, so weak overlap hits the estimate directly. We benchmark six estimators across five datasets and two known-effect sweeps, and validate the mechanisms against a non-simulated paired reference. First, weak overlap is governed by logger-target action alignment, not by logging sharpness alone: what governs support is the logger's probability of the target's actions. Sharpening a logger built from the target's own score barely moves overlap over the tested range; action-level disagreement collapses it. Effective sample size ranks this risk across logging environments, but is weak at ranking candidates within the single log a practitioner holds, and its cut point does not transfer. Second, the optimizer's curse is not fixed by cross-fitting the outcome nuisance. When the rule is fit on the data used to evaluate it, cross-fitting the nuisance alone leaves the reuse bias in place and makes it worse. Honest policy-level splitting avoids the reuse by targeting the learning procedure's value -- a change of estimand, not a de-biasing of the full-sample policy. Third, propensity-estimation error is the largest degradation we measure: an out-of-fold estimate hurts IPS more than any other stress we apply, leaves doubly-robust estimation almost unchanged, and can invert the overlap diagnostic itself. Logging is synthesized and propensities floored at 0.02, so every failure occurs with bounded weights; the floor also reduces the two tuned hybrids to their untuned parents, leaving four practically distinct estimators, and all exact-value surfaces are synthetic or semi-synthetic. We release the benchmark; public data only.

离策略评估预算分配倾向性估计评估可靠性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。