揭露提升评估中指标与目标不匹配问题,验证不同评价指标效果差异。
UpliftBench: Revealing Outcome-Regime and Objective Mismatch in Uplift Evaluation

- 构建多目标、隔离测试的基准框架,统一评估12种提升模型。
- Qini指标与实际效果相关性极弱,而AUUC更贴近真实效果表现。
- 指标选择需匹配政策需求,否则可能导致严重决策偏差。
提升建模(条件平均处理效应估计)推动个性化投放,但现有提升基准常对最优估计器意见不一;我们发现分歧主要源于评价指标而非模型本身。UpliftBench在七个数据集族上,采用外测隔离、多目标协议评估12种提升估计器。当存在参考目标时,两个核心发现被识别:在标准连续基准IHDP上为F1,于Jobs数据集的样本内案例研究中为F2。在IHDP基准上,Qini与效果准确性的关联几乎不存在——100次实现实验中其均值秩相关为+0.07(95%置信区间[-0.03, +0.16]),而AUUC始终更一致(配对前缀-均值-AUUC-超-Qini差值+0.49 [+0.40, +0.59];发布的累积收益AUUC表现更优,达+0.73)。在Jobs上,排名类指标对符号阈值策略无效,因其忽略得分水平;实证显示,直接基于政策风险选择模型的基准遗憾低于随机选型(14-15%),而Qini、AUUC及uplift-at-$k$无改善作用。校准决策阈值可消除81%的Qini选择带来的遗憾。以上结论为受限非普遍:F1未在其他验证集(ACIC和Revenue-Synthetic)中检测到,而F2在预算价值目标下消失,此时排名已足够。UpliftBench发布版本化加载器、固定协议、结果产物及可复现的动态排行榜,公开仓库随论文同步发布。
原文摘要 · Abstract (English)
Uplift modeling (conditional-average-treatment-effect estimation) drives personalized targeting, yet published uplift benchmarks frequently disagree on which estimator performs best; we show the disagreement is substantially about metrics, not models. UpliftBench evaluates 12 uplift estimators under an outer-test-isolated, multi-objective protocol across seven dataset families; its two findings are identified where a reference objective exists -- F1 on the standard continuous benchmark (IHDP), F2 in a within-sample case study on Jobs. On that benchmark, Qini shows no detectable alignment with effect accuracy -- across all 100 IHDP realizations its mean rank correlation with effect accuracy is +0.07, 95% CI [-0.03, +0.16] -- while AUUC is consistently more aligned (paired prefix-mean-AUUC-over-Qini gap +0.49 [+0.40, +0.59]; the shipped cumulative-gain AUUC aligns better still, +0.73). On Jobs, ranking metrics are structurally insufficient for a sign-threshold policy because they discard the score level; empirically, within the released split-rotation analysis direct policy-risk selection yields lower benchmark regret than random model selection while Qini, AUUC, and uplift-at-$k$ do not (14-15% regret). Calibrating the decision threshold removes 81% of the Qini-selection regret. Both findings are bounded, not universal: F1 is not detected on either validation family (the ACIC and Revenue-Synthetic gaps are both indistinguishable from zero), and F2 vanishes under a budgeted-value objective where rank suffices. UpliftBench releases versioned loaders, fixed protocols, result artifacts, and a reproducible living leaderboard; the public repository accompanies the paper.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。