为线上实验中的代理指标提供可靠性评分,避免因用户分群差异导致误判。
PROXIMA: A Reliability Scoring Framework for Proxy Metrics in Online Controlled Experiments

- 通过相关性、方向准确性和分段脆弱性三维度评估代理指标可靠性。
- 在广告和推荐场景中,早期互动指标决策一致率达98.4%。
- 揭示推荐领域分段异质性高达68%,远高于广告的13%。
大规模在线A/B测试依赖代理指标——即短期易获取信号,替代反应缓慢的长期结果。当代理与结果的关系在不同用户群体中存在异质性时,整体相关性可能掩盖方向性错误,类似辛普森悖论,导致代价高昂的发布/不发布误判。我们提出PROXIMA(在线实验代理指标验证框架),一个轻量级诊断工具,通过归一化效应相关性、方向准确性和分段脆弱性三个互补维度对代理指标的可靠性进行评分。不同于预测长期处理效应的替代指标方法,PROXIMA直接检验候选代理是否导致正确发布决策,并标出失效的用户群体。我们在两个公开数据集上验证:Criteo Uplift语料库(1400万样本,广告)和KuaiRec(7000用户,视频推荐),共使用80次模拟A/B测试。早期参与度指标在Criteo上复合可靠性得分为0.80,在KuaiRec上为0.62,与理想策略的平均决策一致率达98.4%。脆弱性分析显示,推荐领域分段异质性(68%脆弱性)显著高于广告领域(13%),但两者方向准确性均超过96%。权重空间敏感性分析表明,单一维度不足,复合评分比相关性单独评估更能有效区分可靠与不可靠代理。代码与复现脚本见:https://github.com/Avinash-Amudala/PROXIMA。
原文摘要 · Abstract (English)
Online A/B testing at scale relies on proxy metrics -- short-term, easily-measured signals used in place of slow-moving long-term outcomes. When the proxy-outcome relationship is heterogeneous across user segments, aggregate correlation can mask directional failures akin to Simpson's Paradox, leading to costly ship/no-ship errors. We introduce PROXIMA (Proxy Metric Validation Framework for Online Experiments), a lightweight diagnostic framework that scores proxy reliability through a composite of three complementary dimensions: normalised effect correlation, directional accuracy, and segment-level fragility rate. Unlike surrogate-index approaches that predict long-term treatment effects, PROXIMA directly audits whether a candidate proxy leads to correct launch decisions and flags the user segments where it fails. We validate PROXIMA on two public datasets -- the Criteo Uplift corpus (14M observations, advertising) and KuaiRec (7K users, video recommendation) -- using 80 simulated A/B tests. Early engagement metrics achieve a composite reliability of 0.80 on Criteo and 0.62 on KuaiRec, yielding 98.4% average decision agreement with an oracle policy. Fragility analysis reveals that recommendation domains exhibit substantially higher segment-level heterogeneity (68% fragility) than advertising (13%), yet directional accuracy remains above 96% in both cases. A sensitivity analysis over the weight space confirms that no single component suffices and that the composite provides substantially better discrimination between reliable and unreliable proxies than correlation alone. Code and reproduction scripts are available at: https://github.com/Avinash-Amudala/PROXIMA
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。