arXiv:2607.24401stat.MLcs.LG2026-07

解决代理指标估算偏差问题,提升实验推断可靠性。

proxymate: Diagnosis and Adjustment of Proxy Estimates for Reliable Inference

论文配图:proxymate: Diagnosis and Adjustment of Proxy Estimates for Reliable Inference
图 1 · 摘自论文原文
  • 分四层诊断代理指标有效性:人群、测量、决策、跨域
  • 通过诊断与修正策略,纠正数百万次代理-主指标对比偏差
  • 适用于实验评估、流行病监测等需快速决策的场景

代理结果(如短期行为信号、模型预测或替代终点)常被用作难以直接测量、成熟慢或罕见的主结果的替代。但代理结果的有效性并不保证主结果推断的有效性,因为代理估计可能产生难以预测的系统性偏差,导致置信区间校准错误。本文提出 proxymate,一个用于代理验证与调整的框架及开源 Python 工具包。该框架分为四个层级:代表性层级(人群有效性)、单元层级(测量质量)、估计层级(决策有效性)和领域层级(跨领域可迁移性)。每一层级提供诊断检查与针对性调整策略,将具体失效模式映射至相应修正方法。在 Meta 内部,proxymate 已被多个应用场景采用,涵盖实验设计、流行率估计与监控等,应对不同挑战(如人工审核时间有限、结果成熟周期长、检测率低),展现出高度模块化。该工具已评估并修正了数百万次代理与主指标单位的对比,支持数千个实验的快速决策,推动多个业务线落地。

原文摘要 · Abstract (English)

Proxy outcomes (such as short-term behavioral signals, model predictions, or surrogate endpoints) are frequently used in place of primary outcomes that are too slow to mature, rare, or challenging to measure directly. But valid inference on a proxy does not guarantee valid inference on the primary estimate as proxy-based estimates can be systematically biased in ways that are difficult to predict, leading to improperly calibrated confidence intervals. We present proxymate, a framework and open-source Python package for proxy validation and adjustment. proxymate organizes into four levels: The Representativity Level (population validity), the Unit Level (measurement quality), the Estimate Level (decision validity), and the Domain Level (cross-domain transportability). Within each level, proxymate provides diagnostic checks, and targeted adjustment strategies that map specific failures to appropriate corrections. At Meta, proxymate has been adopted by many different use cases, spanning experimentation, prevalence estimation, and monitoring use cases, all facing different proxy challenges (limited human review time, long maturation window of outcomes, low detectability) and showcasing the modularity of the framework. Across all products, proxymate assessed and corrected millions of proxy, primary unit comparisons. It has facilitated launches across multiple work streams including enabling quick decision making on thousands of experiments.

代理指标实验评估偏差修正可信赖推断

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。