arXiv:2605.06484stat.MEcs.LG2026-05KDD被引 1

用历史数据校准代理变量推断,提升复杂场景下的统计可靠性。

Estimate Level Adjustment For Inference With Proxies Under Random Distribution Shifts

论文配图:Estimate Level Adjustment For Inference With Proxies Under Random Distribution Shifts
图 1 · 摘自论文原文
  • 将代理与主结果的偏差建模为参数层面的随机效应,基于历史数据估计分布
  • 无需个体响应数据即可校准,支持现有方法叠加使用以应对多重偏差
  • 适用于科学实验、跨时段分析等存在分布偏移的场景,尤其适合小样本历史数据

在众多科学领域中,研究者常依赖代理结果进行快速频繁观测,尤其当主要结果难以直接测量时。尽管代理变量提供更易获取的观测,但其与主结果间存在不完美性。现有统计推断方法通常依赖严格识别假设(如代理性、协变量/标签偏移或缺失性假设),这些假设难以验证且易受多种分布偏移影响,导致参数估计偏差和不确定性量化失准。本文提出一种估计层级框架,受领域自适应启发,通过聚合历史域(如实验、时间段或不同子群体)的观测数据,经验性地校准代理变量推断。该框架将代理-主指标差异建模为参数层面的随机效应,避免对个体级响应数据的依赖。此外,该校正可叠加于现有代理校正方法(如预测驱动推断或重要性加权)之上,以处理未被覆盖的额外偏差。针对历史域数量有限带来的不确定性,我们提供矩估计法和领域自助法。通过公开数据集和真实实验验证了该方法的有效性。

原文摘要 · Abstract (English)

In many scientific domains, including experimentation, researchers rely on measurements of proxy outcomes to achieve faster and more frequent reads, especially when the primary outcome of interest is challenging to measure directly. While proxies offer a more readily accessible observation for inference, the ultimate goal is to draw statistical inferences about the primary outcome parameter and proxy data are typically imperfect in some ways. To correct for these imperfections, current statistical inference methods often depend on strict identifying assumptions (such as surrogacy, covariate/label shift, or missingness assumptions). These assumptions can be difficult to validate and may be violated by various additional sources of distribution shift, potentially leading to biased parameter estimates and miscalibrated uncertainty quantification. We introduce an estimate-level framework, inspired by domain adaptation techniques, to empirically calibrate proxy-based inference. This framework models the proxy-primary metric discrepancy as a random effect at the parameter level, estimating its distribution from aggregated historical observations across past domains (e.g., experiments, time periods, or distinct segments). This method avoids the requirement for retaining individual-level response data. Additionally, this adjustment can be layered on top of existing proxy-correction methods (such as prediction-powered inference or importance weighting) to account for additional biases not addressed by those corrections. To manage uncertainty when the number of historical domains is limited, we provide both a method-of-moments estimator and a domain bootstrap procedure. We further validate this approach using publicly available datasets and real-world experiments.

统计推断代理变量分布偏移领域自适应

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。