arXiv:2605.12051cs.LG2026-05

从数据中学习可直接替换主结果的代理指标,提升实验效率。

Learning plug-in surrogate endpoints for randomized experiments

论文配图:Learning plug-in surrogate endpoints for randomized experiments
图 1 · 摘自论文原文
  • 通过建模治疗后变量构造可直接使用的代理指标
  • 在真实与模拟实验中均显著提升代理效果预测力
  • 适合需要快速评估治疗效果的研究者使用

在随机实验中,当观测长期结果成本过高或不切实际时,常使用短期代理终点替代。一个好的代理终点应使基于该代理的实验结果能预测真实结局实验的结果。尽管已有大量研究从因果角度形式化这一性质,但多数标准不可识别,难以转化为可操作的数据学习算法。为此,本文研究了插件型复合代理终点——即基于治疗后变量的函数,可直接替代主结果用于随机实验。我们提出了两种最大化效应可预测性的学习方法,并在代表性场景下刻画了获得无偏效应估计的可能性。在具有已知效应的合成实验以及真实世界实验数据中,基于直接建模代理效应的方法所得到的插件代理终点,比现有方法更准确地预测主效应。

原文摘要 · Abstract (English)

Surrogate endpoints are used in place of long-term outcomes in randomized experiments when observing the real outcome for a large enough cohort is prohibitively expensive or impractical. A short-term surrogate is good if the result of an experiment using the surrogate is predictive of the result of a hypothetical study using the real outcome. Much attention has been paid to formalizing this property in causal terms, but most criteria are unidentifiable and cannot be turned into practical algorithms for learning surrogate endpoints from data. To address this, we study plug-in composite surrogates, functions of post-treatment variables that may be substituted directly for the primary outcome in a randomized experiment. We propose two methods for learning plug-in surrogates that maximize effect predictiveness, and characterize the possibility of finding endpoints that yield unbiased effect estimates in representative scenarios. Finally, in both synthetic experiments with known effects and in data from a real-world experiment, we find that our method, based on directly modeling the surrogate effect, returns plug-in endpoints more predictive of the primary effect than established methods.

代理终点因果推断实验设计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。