arXiv:2609.06294cs.LG2026-09

用历史多任务数据学低维表示,提升小样本下因果效应估计效率

Representation Learning for Sample-Efficient CATE Estimation by Leveraging Multiple Outcomes

论文配图:Representation Learning for Sample-Efficient CATE Estimation by Leveraging Multiple Outcomes
图 1 · 摘自论文原文
  • 利用历史多结局数据学习协变量的低维表示,缓解高维困境
  • 理论证明该表示可保留处理异质性,使小样本下估计更准确
  • 适合医疗、政策等实验样本少但历史数据丰富的领域

估计条件平均处理效应(CATE)可实现干预的精准投放,但许多应用面临实验样本有限的问题,难以从高维协变量中估计异质效应。此时,政策制定者和医疗从业者常陷入维度诅咒,或使用不保处理异质性的通用降维方法。然而这些领域通常拥有大量测量多种结果的历史数据——这一监督信号在实践中极少被利用。基于因果表示学习,我们假设高维协变量背后存在低维动态。可通过历史数据中的多样结果学习协变量的低维表示。理论上,当辅助结果满足一组代理条件且表示保留相关协变量信息时,原始CATE在用学习表示替代高维协变量后仍可识别。结合现有CATE估计的维度依赖率,结果表明相同实验样本下具有更高样本效率。此外,我们刻画了假设不完全成立时的偏差-方差权衡,发现当估计器方差下降超过压缩带来的偏差时,表示基估计器仍能实现更低误差。我们在合成数据和半合成医学数据上进行了实证评估。

原文摘要 · Abstract (English)

Estimating conditional average treatment effects (CATE) enables efficient targeting of interventions, but many applications have limited experimental samples, making it difficult to estimate heterogeneous effects from high-dimensional covariates. In such settings, policymakers and medical practitioners often succumb to the curse of dimensionality or apply off-the-shelf dimension reduction methods that may not preserve treatment heterogeneity. Yet these domains often come with large historical datasets measuring a wide range of outcomes -- a source of supervision that is rarely exploited in practice. Following causal representation learning, we hypothesize that such domains with high-dimensional covariates have lower-dimensional underlying dynamics. We can thus leverage the diverse outcomes measured in historical data to learn a lower-dimensional representation of the covariates. Theoretically, we prove that when the auxiliary outcomes satisfy a set of surrogacy conditions and the representation retains relevant covariate information, the original CATE is identified when the high-dimensional covariates are replaced by the learned representation. Combined with existing dimension-dependent rates for CATE estimation, the result implies greater sample-efficiency on the same experimental sample. Additionally, we characterize the bias-variance tradeoff when the assumptions do not hold perfectly, and show that the representation-based estimator can still achieve lower error when the reduction in estimator variance outweighs the bias due to compression. Empirically, we evaluate the method on synthetic data and semi-synthetic medical data.

因果推断表示学习小样本医疗分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。