arXiv:2502.11331stat.MEcs.LG2025-02被引 2

用核岭回归实现自适应迁移学习,提升异质治疗效应估计精度

Transfer Learning of CATE with Kernel Ridge Regression

  • 分两阶段训练:先建模再选最优,利用伪结果策略优化
  • 在弱重叠和复杂效应函数下仍保持低均方误差
  • 适合处理源数据与目标人群差异大、无直接观测的场景

数据激增推动了跨研究迁移治疗效应估计的需求,但常受协变量分布差异和处理组与对照组重叠度低的限制。本文提出一种基于核岭回归的条件平均治疗效应(CATE)自适应迁移学习方法。将有标签源数据分为两部分:第一部分用于构建基于回归调整和伪结果的候选模型;第二部分结合无标签目标数据,通过伪结果策略筛选最优模型。理论分析给出紧致的非渐近均方误差界,证明方法对弱重叠和复杂CATE函数的自适应性。大量数值实验表明,该方法在有限样本下具有更优效率和适应性。最终在两个真实数据集上验证了其有效性和优越性。

原文摘要 · Abstract (English)

The proliferation of data has sparked significant interest in leveraging findings from one study to estimate treatment effects in a different target population without direct outcome observations. However, the transfer learning process is frequently hindered by substantial covariate shift and limited overlap between (i) the source and target populations, as well as (ii) the treatment and control groups within the source. We propose a novel method for overlap-adaptive transfer learning of conditional average treatment effect (CATE) using kernel ridge regression (KRR). Our approach involves partitioning the labeled source data into two subsets. The first one is used to train candidate CATE models based on regression adjustment and pseudo-outcomes. An optimal model is then selected using the second subset and unlabeled target data, employing another pseudo-outcome-based strategy. We provide a theoretical justification for our method through sharp non-asymptotic MSE bounds, highlighting its adaptivity to both weak overlaps and the complexity of the CATE function. Extensive numerical studies confirm that our method achieves superior finite-sample efficiency and adaptability. We conclude by demonstrating the effectiveness and superiority of our approach on two real-world datasets.

因果推断迁移学习核方法治疗效应

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。