用部分参数初始化缓解隐藏混杂下的CATE估计过拟合问题
A Partial Initialization Strategy to Mitigate the Overfitting Problem in CATE Estimation with Hidden Confounding
- 两阶段框架:先用大规模观测数据学基础表示,再融合调整隐藏混杂
- 在两个数据集上验证,相比全量训练减少过拟合,提升估计稳定性
- 适合有小样本随机试验但存在隐藏混杂的场景,如医疗或电商决策
从观测数据中估计条件平均处理效应(CATE)在电商、医疗和经济学等领域至关重要。现有方法多依赖无法检验的不可忽略性假设,而随机对照试验(RCT)虽无混杂但样本量小,易导致过拟合。为此,本文提出一种两阶段预训练-微调(TSPF)框架,结合部分参数初始化策略,在存在隐藏混杂时估计CATE。第一阶段利用大规模观测数据训练协变量的基础表示以预测反事实结果;第二阶段通过拼接基础表示与增强表示来校正隐藏混杂,并将第一阶段的部分预测头用于初始化,避免从头训练。在两个数据集上的大量实验验证了该方法的有效性。
原文摘要 · Abstract (English)
Estimating the conditional average treatment effect (CATE) from observational data plays a crucial role in areas such as e-commerce, healthcare, and economics. Existing studies mainly rely on the strong ignorability assumption that there are no hidden confounders, whose existence cannot be tested from observational data and can invalidate any causal conclusion. In contrast, data collected from randomized controlled trials (RCT) do not suffer from confounding but are usually limited by a small sample size. To avoid overfitting caused by the small-scale RCT data, we propose a novel two-stage pretraining-finetuning (TSPF) framework with a partial parameter initialization strategy to estimate the CATE in the presence of hidden confounding. In the first stage, a foundational representation of covariates is trained to estimate counterfactual outcomes through large-scale observational data. In the second stage, we propose to train an augmented representation of the covariates, which is concatenated with the foundational representation obtained in the first stage to adjust for the hidden confounding. Rather than training a separate network from scratch, part of the prediction heads are initialized from the first stage. The superiority of our approach is validated on two datasets with extensive experiments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。