解决真实世界数据与随机试验数据的特征差异,提升治疗效果估计精度。
Improving RCT-Based Treatment Effect Estimation Under Covariate Mismatch via Calibrated Alignment
- 通过学习统一表示空间,对齐不同来源的特征
- 在51组模拟中显著降低方差,非线性场景下表现最优
- 适合医学研究中融合小规模试验与大规模观察数据
随机对照试验(RCT)是评估治疗效果的金标准,但常因样本量不足难以检测效果异质性。大型观察研究(OS)可补充用于条件平均治疗效应(CATE)估计,但主要障碍是协变量不匹配:两数据源测量的协变量不同且部分重叠。本文提出CALM(Calibrated ALignment under covariate Mismatch),学习将各源特征映射到共同表示空间的嵌入。将OS结果模型迁移至RCT嵌入空间,并用试验数据校准,保留随机化带来的因果识别。有限样本风险界分解为对齐误差、结果模型复杂度和校准复杂度项,明确指出了嵌入足够准确时可降低方差的条件。我们实现两种形式:闭式线性版本CALM-Lin和神经表示学习版本CALM-NN。在51组模拟设置中,基于校准的线性方法在线性-CATE情形下表现相当;而在全部22个非线性-CATE场景中,CALM-NN大幅领先。此外,在两个真实数据研究中,CALM-NN相较仅使用试验数据的基线取得最大提升。
原文摘要 · Abstract (English)
Randomized controlled trials (RCTs) are the gold standard for estimating treatment effects, yet they are often underpowered for detecting effect heterogeneity. Large observational studies (OS) can supplement RCTs for conditional average treatment effect (CATE) estimation, but a key barrier is covariate mismatch: the two sources measure different, only partially overlapping, covariates. We propose CALM (Calibrated ALignment under covariate Mismatch), which learns embeddings that map each source's features into a common representation space. OS outcome models are transferred to the RCT embedding space and calibrated using trial data, preserving causal identification from randomization. Finite-sample risk bounds decompose into alignment error, outcome-model complexity, and calibration complexity terms, making explicit when the learned embedding is accurate enough to reduce variance. We instantiate CALM in two forms: a closed-form linear version, CALM-Lin, and a neural representation-learning version, CALM-NN. Across 51 simulation settings, calibration-based linear methods are effectively tied in linear-CATE regimes, while CALM-NN wins all 22 nonlinear-CATE settings by wide margins. Moreover, on two real-data studies CALM-NN delivers the largest gains over the trial-only baseline.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。