融合不完整观察数据与随机实验数据,提升治疗效果差异估计精度。
Combining Incomplete Observational and Randomized Data for Heterogeneous Treatment Effects
- 通过伪实验组与伪对照组构建混淆偏差函数
- 在部分缺失的观察数据下仍可准确估计异质治疗效应
- 适合处理真实世界中常缺样本的医疗研究场景
观察性研究(OS)数据广泛易得,但常含混杂偏倚;随机对照试验(RCT)数据虽能降低偏倚,却因成本高而样本量小。现有融合方法要求观察数据必须完整(包含治疗与未治疗者),限制了实际应用。本文提出CIO方法,可在观察数据不完整时仍有效估计异质治疗效应(HTE)。具体地,利用观测数据中的伪实验组与随机试验中的伪对照组,通过效应估计推导出混淆偏差函数,并将其作为校正残差,修正观察数据的观测结果。最终结合可用的观察数据与全部随机数据进行HTE估计。我们在一个合成数据集和两个半合成数据集上验证了该方法的有效性。
原文摘要 · Abstract (English)
Data from observational studies (OSs) is widely available and readily obtainable yet frequently contains confounding biases. On the other hand, data derived from randomized controlled trials (RCTs) helps to reduce these biases; however, it is expensive to gather, resulting in a tiny size of randomized data. For this reason, effectively fusing observational data and randomized data to better estimate heterogeneous treatment effects (HTEs) has gained increasing attention. However, existing methods for integrating observational data with randomized data must require \textit{complete} observational data, meaning that both treated subjects and untreated subjects must be included in OSs. This prerequisite confines the applicability of such methods to very specific situations, given that including all subjects, whether treated or untreated, in observational studies is not consistently achievable. In our paper, we propose a resilient approach to \textbf{C}ombine \textbf{I}ncomplete \textbf{O}bservational data and randomized data for HTE estimation, which we abbreviate as \textbf{CIO}. The CIO is capable of estimating HTEs efficiently regardless of the completeness of the observational data, be it full or partial. Concretely, a confounding bias function is first derived using the pseudo-experimental group from OSs, in conjunction with the pseudo-control group from RCTs, via an effect estimation procedure. This function is subsequently utilized as a corrective residual to rectify the observed outcomes of observational data during the HTE estimation by combining the available observational data and the all randomized data. To validate our approach, we have conducted experiments on a synthetic dataset and two semi-synthetic datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。