arXiv:2604.23904stat.MEcs.AI2026-04

生成数据常误导因果推断,新方法分离处理与结果建模以提升准确性。

Generative Synthetic Data for Causal Inference: Pitfalls, Remedies, and Opportunities

论文配图:Generative Synthetic Data for Causal Inference: Pitfalls, Remedies, and Opportunities
图 1 · 摘自论文原文
  • 提出混合生成框架,分开建模协变量、处理和结果机制。
  • 在模拟实验中,混合方法比全生成模型更准确保留平均处理效应(ATE)。
  • 适合需要可靠因果推断的医疗、政策评估等实证研究者使用。

合成表格数据通常通过分布相似性、隐私距离或用合成数据训练、真实数据测试的预测性能来评估,但这些标准无法保证因果推断的有效性。我们发现,包括GAN和LLM-based在内的全生成合成器虽能保持预测效用,却会扭曲平均处理效应(ATE)估计。其失败源于结构性缺陷:ATE保真要求协变量分布真实且处理效应对比准确,而预测损失仅通过重叠加权项惩罚处理效应误差。因此,在不平衡或重叠有限的情况下,生成器可能复制主导观测结果,却忽视干预相关对比。我们通过敏感性分析和损失分解形式化这一偏差。基于此因果分析与直觉,提出一种用于因果推断的混合合成框架,生成协变量的同时分别建模处理和结果机制。我们在三种场景下评估该框架:全生成与混合合成对ATE的保真度比较、实际正性问题下的数据增强、以及用于比较OR、IPW、AIPW和TMLE的诊断模拟引擎。还通过改变重叠程度、协变量维度、种子样本量和处理效应复杂度进行压力测试,包括逻辑回归模型误设检查。在受控模拟实验中,混合合成相对全生成基线显著提升因果保真度;ACTG应用显示预测保真度改善,并具备小样本估计器基准化潜力。在可评估因果保真度的场景中,基于LLM的混合合成通常比CTGAN更忠实。

原文摘要 · Abstract (English)

Synthetic tabular data are often evaluated by distributional similarity, privacy distance, or train-on-synthetic-test-on-real predictive performance, but these criteria do not ensure validity for causal inference. We show that fully generative tabular synthesizers, including GAN- and LLM-based models, can preserve predictive utility while distorting average treatment effect (ATE) estimates. The failure is structural: ATE preservation requires both a realistic covariate law and an accurate treatment-effect contrast, whereas prediction loss penalizes treatment-effect error only through an overlap-weighted term. Thus, under imbalance or limited overlap, a generator may reproduce dominant observed outcomes while underlearning intervention-relevant contrasts. We formalize this mismatch through sensitivity and loss-decomposition results. Motivated by this causal analysis and intuition, we propose a hybrid synthetic-data framework for causal inference that generates covariates while modeling treatment and outcome mechanisms separately. We evaluate the framework in three settings: ATE preservation under fully generative versus hybrid synthesis, augmentation for practical positivity problems, and diagnostic simulation engines for comparing OR, IPW, AIPW, and TMLE before real-data analysis. We also stress-test the hybrid construction across settings that vary overlap, covariate dimension, seed sample size, and treatment-effect complexity, including a logistic outcome-model misspecification check. Across controlled simulation experiments, hybrid synthesis improves causal fidelity relative to fully generative baselines; the ACTG application shows improved predictive fidelity and potential for finite-sample estimator benchmarking. LLM-based hybrid synthesis is often more faithful than CTGAN in settings where causal fidelity can be assessed.

因果推断合成数据生成模型混合建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。