为医学因果推断设计更真实的合成数据生成方法
Improving the Generation and Evaluation of Synthetic Data for Downstream Medical Causal Inference
- 基于治疗效应分析需求,提出三类数据生成标准
- 新方法STEAM在复杂场景下优于现有生成模型
- 适合医疗因果推断研究者与数据隐私敏感场景
因果推断对医疗干预研发至关重要,但真实医疗数据因监管限制难以获取。合成数据可为此类分析提供支持,并促进新推断方法的开发。生成模型能模拟真实数据分布,但现有方法未考虑下游因果推断任务(特别是治疗相关任务)的独特挑战。本文确立了包含治疗效应的合成数据应满足的三大目标:(i) 保持协变量分布,(ii) 保留治疗分配机制,(iii) 模拟结果生成机制。基于此,提出评估合成数据质量的新指标。进一步提出STEAM方法——一种针对医学治疗效应分析的合成数据生成框架,其模仿真实数据生成过程并优化上述目标。实验表明,在复杂真实生成机制下,STEAM在多项指标上达到当前最优性能。
原文摘要 · Abstract (English)
Causal inference is essential for developing and evaluating medical interventions, yet real-world medical datasets are often difficult to access due to regulatory barriers. This makes synthetic data a potentially valuable asset that enables these medical analyses, along with the development of new inference methods themselves. Generative models can produce synthetic data that closely approximate real data distributions, yet existing methods do not consider the unique challenges that downstream causal inference tasks, and specifically those focused on treatments, pose. We establish a set of desiderata that synthetic data containing treatments should satisfy to maximise downstream utility: preservation of (i) the covariate distribution, (ii) the treatment assignment mechanism, and (iii) the outcome generation mechanism. Based on these desiderata, we propose a set of evaluation metrics to assess such synthetic data. Finally, we present STEAM: a novel method for generating Synthetic data for Treatment Effect Analysis in Medicine that mimics the data-generating process of data containing treatments and optimises for our desiderata. We empirically demonstrate that STEAM achieves state-of-the-art performance across our metrics as compared to existing generative models, particularly as the complexity of the true data-generating process increases.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。