对比七种生成电子病历数据的方法,给出选型建议。
Generating Synthetic Electronic Health Record Data: a Methodological Scoping Review with Benchmarking on Phenotype Data and Open-Source Software
- 系统梳理并测试了七种合成电子病历数据方法
- GAN类方法在数据保真度和实用性上表现最佳
- 提供开源工具包,适合医疗数据研究人员使用
本研究开展了一项关于合成电子健康记录(EHR)数据生成方法的范围综述,并通过开源软件对主流方法进行基准测试,为实践者提供指导。从三个学术数据库中检索相关文献,共识别出42项研究,并将其分为五类。选取涵盖所有类别的七种开源方法,在MIMIC-III数据集上训练,并在MIMIC-III或MIMIC-IV上评估其可迁移性。评估指标包括数据保真度、下游任务效用、隐私保护和计算成本。结果显示,基于GAN的方法在MIMIC-III上保真度和实用性表现优异;基于规则的方法在隐私保护方面更优。在MIMIC-IV上,基于GAN的方法进一步优于基线方法。为此开发了名为SynthEHRella的Python包,集成多种方法与评估指标,支持高效探索与比较。研究发现,方法选择取决于下游应用场景中各评估指标的重要性权重,并提供决策树辅助选择:当存在训练与测试人群分布差异时,推荐使用GAN类方法;否则,CorGAN适用于关联建模,MedGAN适用于预测建模。未来研究应关注提升数据保真度的同时控制隐私风险,并开展纵向或条件生成方法的全面基准测试。
原文摘要 · Abstract (English)
We conduct a scoping review of existing approaches for synthetic EHR data generation, and benchmark major methods with proposed open-source software to offer recommendations for practitioners. We search three academic databases for our scoping review. Methods are benchmarked on open-source EHR datasets, MIMIC-III/IV. Seven existing methods covering major categories and two baseline methods are implemented and compared. Evaluation metrics concern data fidelity, downstream utility, privacy protection, and computational cost. 42 studies are identified and classified into five categories. Seven open-source methods covering all categories are selected, trained on MIMIC-III, and evaluated on MIMIC-III or MIMIC-IV for transportability considerations. Among them, GAN-based methods demonstrate competitive performance in fidelity and utility on MIMIC-III; rule-based methods excel in privacy protection. Similar findings are observed on MIMIC-IV, except that GAN-based methods further outperform the baseline methods in preserving fidelity. A Python package, "SynthEHRella", is provided to integrate various choices of approaches and evaluation metrics, enabling more streamlined exploration and evaluation of multiple methods. We found that method choice is governed by the relative importance of the evaluation metrics in downstream use cases. We provide a decision tree to guide the choice among the benchmarked methods. Based on the decision tree, GAN-based methods excel when distributional shifts exist between the training and testing populations. Otherwise, CorGAN and MedGAN are most suitable for association modeling and predictive modeling, respectively. Future research should prioritize enhancing fidelity of the synthetic data while controlling privacy exposure, and comprehensive benchmarking of longitudinal or conditional generation methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。