arXiv:2606.06990cs.LG2026-06

统一框架让合成病历生成模型可复现、可比对。

Accelerating Reproducible Research in Synthetic EHR Generation

论文配图:Accelerating Reproducible Research in Synthetic EHR Generation
图 1 · 摘自论文原文
  • 构建端到端统一流水线,整合数据、训练与评估
  • 在完整ICD-9词表下重实现多个强基线模型并加入GPT-2
  • 提供无架构依赖的隐私-效用评估,支持可复现对比

高保真合成电子健康记录(EHR)对推动医学研究并保护患者隐私至关重要。然而,现有生成模型的横向比较受限于代码库分散、数据加载不兼容、依赖冲突及评估协议不一致。为此,我们提出一个轻量级、端到端的可复现合成EHR评估框架,形成从数据摄入、标准化训练到无架构依赖评估的统一流程。当前实现聚焦于纵向ICD诊断码生成——该领域最常研究的模态,基于社区维护的PyHealth库构建。我们在全ICD-9词汇粒度下重实现了多个强基线模型(MedGAN、CorGAN、PromptEHR、HALO),并引入来自通用序列建模领域的轻量GPT-2基线。贡献了一个严格的、无架构依赖的隐私-效用评估套件,适用于GAN与Transformer类生成器,并报告所有指标的自助法置信区间。进一步分析了现有模型在长尾分布上的表现缺陷,并讨论了框架向诊断码以外扩展的潜力。通过降低运行、扩展与评估的工程门槛,本工作为社区驱动的可复现性与基准测试提供了起点。

原文摘要 · Abstract (English)

The generation of high-fidelity synthetic Electronic Health Records (EHR) is crucial for advancing medical research while preserving patient privacy. However, head-to-head comparison of existing generative models is hindered by disjointed codebases, incompatible data loaders, conflicting library dependencies, and inconsistent evaluation protocols. To address these gaps, we introduce a lightweight, end-to-end benchmarking framework for reproducible synthetic EHR evaluation, organized as a unified pipeline spanning data ingestion, standardized model training, and architecture-agnostic evaluation. Our current implementation targets the generation of longitudinal ICD diagnosis codes -- the most commonly studied modality in this literature -- and is built on the community-maintained PyHealth library. We reimplement and unify strong baselines (MedGAN, CorGAN, PromptEHR, HALO) under full ICD-9 vocabulary granularity, and add a lightweight GPT-2 baseline from the general-purpose sequence-modeling literature. We contribute a rigorous, architecture-agnostic privacy-utility evaluation suite that applies identically to GAN- and transformer-based generators, and report bootstrapped confidence intervals across all metrics. We further analyze the poor long-tailed performance of existing models and discuss the extensibility of our framework beyond diagnosis codes. By lowering the engineering barrier to running, extending, and evaluating under a single pipeline, we introduce a starting point for community-driven reproducibility and benchmarking synthetic EHR models.

合成数据医疗生成可复现性EHR

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。