构建多源医学数据合成基准,评估生成数据真实性和下游任务效果。
CoMedBench: A Multi-Source Benchmark of Synthetic Medical Data Fidelity and Downstream Utility

- 统一框架下对比多种生成器在20个静态表型与17个时间序列任务上的表现
- 合成数据在分类任务中保留90.6%~97.3%的模型性能(AUROC),时间序列任务仍达81.6%~95%
- 覆盖7个公开医疗数据集,适合医疗AI开发者验证合成数据可靠性
临床数据对开发可靠医疗机器学习系统至关重要,但电子病历因隐私法规、机构审查及重识别风险难以直接使用。合成数据可保留统计与临床结构的同时降低敏感信息暴露。现有研究多仅评估单一生成器、数据集或窄范围下游任务,难以判断合成数据何时有效。我们提出CoMedBench,一个可复现的多源基准,统一评估多个生成器在真实临床有效性框架下的表现,并使用共享训练与评估引擎,覆盖静态表型与时序任务。该基准包含37个数据-任务组合,涵盖20个静态表型与17个重症监护时序任务,源自七个公开数据源:MIMIC-III、MIMIC-IV、eICU、UCI机器学习库、CDC BRFSS糖尿病队列(2015)、NHANES(1999–2014)以及pycox生存数据集(GBSG与METABRIC)。通过比较真实与合成数据上训练模型的性能,评估统计保真度与任务效用。结果表明,合成数据在多数任务中保留了主要下游信号:表型任务中参考生成器CoMed-CTGAN的平均AUROC效用为90.6%,最强生成器CoMed-TVAE达97.3%;时序任务更难且对生成器敏感,CoMed-CTGAN在平衡性敏感指标AUPRC下仅保留64.0%,而CoMed-TVAE仍维持约95%(AUROC)。
原文摘要 · Abstract (English)
Access to clinical data is essential for developing reliable healthcare machine learning systems, but direct use of electronic health records is constrained by privacy regulation, institutional review, data-use agreements, and the risk of re-identification. Synthetic data promises a practical alternative: it can preserve useful statistical and clinical structure while reducing exposure of sensitive patient records. Prior studies often evaluate a single generator, one dataset, or a narrow downstream task, making it difficult to know when synthetic data can support model development and when it fails to preserve task-critical signal. We introduce CoMedBench, a reproducible benchmark that evaluates a family of generators under a common clinical-validity framework and one shared training and evaluation engine, spanning static tabular and temporal downstream tasks on established critical-care datasets. In total the benchmark spans 37 dataset-task pairs across two modalities consists of 20 static tabular and 17 temporal ICU time-series-drawn from seven public data sources: three intensive-care databases (MIMIC-III, MIMIC-IV, and eICU) together with the UCI Machine Learning Repository, the CDC BRFSS diabetes cohort (2015), NHANES (1999-2014), and the pycox survival datasets (GBSG and METABRIC). The benchmark evaluates both statistical fidelity and task utility by comparing models trained and tested across real and synthetic data. In these settings, synthetic training data preserves most of the downstream signal: on tabular tasks the reference generator CoMed-CTGAN retains a mean AUROC utility (the synthetic-to-real performance ratio) of 90.6%, rising to 97.3% for the strongest generator, CoMed-TVAE. Temporal ICU tasks are harder and more generator-sensitive: CoMed-CTGAN retains 81.6% (AUROC) and only 64.0% under the imbalance-sensitive AUPRC, whereas CoMed-TVAE still retains ~95% (AUROC).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。