arXiv:2606.08903cs.LG2026-06

现有合成医疗数据评估方法易误判质量,本文提出多维度临床有效性评估框架。

Synthetic but Not Realistic: The Evaluation Challenge in Generative Modelling for Structured Electronic Medical Records

论文配图:Synthetic but Not Realistic: The Evaluation Challenge in Generative Modelling for Structured Electronic Medical Records
图 1 · 摘自论文原文
  • 从流行病学出发,构建描述、预测、因果三重评估维度
  • 在5万患者真实数据集上验证,所有模型均未同时保留亚组结构与因果关系
  • 揭示分布拟合好但临床推断不可靠的隐患,适合医学AI研究者参考

合成医疗数据被广泛视为保护隐私的替代方案,但当前评估仍依赖统计相似性和预测性能,无法反映临床有效性。本文提出基于流行病学的多维度评估框架,涵盖描述性保真度、临床实用性与结构有效性,分别对应描述、预测和因果问题。在包含5万患者的PRIME-CVD真实队列上,评估了四种代表性生成范式(GAN、VAE增强、扩散模型、掩码建模)。尽管所有模型均能复现边际分布,但无一能同时保持亚组结构、效应估计值与依赖关系。尤其值得注意的是,具备强分布保真度的模型可能表现不佳校准且关系扭曲,导致推断不可靠。结果表明,现有评估可能高估合成数据质量,强调需基于支持有效临床与科学结论的能力进行领域导向评估。

原文摘要 · Abstract (English)

Synthetic healthcare data are widely proposed as privacy-preserving substitutes for real patient data, yet their evaluation remains dominated by statistical similarity and predictive performance that do not reflect clinical validity. We introduce a multi-dimensional evaluation framework grounded in epidemiology, assessing descriptive fidelity, clinical utility, and structural validity, corresponding to descriptive, predictive, and causal questions. We evaluate four representative generative paradigms - GAN-based, VAE-boosted, diffusion-based, and masked modelling - using PRIME-CVD, a 50,000-person cohort with known ground-truth structure. While all models reproduce marginal distributions, none simultaneously preserve subgroup structure, effect estimates, and dependency structure. Notably, models with strong distributional fidelity can exhibit poor calibration and distorted relationships, leading to unreliable inference. These results show that current evaluation practices can overestimate synthetic data quality and motivate domain-informed assessment based on the ability to support valid clinical and scientific conclusions.

医疗数据生成模型评估框架合成数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。