生成更符合临床逻辑的合成病历,提升数据可用性与安全性。
From Statistical Fidelity to Clinical Consistency: Scalable Generation and Auditing of Synthetic Patient Trajectories
- 用知识增强生成模型捕捉3.2万种临床事件,确保结构完整。
- 生成1.8万条病历,审计后临床不一致率下降至45%以下。
- 适合医疗数据共享、隐私保护研究者使用。
电子健康记录(EHR)在数字健康研究中受限于隐私法规与机构壁垒。合成EHR可实现安全的数据共享,但现有方法常仅保留整体统计特征,忽略临床过程的一致性。本文提出集成式流程,通过高保真生成与可扩展审计两步,提升合成患者轨迹的临床一致性。基于MIMIC-IV数据库,训练了涵盖近3.2万种临床事件(含人口学、检验、用药、操作、诊断)的知识驱动生成模型,并强制结构完整性。为实现大规模临床一致性审计,引入大语言模型自动检测并过滤违禁用药等矛盾项。共生成18,071条合成病历,源自180,712名真实患者。合成事件概率与真实数据均值偏差接近0,相关系数达R²=0.99;但三位医生随机抽样评估20条记录后发现45%-60%存在临床不一致。自动化审计使真实与合成数据差异显著缩小(科恩效应量d从0.59–1.60降至0.18–0.67)。下游模型在经审计数据上表现匹配甚至超越真实数据训练结果。未发现隐私泄露风险,成员推断性能等同于随机猜测(F1-score=0.51)。
原文摘要 · Abstract (English)
Access to electronic health records (EHRs) for digital health research is often limited by privacy regulations and institutional barriers. Synthetic EHRs have been proposed as a way to enable safe and sovereign data sharing; however, existing methods may produce records that capture overall statistical properties of real data but present inconsistencies across clinical processes and observations. We developed an integrated pipeline to make synthetic patient trajectories clinically consistent through two synergistic steps: high-fidelity generation and scalable auditing. Using the MIMIC-IV database, we trained a knowledge-grounded generative model that represents nearly 32,000 distinct clinical events, including demographics, laboratory measurements, medications, procedures, and diagnoses, while enforcing structural integrity. To support clinical consistency at scale, we incorporated an automated auditing module leveraging large language models to filter out clinical inconsistencies (e.g., contraindicated medications) that escape probabilistic generation. We generated 18,071 synthetic patient records derived from a source cohort of 180,712 real patients. While synthetic clinical event probabilities demonstrated robust agreement (mean bias effectively 0.00) and high correlation (R2=0.99) with the real counterparts, review of a random sample of synthetic records (N=20) by three clinicians identified inconsistencies in 45-60% of them. Automated auditing reduced the difference between real and synthetic data (Cohen's effect size d between 0.59 and 1.60 before auditing, and between 0.18 and 0.67 after auditing). Downstream models trained on audited data matched or even exceeded real-data performance. We found no evidence of privacy risks, with membership inference performance indistinguishable from random guessing (F1-score=0.51).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。