用最小预处理生成真实多表时序电子病历数据
Generating Multi-Table Time Series EHR from Latent Space with Minimal Preprocessing
- 基于文本表示与压缩技术,直接从原始数据生成多表时序病历
- 在两个公开数据集上优于基线模型,保持数据分布与时间动态
- 适合医疗数据合成、隐私保护研究者使用
电子病历(EHR)是记录患者随时间变化的医疗事件的时序关系型数据库,对医疗研究至关重要。但隐私和监管限制阻碍了其共享与利用,亟需生成合成EHR数据。与以往仅生成少数选定特征(如生命体征或结构化编码)的EHR合成方法不同,本文提出首个能生成接近原始多表时序EHR的框架RawMed。通过文本表示与压缩技术,RawMed在最小预处理下捕捉复杂结构与时间动态。我们还提出了针对多表时序合成EHR的新评估框架,涵盖分布相似性、表间关系、时间动态与隐私性。在两个开源EHR数据集上验证,RawMed在保真度与实用性上均优于基线模型。代码已开源:https://github.com/eunbyeol-cho/RawMed。
原文摘要 · Abstract (English)
Electronic Health Records (EHR) are time-series relational databases that record patient interactions and medical events over time, serving as a critical resource for healthcare research and applications. However, privacy concerns and regulatory restrictions limit the sharing and utilization of such sensitive data, necessitating the generation of synthetic EHR datasets. Unlike previous EHR synthesis methods, which typically generate medical records consisting of expert-chosen features (e.g. a few vital signs or structured codes only), we introduce RawMed, the first framework to synthesize multi-table, time-series EHR data that closely resembles raw EHRs. Using text-based representation and compression techniques, RawMed captures complex structures and temporal dynamics with minimal preprocessing. We also propose a new evaluation framework for multi-table time-series synthetic EHRs, assessing distributional similarity, inter-table relationships, temporal dynamics, and privacy. Validated on two open-source EHR datasets, RawMed outperforms baseline models in fidelity and utility. The code is available at https://github.com/eunbyeol-cho/RawMed.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。