arXiv:2411.13428cs.LGcs.AI2024-11被引 8

用解码器模型生成混合类型电子病历,提升数据质量与隐私保护。

SynEHRgy: Synthesizing Mixed-Type Structured Electronic Health Records using Decoder-Only Transformers

  • 设计专用于多类型病历数据的分词策略,支持变量、编码与时间序列融合
  • 在MIMIC-III数据集上生成高质量合成病历,性能优于现有先进模型
  • 适合医疗数据隐私共享与机器学习训练,尤其关注真实病历结构还原

生成合成电子健康记录(EHR)在数据增强、隐私保护的数据共享以及提升机器学习模型训练方面具有巨大潜力。本文提出一种针对结构化EHR数据的新型分词策略,该数据包含协变量、ICD编码和非规则采样的时间序列等多样数据类型。采用类GPT的解码器仅模型,我们展示了高质合成EHR的生成能力。方法在MIMIC-III数据集上进行评估,并与当前最先进的模型在数据保真度、实用性和隐私性方面进行了基准对比。

原文摘要 · Abstract (English)

Generating synthetic Electronic Health Records (EHRs) offers significant potential for data augmentation, privacy-preserving data sharing, and improving machine learning model training. We propose a novel tokenization strategy tailored for structured EHR data, which encompasses diverse data types such as covariates, ICD codes, and irregularly sampled time series. Using a GPT-like decoder-only transformer model, we demonstrate the generation of high-quality synthetic EHRs. Our approach is evaluated using the MIMIC-III dataset, and we benchmark the fidelity, utility, and privacy of the generated data against state-of-the-art models.

电子病历合成数据Transformer隐私保护

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。