综述生成医学文本、时间序列和纵向数据的合成模型,聚焦隐私保护与评估挑战。
A Review on Generative AI Models for Synthetic Medical Text, Time Series, and Longitudinal Data
- 按数据类型分类梳理52篇文献,分析生成方法与研究目标
- 对抗网络、概率模型、大语言模型分别在三类数据中表现最优
- 指出当前缺乏可靠评估指标,适合医疗数据隐私研究者参考
本文首次开展针对合成健康记录(SHRs)生成的系统性综述,涵盖医学文本、时间序列和纵向数据三类。共纳入52篇符合条件的研究,其中时间序列22篇,纵向数据17篇,医学文本13篇。研究发现,隐私保护是主要目标,其次为类别不平衡、数据稀缺与数据填补。生成性能方面,基于对抗网络的模型在纵向数据上表现更优,概率模型适用于时间序列,而大语言模型在医学文本生成中占优。当前最大研究空白在于缺乏可靠的评估指标来量化合成数据的重识别风险。
原文摘要 · Abstract (English)
This paper presents the results of a novel scoping review on the practical models for generating three different types of synthetic health records (SHRs): medical text, time series, and longitudinal data. The innovative aspects of the review, which incorporate study objectives, data modality, and research methodology of the reviewed studies, uncover the importance and the scope of the topic for the digital medicine context. In total, 52 publications met the eligibility criteria for generating medical time series (22), longitudinal data (17), and medical text (13). Privacy preservation was found to be the main research objective of the studied papers, along with class imbalance, data scarcity, and data imputation as the other objectives. The adversarial network-based, probabilistic, and large language models exhibited superiority for generating synthetic longitudinal data, time series, and medical texts, respectively. Finding a reliable performance measure to quantify SHR re-identification risk is the major research gap of the topic.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。