用层次与语义引导生成更真实的病历数据,提升医疗模型可用性。
Generating Clinically Realistic EHR Data via a Hierarchy- and Semantics-Guided Transformer
- 融合临床编码的层级关系与描述语义,构建双信息驱动生成框架。
- 在MIMIC-III/IV数据集上,合成数据统计特性更接近真实病历。
- 适合需要高保真医疗数据的研究者和隐私保护场景使用。
生成真实合成电子健康记录(EHR)在加速医疗研究、推动人工智能模型发展及保障患者隐私方面具有巨大潜力。然而,现有生成方法通常将EHR视为离散医学编码的扁平序列,忽略了临床编码系统的固有层级结构和代码描述提供的丰富语义信息。因此,合成患者序列往往缺乏高临床保真度,在下游临床任务中实用性有限。本文提出层次与语义引导的Transformer(HiSGT),通过构建编码间的层次图并利用图神经网络提取层次感知嵌入,再与预训练临床语言模型(如ClinicalBERT)提取的语义嵌入融合,使基于Transformer的生成器能更准确建模真实EHR中的复杂临床模式。在MIMIC-III和MIMIC-IV数据集上的大量实验表明,HiSGT显著提升了合成数据与真实病历的统计一致性,并支持慢性病分类等下游应用。该方法克服了传统基于原始编码生成模型的局限,为高保真医疗数据生成提供了新范式,适用于数据增强与隐私保护医疗分析。
原文摘要 · Abstract (English)
Generating realistic synthetic electronic health records (EHRs) holds tremendous promise for accelerating healthcare research, facilitating AI model development and enhancing patient privacy. However, existing generative methods typically treat EHRs as flat sequences of discrete medical codes. This approach overlooks two critical aspects: the inherent hierarchical organization of clinical coding systems and the rich semantic context provided by code descriptions. Consequently, synthetic patient sequences often lack high clinical fidelity and have limited utility in downstream clinical tasks. In this paper, we propose the Hierarchy- and Semantics-Guided Transformer (HiSGT), a novel framework that leverages both hierarchical and semantic information for the generative process. HiSGT constructs a hierarchical graph to encode parent-child and sibling relationships among clinical codes and employs a graph neural network to derive hierarchy-aware embeddings. These are then fused with semantic embeddings extracted from a pre-trained clinical language model (e.g., ClinicalBERT), enabling the Transformer-based generator to more accurately model the nuanced clinical patterns inherent in real EHRs. Extensive experiments on the MIMIC-III and MIMIC-IV datasets demonstrate that HiSGT significantly improves the statistical alignment of synthetic data with real patient records, as well as supports robust downstream applications such as chronic disease classification. By addressing the limitations of conventional raw code-based generative models, HiSGT represents a significant step toward clinically high-fidelity synthetic data generation and a general paradigm suitable for interpretable medical code representation, offering valuable applications in data augmentation and privacy-preserving healthcare analytics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。