用多维度预训练提升电子病历预测能力,显著改善诊断与心衰预测效果。
MPLite: Multi-Aspect Pretraining for Mining Clinical Health Records
- 通过轻量神经网络融合检验结果与结构化数据进行多方面预训练
- 在MIMIC-III和MIMIC-IV上实现更高的加权F1与召回率
- 特别适合缺乏未来就诊标注的单次就诊记录建模
医疗数字化带来海量电子健康记录(EHR),为机器学习预测患者健康结局提供宝贵数据。然而,因缺乏下一次就诊信息标注,单次就诊记录常被忽略,限制了模型的预测与表达能力。本文提出MPLite框架,采用基于检验结果的轻量级神经网络进行多方面预训练,增强医学概念表征并预测个体未来健康状况。该方法融合结构化医疗数据与检验结果信息,充分挖掘患者入院记录价值。设计的预训练模块基于检验结果预测医学编码,通过多维度特征融合确保预测鲁棒性。在MIMIC-III和MIMIC-IV数据集上的实验表明,该方法在诊断预测与心衰预测任务中优于现有模型,取得更高加权F1与召回率。本工作揭示了整合多源数据对推动医疗预测建模的潜力。
原文摘要 · Abstract (English)
The adoption of digital systems in healthcare has resulted in the accumulation of vast electronic health records (EHRs), offering valuable data for machine learning methods to predict patient health outcomes. However, single-visit records of patients are often neglected in the training process due to the lack of annotations of next-visit information, thereby limiting the predictive and expressive power of machine learning models. In this paper, we present a novel framework MPLite that utilizes Multi-aspect Pretraining with Lab results through a light-weight neural network to enhance medical concept representation and predict future health outcomes of individuals. By incorporating both structured medical data and additional information from lab results, our approach fully leverages patient admission records. We design a pretraining module that predicts medical codes based on lab results, ensuring robust prediction by fusing multiple aspects of features. Our experimental evaluation using both MIMIC-III and MIMIC-IV datasets demonstrates improvements over existing models in diagnosis prediction and heart failure prediction tasks, achieving a higher weighted-F1 and recall with MPLite. This work reveals the potential of integrating diverse aspects of data to advance predictive modeling in healthcare.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。