通过分层原型学习,提升电子病历预测的准确性与可解释性。
ProtoEHR: Hierarchical Prototype Learning for EHR-based Healthcare Predictions
- 构建医疗代码、就诊记录、患者三层结构的原型学习框架
- 在5个临床任务中超越现有方法,实现更精准稳健预测
- 提供代码、就诊、患者三个层级的可解释分析,适合临床辅助决策
数字医疗系统积累了海量电子健康记录(EHR)数据,为人工智能在医疗预测中的应用提供了基础。然而,现有研究多聚焦于EHR数据的孤立部分,限制了预测性能和可解释性。为此,我们提出ProtoEHR,一种可解释的分层原型学习框架,充分挖掘EHR数据的多层次结构以增强医疗预测能力。具体而言,ProtoEHR建模医疗代码、医院就诊和患者三个层次间的关联关系。首先利用大语言模型提取医疗代码间的语义关系,构建医疗知识图谱作为知识源;在此基础上,设计分层表示学习框架,捕捉跨三个层次的上下文表示,并在每层引入原型信息以捕捉内在相似性,提升泛化能力。我们在两个公开数据集上对ProtoEHR进行评估,涵盖死亡率预测、再入院预测、住院时长预测、药物推荐和表型预测五项临床重要任务。结果表明,ProtoEHR在准确率、鲁棒性和可解释性方面均优于现有基线方法。此外,该模型在代码、就诊和患者层面提供可解释洞察,助力临床决策。
原文摘要 · Abstract (English)
Digital healthcare systems have enabled the collection of mass healthcare data in electronic healthcare records (EHRs), allowing artificial intelligence solutions for various healthcare prediction tasks. However, existing studies often focus on isolated components of EHR data, limiting their predictive performance and interpretability. To address this gap, we propose ProtoEHR, an interpretable hierarchical prototype learning framework that fully exploits the rich, multi-level structure of EHR data to enhance healthcare predictions. More specifically, ProtoEHR models relationships within and across three hierarchical levels of EHRs: medical codes, hospital visits, and patients. We first leverage large language models to extract semantic relationships among medical codes and construct a medical knowledge graph as the knowledge source. Building on this, we design a hierarchical representation learning framework that captures contextualized representations across three levels, while incorporating prototype information within each level to capture intrinsic similarities and improve generalization. To perform a comprehensive assessment, we evaluate ProtoEHR in two public datasets on five clinically significant tasks, including prediction of mortality, prediction of readmission, prediction of length of stay, drug recommendation, and prediction of phenotype. The results demonstrate the ability of ProtoEHR to make accurate, robust, and interpretable predictions compared to baselines in the literature. Furthermore, ProtoEHR offers interpretable insights on code, visit, and patient levels to aid in healthcare prediction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。