融合医学本体与临床数据,生成更准确的医疗编码嵌入。
KEEP: Integrating Medical Ontologies with Clinical Data for Robust Code Embeddings
- 先用本体图生成嵌入,再通过患者数据自适应优化。
- 在英国生物银行和MIMIC IV上优于传统与大模型方法。
- 无需额外训练,适合资源受限环境使用。
医疗机器学习需要有效的结构化医疗编码表示,但现有方法存在权衡:基于知识图的方法能捕捉正式关系,却忽略真实世界模式;数据驱动方法学习经验关联,常忽视医学术语中的结构化知识。我们提出KEEP(知识保持与经验精炼嵌入过程),一种高效框架,通过结合知识图嵌入与临床数据的自适应学习来弥合这一差距。KEEP首先从知识图生成嵌入,然后在患者记录上进行正则化训练,以自适应整合经验模式,同时保持本体关系。重要的是,KEEP生成最终嵌入无需任务特定辅助或端到端训练,支持多种下游应用和模型架构。在英国生物银行和MIMIC IV的结构化电子病历上评估显示,KEEP在捕捉语义关系和预测临床结果方面优于传统方法与基于语言模型的方法。此外,KEEP计算开销极低,特别适合资源受限环境。
原文摘要 · Abstract (English)
Machine learning in healthcare requires effective representation of structured medical codes, but current methods face a trade off: knowledge graph based approaches capture formal relationships but miss real world patterns, while data driven methods learn empirical associations but often overlook structured knowledge in medical terminologies. We present KEEP (Knowledge preserving and Empirically refined Embedding Process), an efficient framework that bridges this gap by combining knowledge graph embeddings with adaptive learning from clinical data. KEEP first generates embeddings from knowledge graphs, then employs regularized training on patient records to adaptively integrate empirical patterns while preserving ontological relationships. Importantly, KEEP produces final embeddings without task specific auxiliary or end to end training enabling KEEP to support multiple downstream applications and model architectures. Evaluations on structured EHR from UK Biobank and MIMIC IV demonstrate that KEEP outperforms both traditional and Language Model based approaches in capturing semantic relationships and predicting clinical outcomes. Moreover, KEEP's minimal computational requirements make it particularly suitable for resource constrained environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。