解决电子病历模型无法处理未知医学编码的问题
MedRep: Medical Concept Representation for General Electronic Health Record Foundation Models
- 基于OMOP标准用大模型生成概念定义并融合本体图谱
- 在多个预测任务中优于传统模型和旧式分词器
- 适合需要跨系统通用性的医疗AI研究者
电子病历(EHR)基础模型在多项医疗任务中表现优异,但存在根本性局限:无法处理词汇表外的未见医学编码,限制了模型的泛化能力与不同词汇体系模型间的整合。为此,我们提出一种基于观测医学结局合作计划(OMOP)公共数据模型(CDM)的新颖医学概念表征方法(MedRep)。通过大语言模型(LLM)提示为每个概念添加最小定义,并结合OMOP词汇本体图谱补充文本表征。该方法在多种预测任务中均优于原始EHR基础模型及先前提出的医学代码分词器。外部验证进一步证明了MedRep的泛化能力。
原文摘要 · Abstract (English)
Electronic health record (EHR) foundation models have been an area ripe for exploration with their improved performance in various medical tasks. Despite the rapid advances, there exists a fundamental limitation: Processing unseen medical codes out of vocabulary. This problem limits the generalizability of EHR foundation models and the integration of models trained with different vocabularies. To alleviate this problem, we propose a set of novel medical concept representations (MedRep) for EHR foundation models based on the observational medical outcome partnership (OMOP) common data model (CDM). For concept representation learning, we enrich the information of each concept with a minimal definition through large language model (LLM) prompts and complement the text-based representations through the graph ontology of OMOP vocabulary. Our approach outperforms the vanilla EHR foundation model and the model with a previously introduced medical code tokenizer in diverse prediction tasks. We also demonstrate the generalizability of MedRep through external validation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。