arXiv:2502.17403cs.LGcs.AI2025-02被引 32

用自然语言描述医疗编码,让通用大模型直接处理病历数据

Large Language Models are Powerful Electronic Health Record Encoders

  • 将医疗代码转为自然语言文本,使通用大模型可直接编码电子病历
  • 在15项临床任务上表现媲美专用EHR模型CLMBR-T-Base
  • 适合缺乏私有数据的机构快速部署,兼具通用性与可迁移性

电子病历(EHR)在临床预测中潜力巨大,但其复杂性和异质性给传统机器学习带来挑战。针对这一问题,基于未标注EHR数据训练的领域专用基础模型已显示出更高的预测准确率和泛化能力。然而,其发展受限于数据访问困难和机构间术语差异。本文将EHR数据中的医疗编码替换为自然语言描述,使通用大语言模型(LLM)无需接触私有医疗训练数据即可生成高维嵌入表示,用于下游预测任务。实验表明,基于LLM的嵌入在EHRSHOT基准的15项临床任务上表现与专用模型CLMBR-T-Base相当。在英国生物银行(UK Biobank)的外部验证中,部分任务上显著优于对比模型,归因于更广的词汇覆盖和稍优的泛化能力。整体揭示了专用模型计算效率与通用嵌入可移植性、数据独立性之间的权衡。

原文摘要 · Abstract (English)

Electronic Health Records (EHRs) offer considerable potential for clinical prediction, but their complexity and heterogeneity challenge traditional machine learning. Domain-specific EHR foundation models trained on unlabeled EHR data have shown improved predictive accuracy and generalization. However, their development is constrained by limited data access and site-specific vocabularies. We convert EHR data into plain text by replacing medical codes with natural-language descriptions, enabling general-purpose Large Language Models (LLMs) to produce high-dimensional embeddings for downstream prediction tasks without access to private medical training data. LLM-based embeddings perform on par with a specialized EHR foundation model, CLMBR-T-Base, across 15 clinical tasks from the EHRSHOT benchmark. In an external validation using the UK Biobank, an LLM-based model shows statistically significant improvements for some tasks, which we attribute to higher vocabulary coverage and slightly better generalization. Overall, we reveal a trade-off between the computational efficiency of specialized EHR models and the portability and data independence of LLM-based embeddings.

大模型电子病历临床预测嵌入表示

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。