用大模型从零散病历中提取结构化语义,提升医疗决策支持
Structured Semantics from Unstructured Notes: Language Model Approaches to EHR-Based Decision Support
- 利用大模型解析自由文本病历,生成语义丰富的表示
- 融合临床文本与结构化数据,提升跨机构数据一致性
- 关注医疗编码与模型公平性,推动AI在医疗中的可信应用
大型语言模型(LLMs)为分析复杂非结构化数据开辟了新路径,尤其在医疗领域。电子健康记录(EHR)包含自由文本临床笔记、结构化检验结果和诊断编码等多种信息。本文探讨如何利用先进语言模型整合这些异构数据源,以改进临床决策支持。我们讨论文本特征在传统高维EHR分析中常被忽略,但能提供语义丰富表示,并有助于不同医疗机构间的数据协调。此外,本文深入分析了引入医疗编码的挑战与机遇,以及确保AI模型泛化能力和公平性的关键问题。
原文摘要 · Abstract (English)
The advent of large language models (LLMs) has opened new avenues for analyzing complex, unstructured data, particularly within the medical domain. Electronic Health Records (EHRs) contain a wealth of information in various formats, including free text clinical notes, structured lab results, and diagnostic codes. This paper explores the application of advanced language models to leverage these diverse data sources for improved clinical decision support. We will discuss how text-based features, often overlooked in traditional high dimensional EHR analysis, can provide semantically rich representations and aid in harmonizing data across different institutions. Furthermore, we delve into the challenges and opportunities of incorporating medical codes and ensuring the generalizability and fairness of AI models in healthcare.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。