评测大模型理解医疗表格数据能力,提升病患信息提取准确率。
Evaluating LLM Abilities to Understand Tabular Electronic Health Records: A Comprehensive Study of Patient Data Extraction and Retrieval
- 通过优化特征选择与序列化方法,性能提升26.79%。
- 引入相关示例的上下文学习使提取效果提高5.95%。
- 为医疗搜索类大模型设计提供可落地指南。
电子健康记录(EHR)表格因其高维稀疏性及隐藏的上下文依赖关系而带来独特挑战。本研究首次系统评估大语言模型(LLM)在理解EHR方面的能力,聚焦患者数据提取与检索任务。基于两个骨干模型Llama2和Meditron,利用MIMICSQL数据集开展大量实验,探究提示结构、指令、上下文与示范对任务表现的影响。定量与定性分析表明,最优的特征选择与序列化方法可使任务性能相比朴素方法提升最高达26.79%;采用相关示例的上下文学习设置,数据提取性能提升5.95%。基于研究发现,提出一套有助于设计支持医疗搜索的LLM模型的指导原则。
原文摘要 · Abstract (English)
Electronic Health Record (EHR) tables pose unique challenges among which is the presence of hidden contextual dependencies between medical features with a high level of data dimensionality and sparsity. This study presents the first investigation into the abilities of LLMs to comprehend EHRs for patient data extraction and retrieval. We conduct extensive experiments using the MIMICSQL dataset to explore the impact of the prompt structure, instruction, context, and demonstration, of two backbone LLMs, Llama2 and Meditron, based on task performance. Through quantitative and qualitative analyses, our findings show that optimal feature selection and serialization methods can enhance task performance by up to 26.79% compared to naive approaches. Similarly, in-context learning setups with relevant example selection improve data extraction performance by 5.95%. Based on our study findings, we propose guidelines that we believe would help the design of LLM-based models to support health search.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。