构建首个结构化病历大模型评估基准,助力医疗AI可信落地
EHRStruct: A Comprehensive Benchmark Framework for Evaluating Large Language Models on Structured Electronic Health Record Tasks
- 设计11项临床任务,基于2200个样本评估大模型处理病历数据能力
- 发现多数任务需强推理能力,通用模型表现远低于专业需求
- 提出代码增强方法EHRMaster,显著提升性能,适合医疗AI研究者参考
结构化电子健康记录(EHR)数据以关系表形式存储患者信息,在临床决策中发挥核心作用。近期研究尝试用大语言模型(LLM)处理此类数据,在多项临床任务中展现出潜力。然而,缺乏标准化评估框架与明确任务定义,导致难以系统性评估和比较模型在结构化EHR数据上的表现。为此,我们提出EHRStruct,一个专为评估LLM在结构化EHR任务中表现而设计的基准框架。EHRStruct定义了11项覆盖多样化临床需求的代表性任务,包含从两个广泛使用的EHR数据集提取的2,200个任务特定评估样本。我们使用EHRStruct评估了20个先进且具有代表性的大模型,涵盖通用与医学专用模型。进一步分析输入格式、少样本泛化能力及微调策略等关键因素对模型表现的影响,并与11种最先进的基于LLM的结构化数据推理增强方法进行对比。结果表明,许多结构化EHR任务对模型的理解与推理能力提出了极高要求。针对此问题,我们提出EHRMaster——一种代码增强方法,在多项任务上达到最先进性能,为未来研究提供实用指导。
原文摘要 · Abstract (English)
Structured Electronic Health Record (EHR) data stores patient information in relational tables and plays a central role in clinical decision-making. Recent advances have explored the use of large language models (LLMs) to process such data, showing promise across various clinical tasks. However, the absence of standardized evaluation frameworks and clearly defined tasks makes it difficult to systematically assess and compare LLM performance on structured EHR data. To address these evaluation challenges, we introduce EHRStruct, a benchmark specifically designed to evaluate LLMs on structured EHR tasks. EHRStruct defines 11 representative tasks spanning diverse clinical needs and includes 2,200 task-specific evaluation samples derived from two widely used EHR datasets. We use EHRStruct to evaluate 20 advanced and representative LLMs, covering both general and medical models. We further analyze key factors influencing model performance, including input formats, few-shot generalisation, and finetuning strategies, and compare results with 11 state-of-the-art LLM-based enhancement methods for structured data reasoning. Our results indicate that many structured EHR tasks place high demands on the understanding and reasoning capabilities of LLMs. In response, we propose EHRMaster, a code-augmented method that achieves state-of-the-art performance and offers practical insights to guide future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。