用大模型从原始病历中提取临床数据,发现越模糊的问题准确率越低。
An ambiguity taxonomy for evaluating large language model performance on clinical registry abstraction: a multi-site prospective study
- 构建六级模糊度分类体系,评估大模型在真实病历中的表现
- 模型整体准确率91.5%,但面对复杂问题准确率降至62%
- 适合研究医疗AI在高不确定性场景下的可靠性
目标:评估大语言模型(LLM)在未经处理的电子病历(EMR)数据上进行临床注册信息抽取的表现。方法:在一家医学中心开展试点研究,模型为每个注册问题识别候选数据源,由人工摘要员据此定义问题特定文档集;在第二家中心的验证研究中,模型使用该文档集回答问题。在查看输出前,两名摘要员独立建立基准答案,并将每个问题归入六个类别(按模糊度与临床推理需求排序):用药/事件标记、二元临床存在、行政类、定量检验/生理指标、临床判断和事件时间。结果:分析样本包含9,430个摘要员答案,对应4,715个共识答案(试点501例,验证4,214例)。试点中,各问题的候选数据源平均数为14.6(标准差13.9,人口学)至89.2(标准差56.1,病史与危险因素)。验证中,人工一致性约98%;模型答案与共识完全一致占87%,部分一致2%,不一致9%。157个至少有20个答案的问题,平均准确率为91.5%(标准差13.4%),随着模糊度上升而下降,从用药/事件标记的96%降至事件时间问题的62%。结论:大模型在原始病历上回答临床注册问题的准确率远低于人类,且准确率随模糊度和所需临床推理程度增加而显著降低。
原文摘要 · Abstract (English)
Objective: To evaluate large language model (LLM) performance on unprocessed electronic medical record (EMR) data for clinical registry abstraction. Methods: We evaluated LLM performance answering registry questions for the American College of Cardiology National Cardiovascular Data Registry (ACC NCDR). In a pilot study at an academic medical center, the model identified candidate data sources for each registry question and experienced abstractors used these results to define question-specific document sets. In a validation study at a second center with a second ACC NCDR registry, the LLM answered questions using the question-specific document sets. Before reviewing any output, two abstractors independently established the ground truth and assigned each question to one of six categories, ordered by the ambiguity and clinical reasoning required to resolve it: Medication/Event Flag, Binary Clinical Presence, Administrative, Quantitative Laboratory/Physiologic, Clinical Interpretation, and Event Timing. Results: The analytical sample comprised 9,430 abstractor answers reconciled to 4,715 consensus answers (501 pilot; 4,214 validation). In the pilot, candidate data sources per question averaged between 14.6 (SD 13.9) for demographics and 89.2 (SD 56.1) for history and risk factors. In validation, human inter-rater agreement was approximately 98\% while 87\% of LLM answers exactly matched consensus, 2\% partially, and 9\% did not. Mean question-level accuracy was 91.5\% (SD 13.4\%) across 157 questions with at least 20 answers, and declined as ambiguity increased, from 96\% for Medication/Event Flag to 62\% for Event Timing questions. Conclusions: LLMs answering clinical registry questions on unprocessed EMR data achieved far lower accuracy than human abstractors. LLM accuracy fell steadily as ambiguity and the level of required clinical reasoning increased.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。