RAG在医疗记录推理中比长上下文输入更高效,尤其擅长影像和用药时间线提取。
Evaluating Retrieval-Augmented Generation vs. Long-Context Input for Clinical Reasoning over EHRs
- 用检索增强生成(RAG)替代近期病历输入,提升临床推理效率。
- 在影像与用药任务上,<8K token的RAG性能优于长上下文方法,F1提升0.17-9.83。
- 诊断生成任务受限于记录不全,表现稳定但提升有限,适合关注结构化信息提取者。
目的:评估检索增强生成(RAG)是否可作为电子健康记录(EHR)临床推理中长上下文提示的有效替代方案。方法:定义了三项可在不同医疗系统复现且推理复杂度各异的EHR任务:1)提取影像检查(模态、日期、解剖部位),2)生成抗生素使用时间线,3)识别住院关键诊断。基于美国某学术医疗系统的住院临床笔记,测试三种大模型(GPT-5.4-mini, Mistral Medium 3, DeepSeek V3.1)在不同上下文长度下的表现,比较针对性检索与仅使用最新病历的效果。结果:在影像提取任务中,RAG显著优于最近病历输入,并超越长上下文表现(所有模型F1提升0.17–9.83),仅需少于8K token;在抗生素时间线任务中,<8K检索文本即可达到或超过长上下文近期病历性能(Jaccard得分-3.26至+3.24)。错误分析显示,因跨院转诊导致信息缺失,限制了部分任务表现。然而,诊断生成任务在各类方法与模型间表现基本持平。讨论:RAG在多数任务中展现强令牌效率,尤其在影像提取与用药时间线重建中效果最显著。诊断生成为最挑战任务,可能受文档变异性和评估限制影响。结论:即使新模型支持更长文本处理,RAG仍是大体量EHR临床任务中具有竞争力且高效的方案。
原文摘要 · Abstract (English)
Objective: To evaluate whether retrieval-augmented generation (RAG) can serve as an efficient alternative to long-context prompting for clinical reasoning over electronic health records (EHRs). Methods: We defined three EHR-based tasks that are replicable across health systems and vary in reasoning complexity: 1) extracting imaging procedures (modality, date, and anatomic site), 2) generating timelines of therapeutic antibiotic use, and 3) identifying the key diagnoses for a hospitalization. Using real inpatient clinical notes from a US academic health system, we evaluated three large language models (GPT-5.4-mini, Mistral Medium 3, DeepSeek V3.1) with varying amounts of provided context, comparing targeted retrieval to using the most recent clinical notes. Results: For Imaging Procedures, RAG strongly outperformed recent-note inputs and exceeded long-context performance (by 0.17-9.83 F1 across all models) using fewer than 8K tokens. Similar benefits were observed for Antibiotic Timelines, where <8K of retrieved tokens matched long-context recent-notes performance (between -3.26 to +3.24 Jaccard). Error analysis revealed that missing information in the clinical notes--often due to inter-hospital transfers--limited performance to some extent. However, performance on the Diagnosis Generation task remains largely static across methods and models. Discussion: RAG demonstrated strong token efficiency across tasks, with the clearest and most consistent gains observed for imaging extraction and antibiotic timeline reconstruction. Diagnosis generation proved the most challenging task, suggesting ceiling effects imposed by documentation variability and evaluation constraints. Conclusion: Our results suggest that RAG remains a competitive and efficient approach for clinical tasks over large amounts of EHR, even as newer models become capable of handling increasingly longer amounts of text.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。