用大模型帮医生快速查病历,省时又减负。
A Randomized Controlled Trial and Pilot of Scout: an LLM-Based EHR Search and Synthesis Platform
- 用自然语言查询病历,结果附带原始数据链接可追溯。
- 任务耗时减少37.6%,心理负担、努力程度等显著下降。
- 适合临床和行政人员,尤其需频繁查病历的场景。
电子健康记录(EHR)中的临床文档与数据检索是医生工作负荷和职业倦怠的重要来源。为此,我们开发了Scout——一个基于大语言模型的EHR搜索与综合平台,使医生可通过自然语言查询病历数据,每个回答均包含指向原始数据源的引用,便于内容验证。我们在七个临床专科中开展前瞻性随机、评估者盲法交叉试验,共20名参与者完成200个结构化临床案例任务。比较使用Scout与仅用EHR完成任务的表现,评估指标包括任务完成时间、NASA任务负荷指数(TLX)工作量评分,以及盲法专家对准确性、完整性与相关性的评判。结果显示,Scout使任务完成时间减少37.6%,显著降低感知工作量,尤其在心理需求、努力程度和时间压力方面。非劣效性分析表明,使用Scout完成的任务在准确性、完整性和相关性上不劣于仅用EHR。同期试点部署覆盖200多名用户、20多个专科,三个月内产生超过6,600次交互,揭示多样化的临床与行政应用场景。采用大模型作为评判者框架的自动评估显示错误率低;后续人工抽查发现,多数被自动标记为错误的条目实际上在患者病历中有据可依,凸显人工验证的重要性。这些结果提供了早期实证证据,表明基于大模型的EHR工具可显著减轻临床与行政工作负荷,同时保持输出质量。
原文摘要 · Abstract (English)
Clinical documentation and data retrieval within Electronic Health Records (EHRs) contribute substantially to clinician workload and burnout. To address this, we developed Scout, an LLM-based EHR search and synthesis platform that enables clinicians to query EHR data using natural language. Each response includes citations linking each claim to the original data source, facilitating easy verification of generated content. We conducted a prospective randomized, evaluator-blinded crossover trial across seven clinical specialties (20 participants, 200 structured cases). Participants completed realistic clinical tasks using either Scout or the EHR alone, with outcomes including time to completion, NASA Task Load Index workload scores, and blinded expert adjudication of accuracy, completeness, and relevance. Scout reduced task completion time by 37.6% and significantly decreased perceived workload, with the largest reductions in mental demand, effort, and temporal demand. Non-inferiority analyses showed that tasks completed with Scout maintained accuracy, completeness, and relevance relative to tasks completed with the EHR-only. A concurrent pilot deployment across over 200 users and more than 20 specialties generated over 6,600 interactions in three months, revealing diverse clinical and administrative use cases. Automated evaluation using an LLM-as-judge framework identified errors at low rates. Subsequent manual review of a subset of outputs revealed that most claims flagged by the automated judge as errors were in fact supported by the patient chart, demonstrating the importance of human validation. These findings provide early trial-based evidence that LLM-powered EHR tools can meaningfully reduce clinical and administrative workloads while maintaining output quality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。