用AI评估医疗摘要真实性,准确率接近专家水平。
MedFactEval and MedAgentBrief: A Framework and Workflow for Generating and Evaluating Factual Clinical Summaries
- 用多模型投票机制自动检测生成摘要中的关键事实缺失。
- 评估框架与7位医生共识一致,准确性媲美单个专家。
- 适合医疗AI研发者和临床系统部署人员参考。
大型语言模型生成的临床文本在真实性和准确性方面缺乏可靠评估,而人工评审难以规模化。为此,我们提出两个互补方案:一是MedFactEval,一种可扩展的事实导向评估框架,由临床医生定义关键信息点,再通过“多模型投票”(LLM Jury)判断其是否被包含在生成摘要中;二是MedAgentBrief,一种不依赖特定模型、多步骤的生成工作流,用于生成高质量、准确的出院摘要。为验证评估框架的有效性,我们基于七名医生对住院病例中关键事实的多数意见建立了金标准。结果显示,MedFactEval的LLM Jury与该专家小组达成几乎完美的一致性(Cohen's kappa=81%),其性能在统计上不劣于单个医生(kappa=67%,P<0.001)。本研究提供了可靠的评估框架(MedFactEval)和高效的生成流程(MedAgentBrief),推动生成式AI在临床工作流中的负责任应用。
原文摘要 · Abstract (English)
Evaluating factual accuracy in Large Language Model (LLM)-generated clinical text is a critical barrier to adoption, as expert review is unscalable for the continuous quality assurance these systems require. We address this challenge with two complementary contributions. First, we introduce MedFactEval, a framework for scalable, fact-grounded evaluation where clinicians define high-salience key facts and an "LLM Jury"--a multi-LLM majority vote--assesses their inclusion in generated summaries. Second, we present MedAgentBrief, a model-agnostic, multi-step workflow designed to generate high-quality, factual discharge summaries. To validate our evaluation framework, we established a gold-standard reference using a seven-physician majority vote on clinician-defined key facts from inpatient cases. The MedFactEval LLM Jury achieved almost perfect agreement with this panel (Cohen's kappa=81%), a performance statistically non-inferior to that of a single human expert (kappa=67%, P < 0.001). Our work provides both a robust evaluation framework (MedFactEval) and a high-performing generation workflow (MedAgentBrief), offering a comprehensive approach to advance the responsible deployment of generative AI in clinical workflows.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。