arXiv:2501.16672cs.AIcs.CL2025-01被引 20

用病历数据验证大模型生成的临床文本是否真实可靠。

VeriFact: Verifying Facts in LLM-Generated Clinical Text with Electronic Health Records

  • 结合检索增强生成与大模型评判,自动比对生成文本与患者病历。
  • 在新数据集上达到92.7%准确率,超过普通医生水平。
  • 适合开发医疗大模型应用的研究者与临床工程师使用。

目前缺乏确保大语言模型(LLM)在临床医学中生成文本事实准确性的方法。VeriFact是一种人工智能系统,结合检索增强生成与大模型作为裁判(LLM-as-a-Judge),验证LLM生成的临床文本是否被患者的电子健康记录(EHR)支持。为评估该系统,我们引入了VeriFact-BHC数据集,将出院小结中的简要住院经过分解为若干简单陈述,并由临床医生标注每条陈述是否被患者EHR临床记录支持。尽管医生间最高一致性为88.5%,但VeriFact与去噪并仲裁后的平均人类医生基准相比,最高达成92.7%的一致性,表明其在比对文本与病历时已超越平均水平。VeriFact有望通过消除当前评估瓶颈,加速基于LLM的EHR应用发展。

原文摘要 · Abstract (English)

Methods to ensure factual accuracy of text generated by large language models (LLM) in clinical medicine are lacking. VeriFact is an artificial intelligence system that combines retrieval-augmented generation and LLM-as-a-Judge to verify whether LLM-generated text is factually supported by a patient's medical history based on their electronic health record (EHR). To evaluate this system, we introduce VeriFact-BHC, a new dataset that decomposes Brief Hospital Course narratives from discharge summaries into a set of simple statements with clinician annotations for whether each statement is supported by the patient's EHR clinical notes. Whereas highest agreement between clinicians was 88.5%, VeriFact achieves up to 92.7% agreement when compared to a denoised and adjudicated average human clinican ground truth, suggesting that VeriFact exceeds the average clinician's ability to fact-check text against a patient's medical record. VeriFact may accelerate the development of LLM-based EHR applications by removing current evaluation bottlenecks.

医疗AI事实验证大模型评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。