构建临床笔记事实性评估数据集,检验大模型分解医学文本能力。
FactEHR: A Dataset for Evaluating Factuality in Clinical Notes Using LLMs
- 用大模型将临床笔记拆解为细粒度可验证的原子事实
- 生成98万+蕴含对,覆盖2168份多类型病历
- 发现不同模型表现差异大,提示需提升医疗文本理解能力
在医疗领域安全使用大语言模型(LLMs)需要验证和溯源事实。事实性评估的核心是事实分解,即将复杂临床陈述拆分为单一信息单元以供验证。现有方法利用大模型将原文重写为简洁句子,实现细粒度验证。但临床文档因术语密集、记录类型多样,使事实分解面临独特挑战,研究仍不足。为此,我们提出FactEHR,一个基于自然语言推理(NLI)的数据集,涵盖来自三家医院系统的四类临床笔记共2,168份,生成987,266个蕴含对。我们从大模型的蕴含判断到定性分析进行多维度评估。包括临床医生评审的结果显示,大模型在事实分解上表现差异显著:Gemini-1.5-Flash始终生成相关且准确的事实,而Llama-3 8B则产出更少且一致性差。结果凸显了提升大模型在临床文本中支持事实验证能力的必要性。
原文摘要 · Abstract (English)
Verifying and attributing factual claims is essential for the safe and effective use of large language models (LLMs) in healthcare. A core component of factuality evaluation is fact decomposition, the process of breaking down complex clinical statements into fine-grained atomic facts for verification. Recent work has proposed fact decomposition, which uses LLMs to rewrite source text into concise sentences conveying a single piece of information, to facilitate fine-grained fact verification. However, clinical documentation poses unique challenges for fact decomposition due to dense terminology and diverse note types and remains understudied. To address this gap and explore these challenges, we present FactEHR, an NLI dataset consisting of document fact decompositions for 2,168 clinical notes spanning four types from three hospital systems, resulting in 987,266 entailment pairs. We assess the generated facts on different axes, from entailment evaluation of LLMs to a qualitative analysis. Our evaluation, including review by the clinicians, reveals substantial variability in LLM performance for fact decomposition. For example, Gemini-1.5-Flash consistently generates relevant and accurate facts, while Llama-3 8B produces fewer and less consistent outputs. The results underscore the need for better LLM capabilities to support factual verification in clinical text.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。