arXiv:2606.18471cs.CL2026-06被引 3

评估大模型在临床文本中保留诊断不确定性能力,发现其准确率不足一半。

Possible or Definite? A Benchmark for Evaluating Diagnostic Uncertainty Preservation in Clinical Text

论文配图:Possible or Definite? A Benchmark for Evaluating Diagnostic Uncertainty Preservation in Clinical Text
图 1 · 摘自论文原文
  • 构建含9184个标注的临床文本基准,涵盖五级诊断不确定性。
  • 大模型仅在不足一半情况下正确保留原始不确定性表达。
  • 尤其难以区分相邻级别的不确定性,影响临床决策安全。

大型语言模型(LLMs)在临床文本任务如摘要与修订中应用日益广泛。尽管多数研究关注生成文本的流畅性与连贯性,但对模型是否正确保留诊断不确定性仍缺乏深入探索。在临床实践中,'可能肺炎'这类表述传递证据强度,直接影响随访检查与治疗决策。改变这些不确定性表达会彻底改变临床意义。本文通过两步系统评估该问题:首先,构建包含1,200份临床文档、9,184个不确定性标注的基准,覆盖五个级别;其次,评估三种大模型在此基准上的表现。结果表明:(1)大模型在保留原始不确定性线索方面表现不佳,正确率常低于一半;(2)在相邻等级间细微差异的区分上尤为困难。该研究揭示了现有评估指标未捕捉到的失效模式,为大模型在临床工作流中的安全部署提供重要启示。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly used for clinical text tasks such as summarization and revision. While most studies evaluate the fluency and coherence of LLM-generated text, whether LLMs correctly preserve diagnostic uncertainty remains underexplored. In clinical practice, phrases such as ``possible pneumonia'' communicate the strength of available evidence and directly guide decisions about follow-up testing and treatment. Altering these uncertainty expressions can change the clinical meaning entirely. In this paper, we systematically evaluated this problem in two steps. First, we constructed a benchmark of 1,200 clinical documents with 9,184 uncertainty annotations across five levels. Second, we evaluated three LLMs on this benchmark. Our results show that (1) LLMs preserve the original uncertainty cues poorly, often less than half the time; (2) LLMs struggle with nuanced distinctions between adjacent levels. This work reveals a failure mode not captured by standard evaluation metrics and provides implications for the safe deployment of LLMs in clinical workflows.

临床AI不确定性大模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。