arXiv:2608.09080cs.CLcs.AI2026-08

LLM在临床信息缺失时仍盲目自信,可能误导医疗决策。

When Confidence Fails: Overconfidence in LLMs under Uncertainty and Missing Clinical Information

论文配图:When Confidence Fails: Overconfidence in LLMs under Uncertainty and Missing Clinical Information
图 1 · 摘自论文原文
  • 通过修改提示和移除正确答案,模拟临床不确定性场景。
  • 500题测试中,模型准确率下降但自信度不降,误判率飙升。
  • 部分模型无法识别信息不足,持续输出高置信幻觉答案。

大型语言模型(LLMs)在医学问答与临床推理任务中表现优异,但在不确定情境下的可靠性尚不明确,这对其在高风险临床环境中的部署构成严重隐患。错误预测本身已具风险,而高自信的错误预测更可能误导临床决策。本文针对临床信息不确定性对LLMs进行系统性行为分析,基于MedMCQA数据集构建双模评估框架:一是通过提示修改引入语言层面的模糊性线索,模拟模糊临床情境;二是构造答案移除场景,刻意排除正确选项,迫使模型识别信息不足并选择不回答。我们采用校准差距、期望校准误差(ECE)、非安全自信错误率(UCER)等指标,在500道医学问题上分析模型准确率与置信度行为。结果表明,尽管准确率随不确定性上升而下降,模型置信度仍与之严重错位,导致非安全自信错误显著增加,说明模型置信度对临床关键信息丢失反应迟钝。此外,不同模型在无法获取正确答案时的拒答能力差异显著,部分模型仍持续生成高置信的虚构答案。研究揭示当前LLMs在认知可靠性上的重大缺陷,强调部署前必须建立具备不确定性感知能力的评估体系。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have achieved strong performance in medical question answering and clinical reasoning tasks. However, their reliability under uncertainty remains poorly understood which raises critical concerns for deployment in high-stakes clinical settings. In such environments, incorrect predictions are inherently risky, but confident incorrect predictions can be particularly harmful as they may mislead clinical decision-making. In this paper, we conduct a systematic behavioral analysis of LLMs under clinical information uncertainty. We propose an evaluation framework based on the MedMCQA dataset consisting of two complementary uncertainty settings. First, we introduce linguistic uncertainty cues through prompt modifications to simulate ambiguous clinical contexts. Second, we construct an answer removal setting, wherein the correct option is deliberately excluded mandating the model to recognize insufficient information and abstain. We analyze both model accuracy and confidence behavior using multiple calibration metrics including calibration gap, Expected Calibration Error (ECE), and Unsafe Confident Error Rate (UCER) across 500 medical questions. Our results reveal a consistent failure mode, i.e., although accuracy degrades under increasing uncertainty, model confidence remains misaligned with accuracy. This leads to a substantial increase in unsafe confident errors, indicating that model confidence remains largely insensitive to clinically meaningful information loss. Furthermore, we observe significant variation across models in their ability to abstain when the correct answer is unavailable, with some models persistently producing high confidence hallucinated answers. These findings expose critical limitations in the epistemic reliability of current LLMs and highlight the need for uncertainty aware evaluation methods prior to their deployment in clinical workflows.

大模型临床决策置信度不确定性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。