arXiv:2605.27016cs.CLcs.AI2026-05

检验大模型不确定性估计与幻觉的关系,发现二者关联很弱且不稳定。

Evaluating the Relevance of Uncertainty Estimators for LLM Hallucination

论文配图:Evaluating the Relevance of Uncertainty Estimators for LLM Hallucination
图 1 · 摘自论文原文
  • 对比多种不确定性估算方法在不同幻觉场景下的表现
  • 发现不确定性与幻觉关联性弱,且随模型和幻觉类型变化
  • 提醒别把不确定性当幻觉预警,适合研究幻觉机制的学者

大语言模型(LLMs)容易产生幻觉,即输出与输入或训练数据无关的陈述,影响可靠部署。尽管已有众多不确定性估计(UE)方法被提出以量化模型置信度,常被视为模型失败的代理信号,但其与幻觉的关系尚未充分阐明。本文对LLM中不确定性估计器与幻觉之间的关联进行了系统的实证研究。我们不预设这种关联,而是直接评估其存在条件与程度。考察了信息论、采样型和反思型等多种不确定性估计方法,在包括RAGTruth和HalluLens在内的四个互补基准上,分析其在内在幻觉(违背输入忠实性)和外在幻觉(与训练数据无关的主张)中的行为。结果表明,该关联高度可变且常较弱,取决于幻觉类型及具体模型。这些发现挑战了将不确定性作为幻觉直接信号的做法,并明确了其提供有效信息的条件。

原文摘要 · Abstract (English)

Large language models (LLMs) are prone to hallucinations, i.e., statements unsupported by the input or training data, hindering reliable deployment. In parallel, numerous uncertainty estimation (UE) methods have been proposed to quantify model confidence and are often implicitly treated as proxies for model failure. However, the relationship between uncertainty and hallucinations remains insufficiently characterized. We present a systematic empirical study of the association between uncertainty estimators and hallucinations in LLMs. Rather than assuming this association, we evaluate directly when and to what extent it holds. We consider a diverse set of uncertainty estimators, including information-theoretic, sampling-based, and reflexive estimators, and examine their behavior across hallucination settings. Our experiments cover both intrinsic hallucinations (violations of input faithfulness) and extrinsic hallucinations (unsupported claims relative to training data), using four complementary benchmarks, including RAGTruth and HalluLens. We find that the association is highly variable and often weak, depending on the hallucination type and the LLM under evaluation. These results challenge the use of uncertainty as a direct signal of hallucination and clarify when it provides actionable information.

大模型幻觉不确定性估计可信推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。