arXiv:2511.04418cs.LGcs.CL2025-11被引 16

现有大模型不确定性量化在模糊语境下失效,导致结果近乎随机。

The Illusion of Certainty: Uncertainty Quantification for LLMs Fails under Ambiguity

  • 构建首个带真实答案分布的模糊问答数据集,揭示模型在歧义下的表现退化
  • 多种量化方法在模糊数据上准确率降至接近随机水平(约50%)
  • 适用于关注大模型可信度、部署风险的研究者与工程师

大语言模型(LLM)的不确定性量化(UQ)对可信部署至关重要。然而现实语言天然存在歧义,反映为偶然性不确定性,而现有UQ方法通常在无歧义任务上评估。本文发现,当前不确定性估计器在无歧义假设下表现良好,但在模糊数据上性能退化至接近随机。为此,我们提出MAQA*和AmbigQA*,首个配备基于事实共现估计的真实答案分布的模糊问答数据集。该现象在不同估计范式中均一致:使用预测分布、模型内部表示及模型集成。理论上证明,预测分布与集成类估计器在歧义下存在根本局限。研究揭示了当前LLM UQ方法的关键缺陷,呼吁重构建模范式。

原文摘要 · Abstract (English)

Accurate uncertainty quantification (UQ) in Large Language Models (LLMs) is critical for trustworthy deployment. While real-world language is inherently ambiguous, reflecting aleatoric uncertainty, existing UQ methods are typically benchmarked against tasks with no ambiguity. In this work, we demonstrate that while current uncertainty estimators perform well under the restrictive assumption of no ambiguity, they degrade to close-to-random performance on ambiguous data. To this end, we introduce MAQA* and AmbigQA*, the first ambiguous question-answering (QA) datasets equipped with ground-truth answer distributions estimated from factual co-occurrence. We find this performance deterioration to be consistent across different estimation paradigms: using the predictive distribution itself, internal representations throughout the model, and an ensemble of models. We show that this phenomenon can be theoretically explained, revealing that predictive-distribution and ensemble-based estimators are fundamentally limited under ambiguity. Overall, our study reveals a key shortcoming of current UQ methods for LLMs and motivates a rethinking of current modeling paradigms.

大模型不确定性可信度问答系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。