arXiv:2505.23854cs.CLcs.AI2025-05被引 24

对比80个大模型,发现语言类不确定性度量更可靠。

Revisiting Uncertainty Estimation and Calibration of Large Language Models

  • 用三种黑盒方法评估大模型不确定性,聚焦推理与知识任务差异。
  • 语言类度量(LVU)在校准和区分上优于其他方法,且更易理解。
  • 模型规模和量化不直接影响可靠性,需多角度评估不确定性的真伪。

随着大语言模型在高风险场景中广泛应用,可靠的不确定性估计对确保其安全可信部署至关重要。本文对不确定性估计进行了迄今最全面的研究,评估了80个模型,涵盖开源与闭源模型、密集与专家混合(MoE)架构、推理与非推理模式、量化变体及参数量从0.6B到671B的多种配置。重点考察三种典型的单次前向传播黑盒方法:基于标记概率的不确定性(TPU)、数值表述不确定性(NVU)和语言表述不确定性(LVU),并利用具有挑战性的MMLU-Pro基准测试进行不确定性校准与选择性分类评估,该基准覆盖推理密集型与知识型任务。结果表明,LVU在整体性能上持续优于TPU与NVU,具备更强的校准能力与区分能力,且更具可解释性。研究还发现,高准确率并不等同于可靠不确定性;模型规模、后训练、推理能力及量化均影响估计表现。值得注意的是,模型在推理任务上的不确定性估计优于知识密集型任务,且良好校准未必带来有效错误排序。这些发现强调了多维度评估的必要性,并将LVU定位为提升大模型在真实场景中可靠性的实用工具。

原文摘要 · Abstract (English)

As large language models (LLMs) are increasingly deployed in high-stakes applications, robust uncertainty estimation is essential for ensuring the safe and trustworthy deployment of LLMs. We present the most comprehensive study to date of uncertainty estimation in LLMs, evaluating 80 models spanning open- and closed-source families, dense and Mixture-of-Experts (MoE) architectures, reasoning and non-reasoning modes, quantization variants and parameter scales from 0.6B to 671B. Focusing on three representative black-box single-pass methods, including token probability-based uncertainty (TPU), numerical verbal uncertainty (NVU), and linguistic verbal uncertainty (LVU), we systematically evaluate uncertainty calibration and selective classification using the challenging MMLU-Pro benchmark, which covers both reasoning-intensive and knowledge-based tasks. Our results show that LVU consistently outperforms TPU and NVU, offering stronger calibration and discrimination while being more interpretable. We also find that high accuracy does not imply reliable uncertainty, and that model scale, post-training, reasoning ability and quantization all influence estimation performance. Notably, LLMs exhibit better uncertainty estimates on reasoning tasks than on knowledge-heavy ones, and good calibration does not necessarily translate to effective error ranking. These findings highlight the need for multi-perspective evaluation and position LVU as a practical tool for improving the reliability of LLMs in real-world settings.

大模型不确定性校准评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。