arXiv:2510.26995stat.MEcs.AI2025-10被引 6

LLM估计常过度自信,新评测发现99%置信区间实际覆盖仅65%

LLMs are Overconfident: Evaluating Confidence Interval Calibration with FermiEval

  • 用费米题基准与严格评分规则评估模型置信区间
  • 多数模型99%置信区间实际覆盖率仅65%,显著低于目标
  • 提出校准方法使覆盖率达99%,并解释过自信的感知隧道机制

大型语言模型在数值估计方面表现优异,但难以正确量化不确定性。我们研究了大模型构建自身答案置信区间的准确性,发现其存在系统性过度自信。为此,我们引入FermiEval,一个包含费米式估算问题的基准,并采用严格的置信区间覆盖率与精度评分规则。在多个现代模型上,名义上的99%置信区间平均仅覆盖真实值65%。通过基于符合预测的方法调整区间,实现准确的99%观测覆盖率,且Winkler区间得分降低54%。我们还提出了直接对数概率诱导与分位数调整方法,在高置信度下进一步缓解过度自信。最后,我们提出感知隧道理论:当模型在不确定条件下推理时,其行为如同从推断分布的截断区域采样,忽略了分布尾部。

原文摘要 · Abstract (English)

Large language models (LLMs) excel at numerical estimation but struggle to correctly quantify uncertainty. We study how well LLMs construct confidence intervals around their own answers and find that they are systematically overconfident. To evaluate this behavior, we introduce FermiEval, a benchmark of Fermi-style estimation questions with a rigorous scoring rule for confidence interval coverage and sharpness. Across several modern models, nominal 99\% intervals cover the true answer only 65\% of the time on average. With a conformal prediction based approach that adjusts the intervals, we obtain accurate 99\% observed coverage, and the Winkler interval score decreases by 54\%. We also propose direct log-probability elicitation and quantile adjustment methods, which further reduce overconfidence at high confidence levels. Finally, we develop a perception-tunnel theory explaining why LLMs exhibit overconfidence: when reasoning under uncertainty, they act as if sampling from a truncated region of their inferred distribution, neglecting its tails.

大模型不确定性置信区间校准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。